Per-Ractor GC (Feature #22227) is a big win. On a 16-thread box, JSON parsing in 16 Ractors went from 13.6x slower than one Ractor to 2.2x. Fork gets 2.9x. While benchmarking it I found that a single Ractor got slower again on later master. Fork and the main Ractor didn't.
abac51a1ce, "gc: give a Ractor's own objspace a smaller initial heap size". Part of PR #18806, merged 2026-09-13.
| Nightly | Revision | GCs inside the Ractor |
|---|---|---|
| 2026-08-07 | 66c8e49a18 | 213 |
| 2026-09-13 | 5a02ff8481 | 125 |
| 2026-09-14 | c4e06b4e30 | 4,001 |
| 2026-10-09 | 28050861df | 4,001-7,999 |
I bisected the nightlies down to one day, then built the commit, its parent, master, and master with the change undone:
| Build | GCs in Ractor | Heap pages | JSON, 1 Ractor | Producer/consumer, 1 Ractor |
|---|---|---|---|---|
| parent ab51d15bc3 | 123 | 102 | 0.625s | 0.359s |
| abac51a1ce | 3,084-5,330 | 6-8 | 0.689s | 0.454s |
| master 28050861df | 3,201-4,002 | 6 | 0.667s | 0.583s |
| master, change undone | 122 | 102 | 0.593s | 0.345s |
Source builds lose 10-12% on JSON. The prebuilt --enable-shared --enable-yjit binaries
(ruby-dev-builder, Docker nightlies) lose 30-50% with the same GC counts. Their GCs cost more
each. I haven't tracked down why.
A GC runs when the heap fills up, so the GC count is roughly allocations divided by heap size.
Before this commit every Ractor's heap started at RUBY_GC_HEAP_INIT_BYTES (2.5 MB, about 100
pages). Now a Ractor starts at one slot and grows only when a sweep leaves less than 20% of the heap
free. If everything the Ractor allocates dies young, every sweep frees almost all of it, so the heap
never grows. It sits at 5-8 pages, which is smaller than what one parse of a 30 KB JSON document
allocates. The Ractor ends up collecting once or twice per parse.
The main Ractor isn't affected because the free-slot check uses the initial size as a floor
(MAX(total_slots, init_slots) in gc_sweep_finish_heap), and main still gets 2.5 MB. Workloads
that keep a live set (binary trees) are fine because the live set makes the heap grow.
The commit's goal makes sense. With per-Ractor GC, 64 Ractors would each keep a 4.3 MB heap. Its benchmarks either allocate small garbage or hold a live set. A job worker that parses payloads does neither: it has a big working set per unit of work, and all of it is garbage.
The commit added RUBY_GC_RACTOR_HEAP_INIT_BYTES. Setting it on master (28050861df):
RUBY_GC_RACTOR_HEAP_INIT_BYTES |
GCs in Ractor | Heap pages | Ractor wall | Ractor/main |
|---|---|---|---|---|
| default (one slot) | 4,001-7,999 | 7 | 0.78-0.96s | 1.13-1.53 |
| 65536 | 2,002-7,998 | 8-10 | 0.70-1.02s | 1.03-1.51 |
| 262144 | 1,333 | 13-14 | 0.69s | 1.00-1.09 |
| 1048576 | 296 | 45 | 0.63s | 0.86-0.92 |
| 2621440 (main's value) | 125 | 102 | 0.63s | 0.94-1.01 |
1 MB recovers all of the time with less than half the old heap (45 pages vs 102).
- A larger default, somewhere from 256 KB to 1 MB. 64 idle Ractors at 1 MB is 64 MB, against 275 MB at the old 2.5 MB.
- Grow a Ractor's heap when GCs come too close together (few allocations between collections), not only when a sweep frees too little. That keeps idle Ractors small and lets busy ones size themselves.
- Short term, document
RUBY_GC_RACTOR_HEAP_INIT_BYTESas the knob for allocation-heavy Ractors.
ractor_gc_repro.rb below runs the same JSON work on main and in one Ractor and reads GC.stat
inside each. The GC count tells you whether a build is affected; timing isn't needed.
ruby ractor_gc_repro.rb
RUBY_GC_RACTOR_HEAP_INIT_BYTES=1048576 ruby ractor_gc_repro.rb
Prebuilt master binaries are at ruby/ruby-dev-builder. On an M2 with the 2026-10-09 build:
RUBY_GC_RACTOR_HEAP_INIT_BYTES=(default)
round 1 ractor/main = 1.22
main 0.755s 115 GCs 135 heap pages
ractor 0.921s 4000 GCs 5 heap pages
RUBY_GC_RACTOR_HEAP_INIT_BYTES=1048576
round 2 ractor/main = 0.98
main 0.825s 147 GCs 112 heap pages
ractor 0.812s 291 GCs 45 heap pages
Box for the tables above: AMD Ryzen 7 PRO 8700GE (8 cores, 16 threads), Ubuntu 24.04, gcc 13.3, json 3.0.2, medians of 5 interleaved rounds.