Skip to content
All posts
5 min readtebakopackagingbenchmarksruby

Tebako runtimes under the bench: MRI vs TruffleRuby vs JRuby, and what the JITs actually buy you

The tebako project has measured five Ruby runtime arms through the real tebako dispatch path — MRI 3.3.12, MRI 4.0.6 with YJIT and ZJIT, TruffleRuby native and JVM, and JRuby on OpenJDK 25 — on synthetic kernels and on a real Metanorma document compile. The result is not a single fastest engine; the fastest engine depends on the shape of the process, and the packaging layer is where that choice belongs.

The Tebako team

github.com/tamatebako

Tebako v2 treats the Ruby interpreter as a replaceable runtime line: the same payload can resolve against MRI, TruffleRuby (native or JVM flavor), or JRuby, and the loader neither knows nor cares which one starts (post 3 of this series). Machinery is only interesting once it is measured. This post is the measurement: five runtime arms, two synthetic kernels, one real standards- document compile, all through the real v2 dispatch path — shim → driver → runtime exe + mounted env image, no shortcuts, no host ruby.

All numbers are single-machine (MacBook Pro, Apple M1 Max, macOS 14), quiet-window gated, median-of-rounds; the protocol and the per-cell load guards are described with each experiment below. Absolute seconds are local to the machine; the ratios carry the meaning.

The synthetic ladder: who wins when the code gets hot

The ladder runs two kernels, warmed 3× then timed 7× inside each process, in three interleaved rounds per arm: fib34, a recursive method-call workload that favors a just-in-time (JIT) compiler, and alloc2m, two million object allocations that exercise the garbage collector (GC).

arm fib34 (vs MRI 3.3.12) alloc2m (vs MRI) peak RSS

MRI 3.3.12 (plain)

1.00× (0.507 s)

1.00× (0.269 s)

63 MiB

MRI 4.0.6 (plain)

0.97× (0.495 s)

1.05× (0.282 s)

~same

MRI 4.0.6 +YJIT

0.16× (0.082 s)

1.01× (no gain)

~same

MRI 4.0.6 +ZJIT

0.89× (0.451 s)

1.01× (no gain)

~same

TruffleRuby 34.0.1 native

0.14× (0.067 s)

0.30× (0.080 s)

590 MiB

JRuby 10.1.1 on OpenJDK 25

0.25× (0.126 s)

0.49× (0.133 s)

828 MiB

Three readings fall out of this table.

The version bump is free. Plain 4.0.6 ≈ plain 3.3.12 (0.97× on the call kernel, 1.05× on allocation — noise-class). Moving the interpreter line forward is not a performance decision either way; it is a support-window decision.

Hot code converges to the same class, whatever the JIT. YJIT on 4.0.6, TruffleRuby native, and JRuby after tier-up all land between 0.14× and 0.25× of plain MRI on fib34 — a 4–7× speedup that is the argument for JITted runtimes on long-lived, CPU-bound processes. YJIT holding its own against the GraalVM family (0.16× vs 0.14×) is the genuinely notable cell: a payload no longer needs a 600-MiB runtime to obtain most of the converged-class win.

ZJIT is real but young. On 4.0.6 it gains ~9% inside the short-horizon protocol — honest, repeatable, and nowhere near YJIT’s class yet. This matches upstream’s own posture exactly: Ruby 4.0 ships ZJIT compiled in but off by default, YJIT remains the production compiler, and ZJIT is the long game. Tebako mirrors that posture: neither JIT is enabled by default in its runtimes; both are one environment variable away (RUBY_YJIT_ENABLE=1 / RUBY_ZJIT_ENABLE=1).

The real-workload finding: the JIT can cost you

Synthetic kernels favor JITs. A real Metanorma compile — the ISO/OIML r060 document to HTML, with fonts and the Java conversion pipeline included — is where that advantage is priced. The table lists clean cells only (per-cell load guards):

arm boot (median) r060 HTML compile

MRI 3.3.12

4.2 s

47.0 s — fastest correct arm

MRI 3.3.12 +YJIT

4.0 s (neutral)

71.9 s (1.5× slower)

TruffleRuby native

34.7 s (8.2×)

~275–287 s, output complete, then a teardown crash†

TruffleRuby JVM

41.4 s (9.8×)

died deep in-run (SIGILL‡)

† oracle/truffleruby#4457 — the process produces correct output and then dies in thread teardown, in three of three runs. ‡ the crash is a deterministic illegal instruction on this workload; see the TruffleRuby feedstock.

The headline finding is that on a ~45-second real compile, YJIT makes MRI slower, not faster. It burns roughly ten CPU-seconds JIT-compiling blocks whose amortization the process exits before reaching. Ruby’s own default posture — YJIT enabled for long processes, with no claim that it helps short ones — is vindicated by a real workload rather than by a microbenchmark. The GraalVM variants, meanwhile, pay an 8–10× boot cost before they execute a line of application code; they are engines for persistent services, not for one-shot compiles.

This is the answer the packaging layer was designed around, and the bench supports the design: MRI stays the default runtime; variants are an opt-in selector; JIT enablement is a packaging-time decision. With the interp_env manifest key (merged this week in tebako#560), a payload packager whose workload is long and hot — a relaton sync service, a render farm worker — can ship RUBY_YJIT_ENABLE=1 as the payload’s declared environment, and the user can still override it locally. The packager measures; the user controls; the loader stays ignorant.

The honesty section: what is not green yet

Benchmarks that only report wins are marketing. Three caveats travel with every number above:

  • TruffleRuby’s spawned-subprocess support is broken today. Every host-binary spawn from a TruffleRuby runtime with a payload mounted fails (ENOENT on posix_spawn /bin/sh) — Metanorma’s PDF stage, which shells out to Java, has never completed on TruffleRuby under tebako. This is tebako#553, the TruffleRuby line’s top blocker, and the fix is per-interpreter: the shell-form spawn bridge that cured this class for MRI (runtime 0.16.20+) does not cover TruffleRuby’s own process operations.

  • One earlier result was wrong, and the record says so. An earlier window reported TruffleRuby PDF cells as "rc 0, did not reproduce" — incorrect, because Metanorma’s worker pool swallowed the exception in a thread and exited 0 with no PDF produced. The team corrected the record on the issue and filed metanorma#601 upstream (a failed output stage must not exit 0). The bench rule is now artifact- present and trace-free, never the exit code alone.

  • JRuby’s convergence is longer than the protocol. Tier-up was still visibly in progress at rep 7; a longer-horizon JRuby cell is queued, not claimed.

What is next

  • Python joins the bench. The Python runtime factory shipped 3.11.16/3.12.14/3.13.15/3.14.7 last week, plus CPython JIT flavor runtimes (3.13.15-jit, 3.14.7-jit — PEP 744’s experimental-in-3.13, production-track-in-3.14 tier-1 JIT; musl legs skip loudly because upstream’s own matrix excludes them). The same kernel-and-real-workload protocol runs against those four arms next window.

  • A longer-horizon ZJIT cell is queued, together with a convergent-workload showcase (a persistent service loop), which is the honest shape for GraalVM-family claims.

  • A TruffleRuby re-verdict is due once tebako#553 and the two oracle filings land.

The point of all five arms is not to crown a winner. It is to prove the platform’s claim with instruments instead of adjectives: any runtime, any payload, measured the same way, chosen at packaging time, and overridable at run time. The runtime is a line, and the packager picks the line.