Python 3.15 tracing JIT unlocks 11% performance gain on AArch64.

Sunday 11 October 2026, 02:04 PM

Python 3.15 tracing JIT unlocks 11% performance gain on AArch64.

Python 3.15 adds a tracing frontend to its experimental JIT compiler, delivering an 11% performance boost over the tail-calling interpreter on AArch64.


I spent the first week of October tracking the Python 3.15 release cycle. The unscheduled RC3 drop on October 2 caught my attention, but it turned out the delay wasn't JIT-related. It was a last-minute blocker with PEP 810's explicit lazy imports. When the final build shipped on October 9, I immediately pulled it down to look at the experimental Just-In-Time compiler.

We are finally moving past the initial copy-and-patch proof of concept introduced in 3.13. Under issue gh-139109, core developer Ken Jin contributed a new tracing frontend. The broader compiler community largely abandoned tracing JITs years ago, but Python operates under unique stack-machine constraints where this approach actually works. Instead of estimating execution paths, the frontend actively records them at runtime to optimize hot loops. For those of us running heavy backend services, shifting from estimation to actual path recording changes how we approach bottleneck optimization.

To support this architecture—specifically the enhanced machine code generation and stencil creation—core contributor Savannah Ostrowski bumped the JIT's build-time dependency to LLVM 21 (issue gh-140973). Whenever I see a new LLVM requirement in a changelog, my immediate concern is container bloat and CI/CD pipeline overhead. Fortunately, LLVM remains strictly a build-time requirement here. We don't need to ship it in our production images, which keeps the deployment footprint lean.

Under the hood, the execution speed improvements come down to basic register allocation, reference-count elimination, and in-place numeric operations. The team also wired up native GDB unwinding support, which is a relief for debugging optimized code in staging. I ran the numbers on my local Apple Silicon machine, and the benchmarks align with the official claims: an 11 to 12% geometric mean speedup over the tail-calling interpreter on AArch64 macOS. On x86-64 Linux, it hovers around 7 to 8%. In an ecosystem where AWS Graviton and Apple M-series chips dominate our infrastructure, squeezing an extra 11% out of ARM architectures translates directly to lower cloud compute bills at scale. The runtime manages this without breaking the extensive C-API ecosystem.

Optimizing code is only half the battle; we also need to measure it. Python 3.15 pairs these JIT updates with Tachyon (PEP 799), a new zero-overhead sampling profiler. Profiling async Python in live environments usually means accepting a performance penalty. Tachyon bypasses this by reconstructing async tasks and opcode-level execution without modifying the running process. I plan to deploy this alongside the JIT to gather granular data on hot loops without degrading production throughput.

Right now, the JIT is still an opt-in experimental feature that requires compiling from source. But the trajectory is clear. Back in July 2026, the steering council drafted PEP 836—titled "JIT Go Brrr: The Path to a Supported JIT Compiler for CPython"—to propose making the JIT a fully supported, default component. As we prepare for that transition, I am keeping a close eye on our dependencies. Shifting the execution model exposes us to potential memory-management bugs in third-party C extensions. We also have to consider platform inequality, as non-Tier 1 architectures likely won't see these same optimizations.

For now, I recommend compiling 3.15 from source on an ARM instance. It gives you a practical look at how your specific workloads behave under the new tracing frontend before these architectural shifts become the default standard.


References

Subscribe to our mailing list

We'll send you an email whenever there's a new post

Copyright © 2026 Tech Vogue