Four bytes of object metadata and the costs around them

The current Linux ARM64 performance baseline has ParparVM translating its own compiler in 37.5% less elapsed time and 45.8% less peak memory than JDK 25. Across all 16 recorded machine configurations, self-translation takes 22.5% to 56.6% less elapsed time. The gains vary by OS and processor; the chart and table include every configuration. Your Java source does not need to change to benefit from a smaller runtime and a compiler that removes more work.

Self-translation elapsed time and peak memory relative to JDK 25 on all sixteen recorded machine configurations; lower is better

These are the repository’s checked-in performance baselines, read on October 1. They are recorded reference results used to detect regressions, not new measurements made for this article. Each comparison runs the same translator classes on the two runtimes and verifies the generated files byte for byte.

Self-translation baselineElapsed time, ParparVM / JDK 25Peak memory, ParparVM / JDK 25Recorded calibration runs
Linux ARM64, Neoverse N20.6250.5427
Linux x64, EPYC 77630.4840.5193
Linux x64, EPYC 9V450.4950.5391
Linux x64, EPYC 9V740.5160.5222
Linux x64, Xeon 6973P-C0.5370.5201
Linux x64, Xeon Platinum 8370C0.4340.5093
Linux x64, Xeon Platinum 8573C0.5050.5171
macOS ARM640.5390.7391
Windows ARM64, model D490.7510.5397
Windows ARM64, model D840.7750.5701
Windows x64, AMD family 25/model 10.6060.5427
Windows x64, AMD family 25/model 170.6360.5512
Windows x64, AMD family 26/model 20.6110.5461
Windows x64, Intel family 6/model 1060.6270.5361
Windows x64, Intel family 6/model 1730.6770.5491
Windows x64, Intel family 6/model 2070.6020.5261

Lower is better in both columns. JDK 25 is 1.0. Each baseline is the median of the recorded runs’ median ratios, with each platform using its available CPUs and runtime defaults. Translation time includes the whole process; Linux and macOS memory is peak RSS; Windows memory is peak working set. Some configurations have only one calibration run, as shown. They are separate machines, not a core-count scaling test.

A couple of weeks ago I wrote about making ParparVM compile itself. HotSpot beat us on elapsed time. That was disappointing, though hardly surprising. Since then we have worked on object layout, collections, generated C and garbage collection. The results above are a much better place to be.

The smaller objects matter beyond the compiler benchmark. In Friday’s server comparison, CN1’s native backend peaks at 57.8 MiB RSS, against 117.0 MiB for Spring/GraalVM, and rests at 29.8 MiB after load against 79.6 MiB. That compares complete HTTP stacks. Here we will look at the runtime changes that help make a smaller Java process possible.

What gets faster, and what still needs work?

Self-translation is a substantial Java workload: allocating objects, walking collections, resolving methods and writing generated C. It is also our own compiler, so we should check other workloads before treating it as a verdict on Java performance.

The complete Linux ARM64 baseline contains 13 workloads. ParparVM uses less peak memory in all 13 and less elapsed time in eight. Sequential array access takes 0.362 times the JDK’s measured time; hash-map churn uses 0.121 times its peak process memory. Allocation remains an important weakness: that workload takes 3.043 times the JDK’s time, even while using 0.239 times its peak memory.

All thirteen Linux ARM64 baseline workloads, showing elapsed time and peak memory as ratios to JDK 25; lower is better

Each row combines seven calibration runs; each run compares the two runtimes in alternating order. Translation rows measure the complete process. The smaller workloads time repetitions inside the process after their own warmup, so JVM startup does not count against HotSpot. Memory is whole-process peak RSS in every row, not live heap.

The performance harness checks workload results as well as time and memory. The original self-hosting post measured 1.56 seconds for HotSpot against 1.84 seconds for ParparVM. The opening charts show the later checked-in baselines after the optimization work; they are separate measurements, not a reinterpretation of that earlier result.

Elapsed time and CPU time also answer different questions. Parallel JIT compilation and collection can reduce the wait while consuming more CPU across threads. That distinction helped guide the investigation, but it does not establish a battery saving: energy depends on the machine’s power use over time. The current baseline records elapsed time and peak memory, so those are the claims we can make from it.

Why a tiny Java object needs a header

An object header is runtime metadata stored alongside your fields. It lets the runtime identify the object’s class and track information needed for collection and other operations. On a typical 64-bit HotSpot configuration, the header occupies 12 bytes; without compressed class pointers it can occupy 16. JDK 25’s optional compact object headers reduce that to 8 bytes. Oracle’s JDK 25 GC guide documents the sizes and the -XX:+UseCompactObjectHeaders switch.

This is the work associated with Project Lilliput. Project Leyden addresses startup, warmup and footprint through ahead-of-time work, including AOT caches. The names are easy to mix up, but shrinking each object’s header and caching class-loading work solve different problems. See the Leyden project for that distinction.

Runtime layoutObject header metadataWhat the number excludes
HotSpot, compressed class pointers12 bytesFields, alignment and backing storage
HotSpot, JDK 25 compact headers enabled8 bytesFields, alignment and backing storage
ParparVM release layout4 bytesFields, alignment, side tables and backing storage

These are layout sizes, not a prediction of process memory. A small HTTP server can spend more of its resident memory on stacks, code, allocator pages and library state than on live Java object headers. The server comparison measures the whole process separately.

The release object header carries a 16-bit class index, an 8-bit collection epoch and an 8-bit heap state. The class index addresses a closed-world class table instead of storing a pointer in every object. Objects requiring a legacy heap index use a side table keyed by address.

Four-byte ParparVM metadata: a two-byte class index, one-byte mark epoch and one-byte heap state, with class information held in a shared table

The runtime header is explicit about the distinction between metadata and allocation. The C header struct remains eight-byte aligned. Small objects live in allocator size classes, and fields need alignment too. A four-byte header does not make an empty object a four-byte allocation.

The translator packs eligible fields into the gap after the metadata. It stores fields at their real widths and orders them to reduce padding. A subclass must preserve its superclass layout as a prefix, so it cannot arbitrarily reach back and fill a hole belonging to an ancestor. That restriction sharply reduced the savings predicted by an early census.

Object in the experimentBeforeAfter four-byte header and root-field packing
VarOp32 bytes24 bytes
ArrayList32 bytes24 bytes
ByteCodeMethodArg32 bytes24 bytes
BasicInstruction40 bytes32 bytes

Object sizes from the layout experiment. Backing storage is additional.

The header-layout experiment measured one-core container RSS moving from 618 MB to 605 MB across five interleaved runs. That is useful, but much smaller than “we halved the header” might suggest. The page allocator and native collection buffers still occupy memory.

There was a second trap. A static class table that named every class also kept those classes reachable to the native linker. The current implementation registers most entries as classes are used. Shrinking object metadata should not accidentally force otherwise unused classes into the executable.

Get less overhead from the same Java loop

ParparVM translates bytecode to C. “Lowering” is the step that turns a higher-level operation into a representation the native compiler can optimize directly. We do not get that benefit merely by giving Clang a large pile of C functions.

Consider a collection traversal. The ordinary implementation has an iterator object and virtual calls for hasNext() and next(). When the translator proves the exact collection layout, it can emit a native cursor loop. When it cannot prove the receiver, it can check the exact class once and retain the ordinary iterator as a fallback.

flowchart TD Loop[Collection traversal] --> Proof{Exact supported layout?} Proof -->|Proved locally| Cursor[Native cursor loop] Proof -->|Unknown receiver| Guard[Check exact class once] Guard -->|Matches| Cursor Guard -->|Other class| Iterator[Ordinary iterator] Cursor --> Rules[Keep owner alive and preserve modification checks]

The owner must remain a GC root until the final native-buffer access. A raw pointer into a malloc block does not keep the Java owner alive. The emitted lifetime fence is therefore part of correctness, not optional bookkeeping to strip from a fast loop.

Several related changes attack allocation and dispatch:

  • Collection storage: hash tables combine key, value and metadata slices into one native allocation. Reference slices still need tracing and write barriers. Moving them outside the managed heap does not make their Java references invisible to the collector.
  • Strings and builders: StringBuilder keeps Latin-1 bytes until it needs UTF-16. Proven temporary builders can use native storage while preserving ownership across helper calls and exceptions. StringBuffer retains its synchronization.
  • Short string equality: bounded loads compare short strings without an ordinary Java method call. Longer ranges still use memcmp; mixed encodings preserve their own comparison rules. The implementation must not read past the logical range just because a wide load would be convenient.
  • Immediate array streams: supported Stream.of(array) pipelines with known lambda callbacks can become one C loop. Escaping streams, unknown callbacks, sorted, distinct and unsupported terminals keep the lazy implementation.

For example, this shape is eligible for the stream work when the surrounding method satisfies the proof:

long count = java.util.stream.Stream.of(names)
        .filter(name -> name.length() > 3)
        .count();

Here names is a String[]. Eligibility is narrower than “streams allocate nothing.” Nor does seeing a loop in emitted C prove that the linked binary vectorizes it. The lowering audit checks generated code and linked instructions separately from elapsed-time measurements.

Keep your program correct when removing work

A declared List type does not prove an ArrayList layout. LocalReceiverTypes tracks allocation provenance, and unknown parameters or factory results retain guards. Analysis is restricted to methods that need it, and the compiler drops the temporary frame information after capturing the proof. Optimizing the translated program should not require keeping an unnecessary analysis graph alive throughout translation.

Direct calls also gain small C intrinsics for operations such as list access, string length, hashing and builder appends. They apply to calls whose target is resolved statically, with the ordinary implementation available off the fast path. Together with traversal lowering, this lets the C compiler see useful operations inside a loop instead of opaque calls at every step.

Another pass removes instance fields the closed program never reads. It checks native-source references and conservatively keeps runtime fields and unresolved cases. Removing a dead reference field can save both its storage and the object it would otherwise retain. This pass is disabled for JavaScript output and on-device debugging.

There is a semantic edge worth stating: the current dead-field pass documents that assigning to a removed field on a null receiver no longer produces the JVM’s NullPointerException. Its treatment of an unused store is therefore not a claim of identical behavior for every Java program. An application depending on that exception needs to account for this difference.

Fit the collector to the cores you have

The small-object allocator groups objects into pages by size, often called BiBOP, short for “big bag of pages.” A nonmoving collector cannot return a page while one live object still occupies it. Free slots inside a partial page can be reusable capacity without being releasable memory.

The experiments found 50 to 60 MB of free slots inside partial pages around one peak. Those slots were reused before new pages were allocated. Calling all of that a leak would send the investigation in the wrong direction.

On one core, the generational collector can stop application work and collect young objects. With spare cores, the concurrent collector overlaps tracing with the application. The concurrent path needs grace for allocations that appear after a snapshot. That can retain dead objects longer, and shortening that lifetime requires preserving everything the snapshot could not yet see.

We tried reclaiming dead large objects every single-core cycle. The first attempts were unsound. A growing output buffer could be allocated after the extent snapshot and remain live only in a thread’s local state. If the collector could not resolve that pointer, it could free a live buffer.

The fix registers those large objects from allocation and rebuilds the relevant snapshot at each thread’s pause in that mode. The concurrent collector keeps its different grace rules. Applying one mode’s shortcut to the other added locking to a hot path without providing a benefit.

One-core translation, five interleaved roundsMinimum elapsedMaximum RSS
JDK 2513.74 s557 MB
Four-byte header before reclamation change10.29 s608 MB
With large-object reclamation change9.84 s503 MB

A separate Linux container experiment isolating the reclamation change: minimum elapsed and maximum RSS from five interleaved runs. Lower is better. Its four-byte-header control peaked at 608 MB; the header-layout experiment reported 605 MB in a different set of runs. The reclamation experiment record describes this later comparison. The opening charts show the current baseline.

We then reduced work in minor collections: skip an already-old page object before expensive conservative resolution, and consult the monitor table only for pages that have had monitors. The experiment record includes the unsuccessful alternatives too. Fewer object bytes help memory use. Avoiding thousands of unnecessary collector operations helps execution time too.

What changes for your app?

If your application creates large graphs of small objects, the smaller headers and field packing can reduce their cost without a source rewrite. Proven collection traversals and stream pipelines can avoid temporary objects and dispatch. These are compiler and runtime changes, so you get them by rebuilding with a release that includes this work.

The benefit depends on what your app spends time and memory doing. A form with a large model, an importer that walks collections, and a service that allocates objects continuously stress different parts of the runtime. The allocation result above is a reason to check a busy handler before assuming it will behave like the compiler benchmark.

For an existing Codename One app, rebuild and compare a real interaction on your target device: startup, a large list or model load, and a sustained operation that creates objects. For a backend, compare resident memory at rest and under representative traffic, plus throughput and latency. Keep the input and build settings the same between releases. The backend guide covers packaging a small service you can test this way.

There is more room to improve. Our measurements point toward shorter-lived garbage in the concurrent collector and less overhead in small native collection buffers. A stop-the-world strategy saves memory on one core, but forcing it everywhere can lose too much execution time. The next step is to reduce that retention while preserving useful concurrency. That is an optimization direction, not a promise that one collector will win on every application.


Discussion

When you compare runtimes, do you record CPU time as well as the time you waited?