CratonVM
Architecture
Architecture
This document describes the high-level architecture of CratonVM. If you want to contribute, this is the place to start.
Crate Layout
The workspace has 22 member crates (the fuzz harness is a separate,
standalone workspace, not a member):
cratonvm/
reader/ cratonvm-reader .class file parser
types/ cratonvm-types Shared types (Value, ClassId, ObjectRef)
native-api/ cratonvm-native-api Native capability facade & FD table
native-builtins/ cratonvm-native-builtins java.lang.*, registration & marshalling
native-builtins-crypto/
cratonvm-native-builtins-crypto
Crypto compatibility kernels
native-builtins-security/
cratonvm-native-builtins-security
JDK security & SunEC native pack
native-collections/ cratonvm-native-collections java.util.* native methods
native-io/ cratonvm-native-io java.io/nio native methods
native-awt/ cratonvm-native-awt AWT/Swing/Java2D native peers
jit-api/ cratonvm-jit-api JIT compiler API types
jit/ cratonvm-jit x86-64 / AArch64 JIT compiler
jit-cuda/ cratonvm-jit-cuda Java bytecode -> PTX lowering for GPU offload
cuda-bridge/ cuda-bridge Thin CUDA Driver API bridge for GPU offload
craton-gpu/ craton-gpu Build-time-only: packages the @GpuKernel/@Parallel Java annotation sources into a jar for jit-cuda's build script; no runtime code
classloading/ cratonvm-classloading Class loading & bytecode verification
gc/ cratonvm-gc Garbage collectors (default generational semi-space; G1 region-based; ZgcRealHeap STW mark-sweep behind the default-off `zgc` feature)
jfr/ cratonvm-jfr Java Flight Recorder
vm/ cratonvm-vm VM runtime engine
vm-cli/ cratonvm-cli CLI entry point
libcratonvm/ libcratonvm C-ABI shared library for embedding (cdylib/staticlib libjvm substitute)
cratonvm-embed/ cratonvm-embed Semver-stable Rust facade for embedding CratonVM
difftest/ cratonvm-difftest HotSpot differential-testing harness
The fuzz/ directory is its own standalone workspace (cratonvm-fuzz, nightly-only libFuzzer harness) and is not a member of this workspace — its #![no_main] fuzz_target! expansion trips the production lints, so it builds separately via cargo +nightly fuzz build.
Dependency flow:
vm-cli -> vm -> {classloading, gc, jit, native-builtins, native-collections,
native-io, native-awt, jfr}
-> {reader, types, native-api, jit-api}
-> (gpu-offload feature only) {jit-cuda, cuda-bridge}
-> also forwards gc/gpu-offload, native-builtins/gpu-offload
native-builtins -> {native-builtins-crypto, native-builtins-security}
libcratonvm -> vm (C-ABI / JNI Invocation API embedding shim)
cratonvm-embed -> vm (curated, semver-stable Rust embedding facade)
craton-gpu does not appear above: it is a build-dependency of
jit-cuda only (its build.rs compiles the GPU annotation sources and
exposes their jar path via Cargo's links metadata), never a runtime
dependency of anything.
reader — Class File Parser
Parses .class files per the JVM specification (JVMS Ch. 4).
reader/src/
class_reader.rs Entry point: bytes -> ClassFile
buffer.rs Binary cursor (big-endian reads)
constant_pool.rs 20 CP entry types
instruction.rs 200+ bytecode opcodes
attribute.rs 30+ attribute types
stack_map.rs StackMapTable verification frames
field_type.rs Field descriptor parsing (e.g. "Ljava/lang/String;")
method_descriptor.rs Method descriptor parsing (e.g. "(II)V")
class_access_flags.rs Bitflag types for access modifiers
class_file_version.rs Major/minor version constants (Java 1.1-25)
The reader is a pure parser with no VM dependencies. It can be used
independently to inspect .class files.
vm — Virtual Machine
The VM is the core of the project (~1,350,000 Rust LoC across the 22 workspace
member crates, plus the separate fuzz harness workspace).
It contains six major subsystems (several now extracted into their own
crates).
That figure is the raw line count of every .rs file under the 22 directories
named in the root Cargo.toml [workspace] members list, excluding target/,
excluding the non-member fuzz/ workspace, and excluding vendored third-party
sources under any vendor/ directory (e.g.
native-builtins/vendor/rustls-cbc). Reproduce it with:
find <the 22 member dirs> -name '*.rs' -type f \
-not -path '*/target/*' -not -path '*/vendor/*' -print0 \
| xargs -0 cat | wc -l
which reports roughly 1,350,000 lines across about 700 files.
Rough size distribution, largest first, so newcomers know where the mass actually is:
| Crate | LoC | Crate | LoC |
|---|---|---|---|
native-builtins | 552,000 | native-awt | 18,000 |
vm | 343,000 | types | 17,000 |
jit | 110,000 | native-api | 17,000 |
gc | 64,000 | reader | 14,000 |
classloading | 57,000 | jfr | 14,000 |
native-collections | 55,000 | jit-cuda | 10,000 |
native-io | 55,000 | remaining 9 | < 7,000 each |
Several individual files are far larger than is comfortable. The two worst have been split at the section banners they already carried:
vm/src/runtime/interpreter.rswent ~50,500 → ~24,100 lines, withinterpreter/typecheck.rs(checkcast/instanceof/aastorecompatibility),interpreter/constants.rs(ldcand loader-faithfulCONSTANT_Classresolution),interpreter/field_access.rs(field resolution and invoke argument plumbing), andinterpreter/invoke.rs(method resolution, dispatch, and the native bridge — ~23,200 lines, still the single largest thing here because method invocation genuinely is one subsystem).jit/src/x64.rswent ~44,200 → ~36,200 lines, with each optimization pass in its own module:x64/licm.rs,x64/licm_int.rs,x64/bce.rs,x64/escape_analysis.rs,x64/null_check_elim.rs,x64/simd_analysis.rs,x64/bytecode_compat.rs,x64/reg_encoding.rs,x64/switch_validation.rs, andx64/cpu_features.rs. What remains is the emitter and the compilation entry point, which are not a clean seam.
Nothing moved between modules and nothing became more public than it was: each
child is mod x; pub use x::*;, and a glob re-export caps every item at its own
declared visibility. Note that hot_files_have_no_production_panics enumerates
both directories from disk — a gate that kept scanning only the parent would
have turned this split into "silently stopped checking most of it".
vm/src/vm.rs appears to dwarf both at ~74,200 lines, but that number is
misleading and this is not the file to start reading. Everything from
line 62 onward is a single #[cfg(all(test, feature = "synthetic-jdk"))] mod tests — a test module the default build does not even compile, because
synthetic-jdk is deliberately not a default Cargo feature (vm/Cargo.toml).
The orchestrator itself is the ~61-line module header above it, which declares
and re-exports vm/src/vm/: vm_exec.rs (~21,900 lines), vm_init.rs
(~12,300), vm_util.rs (~4,600), vm_object.rs (~2,100), and realms/.
Those are the files to open.
Two native-builtins files were worse and have also been split:
lib.rs went ~90,000 → ~39,000 across 13 per-domain
modules (util_concurrent_ext, antlr_intrinsics, regex_matcher,
math_bignum, …), and phases_late.rs went ~77,000 → ~8,000 across 18
modules under native-builtins/src/phases_late/ (nio_file, bouncycastle,
concurrent, ssl_security, net_channels, streams, jdbc, …).
phases_late.rs itself now holds only the shared preamble, the per-phase
dispatchers, and cross-domain leftovers.
Splitting a file does not speed up incremental builds, and it cannot. Rust's
compilation unit is the crate, not the file: moving code into modules of the
same crate leaves that crate — and everything downstream of it — rebuilding in
full. Measured over the 90k→39k split above, touch lib.rs && cargo build --release -p cratonvm-cli went 1m58s/2m00s before to 2m03s/2m03s after, i.e.
noise. The payoff of splitting is merge-conflict surface and reviewability
(the two worst files are 57% and 89% smaller, and edits now land in 31 separate
files instead of colliding in two), not build time.
An actual incremental-build win requires separate crates. That work has
started at the heaviest stable seams: cryptographic kernels and JDK
security/SunEC code now compile as native-builtins-crypto and
native-builtins-security. The facade crate retains registration and
Java-object marshalling; other domain modules remain candidates only when a
measured rebuild or ownership benefit justifies another crate boundary.
Runtime (vm/src/runtime/)
The bytecode execution engine.
-
interpreter.rs— Main dispatch loop (~24,100 lines, plus theinterpreter/submodules listed above). Each opcode reads operands, manipulates the operand stack and local variables, and advances the program counter.There are two cooperating dispatch paths:
- the common fast path reads opcodes from verified raw bytecode and fuses
common sequences (for example
iload_X; iload_Y; iadd) as superinstructions; and - unsupported opcodes or guarded edge cases fall through to the shared decoded handler.
Package names are no longer execution-policy inputs: identical verified bytecode takes the same common path whether it belongs to an application, the JDK, or a framework.
--noverifydisables unchecked raw-byte handlers and uses the bounds-checked decoded path. A one-time quickening pass stores aQuickenedCode(reader/src/quickened.rs), whose bitmap/rank index maps a bytecode PC to a decoded instruction in O(1), including taken branches. - the common fast path reads opcodes from verified raw bytecode and fuses
common sequences (for example
-
frame.rs— Stack frame: local variables and operand stack use 8-byte NaN-boxedCompactValueslots. A parallel byte array is retained for the ambiguous rawlong/doublelocal cases; tags are otherwise inline. -
call_stack.rs— Per-thread call stack of frames. -
value_stack.rs— Typed operand stack (SoA encoded). -
exceptions.rs— Java exception creation and throw handling. -
invokedynamic.rs— Lambda/method-ref bootstrap via LambdaMetafactory.
Typed Bootstrap (vm/src/vm/vm_init.rs)
Startup is encoded as consuming typestate transitions:
BootstrapPhase<Allocated>
→ BootstrapPhase<ClassesReady>
→ BootstrapPhase<NativesReady>
→ BootstrapPhase<RuntimeReady>
Each boundary validates the invariant it owns: a nonempty class universe with
java/lang/Object, a nonempty native registry, and wired runtime hooks.
Only RuntimeReady exposes finish. The token is deliberately small; large
subsystems remain locally owned during SharedVm::new, while the token prevents
accidental reordering and gives tests a precise failure boundary.
The end-to-end process lifecycle and the cross-subsystem safety invariants are documented in the manual's Runtime Lifecycle and Runtime Contracts chapters.
Class Loading (classloading/ crate)
Implements JVMS Ch. 5: loading, linking, and initialization. Extracted into the
cratonvm-classloading crate.
class_manager.rs— Central class cache and loading coordinator.loaders.rs— Bootstrap, extension, and application class loaders.class_path.rs— Classpath scanning (directories + JARs viazipcrate).class.rs— Runtime class representation with metadata and hierarchy.verifier.rs— Structural verification (JVMS 4.8-4.9).bytecode_verifier.rs— Type-checking verification (JVMS 4.10).vtype.rs— Verification type lattice.resolution.rs— Symbolic reference resolution with inline caching.access_control.rs— Access checking (JVMS 5.4.4).
Memory (gc/ crate)
Garbage collectors, extracted into the cratonvm-gc crate. The default is the generational collector with the Cheney moving young gen enabled — but see "When the default does not compact" below: requesting a moving cycle is not the same as running one, because each cycle must still carry a per-cycle root-coverage proof. A region-based G1 collector is also present and opt-in selectable via -XX:+UseG1GC (experimental; Generational remains the default safety net during G1 maturation — see docs/feature-designs/concurrent-gc-maturation.md). ZGC is feature-gated (zgc, not part of the default feature set) and holds two things of very different maturity. ZgcRealHeap (gc/src/zgc.rs:1396) is a real, memory-backed collector — Arena storage, real ObjectHeaders, real java.lang.ref processing — and it is wired end to end: GcAlgorithm::Zgc → GcBackend::Zgc → VmHeap::Zgc, so -XX:+UseZGC really selects it. Above it in the same file sits a metadata-only simulation of OpenJDK's colored-pointer model (ZgcCollector / ColoredPointer / LoadBarrier / GenerationalZgc) with no production consumer. What ZGC is not is present in a stock build: the feature is default-off, so GcAlgorithm::Zgc and its parse_gc_algorithm arm do not exist unless you build cargo build --release -p cratonvm-cli --features zgc, and -XX:+UseZGC otherwise warns and falls back to Generational. Nor is it production ZGC — it is stop-the-world, non-moving, non-generational, non-compacting, TLAB-less. It stays default-off for pass-rate parity: on the 1975-class Spring Boot suite, same binary with only the collector toggled, it measures 1860 PASS / 49 HANG / 22 FAIL against the default collector's 1902 / 18 / 11. The path to a real one is docs/feature-designs/zgc-production-implementation-plan.md.
ZgcRealHeapis wired into backend dispatch —gc/src/vm_heap.rsimports it andVmHeap::Zgchas real arms. What gates it is the default-off Cargo feature, restated above. The working build command is-p cratonvm-cli --features zgc;--features cratonvm-vm/zgcalso reaches the gated code.
heap.rs— Object/array layout and allocation (semi-space).gen_heap.rs— Generational heap: young gen (copying) + old gen.gc.rs— Cheney copying collector algorithm.collector.rs— GC coordination and stop-the-world orchestration.card_table.rs— Card marking for cross-generational references.arena.rs— Bump-pointer memory arena.roots.rs— Root scanning and pointer remapping.old_gen.rs— Old generation management.
When the default does not compact. The flag gate is gone; the per-cycle proof is not. Read both before reasoning about allocation-path or GC-pause code.
-
moving_youngis an opt-out flag that defaults to true —types/src/flags.rs::DEFAULT_MOVING_YOUNG, shaped as "CRATONVM_NO_MOVING_YOUNGturns it off,CRATONVM_MOVING_YOUNGis a retained no-op opt-in", and pinned byempty_source_matches_all_documented_defaultsandmoving_young_is_an_opt_out_with_a_compatibility_opt_in. A defaultcargo buildtherefore requests a moving young cycle. -
Requesting one is not running one: every cycle must carry a per-cycle coverage proof before it may relocate.
collect_garbage_innerdiverts to the non-moving sweep ondivert_for_incomplete_moving_coverage, whichvm/src/memory/roots.rscomputes fromconservative_roots::refresh_moving_young_coverage_for_collection()— every live compiled frame's active safepoint must certifymoving_young_coverage_complete, no unregistered JIT frame may be on the native stack, no peer thread may be in JIT, and a frame-band verifier must find no young-resident word the shadow stack did not publish. Each diversion is counted and logged atwarnwith the specific unproven obligation (gc_quiescence::incomplete_reason), and--verbose:gc/CRATONVM_GC_STATSprintmoving_young: cycles=N coverage_fallbacks=Mplus a per-reason histogram. -
Ask the runtime, not a document.
gc_metrics::collector_decision_report()(gc/src/gc_metrics.rs) renders the last collection's actual decision — backend, moving/non-moving verdict, the stable reason code, and the specific unproven obligation — andprint_gc_summaryemits it on every--verbose:gcrun. It is the authority for "did this process compact?"; prose about the flag is not.docs/GC.md's backend table is kept in step with that report; when the two disagree, the runtime report wins.(An earlier
fail_closed_non_moving = is_active() && !allow_moving_youngterm made this unconditional — it meantCRATONVM_MOVING_YOUNG=1alone could never run a moving cycle under a live JIT frame, the only case the feature exists for. That term and theCRATONVM_ALLOW_MOVING_YOUNGflag are gone.)
So on a JIT-warm workload a default build can still spend cycles in the
generational non-moving mark-sweep with selective promotion path — that is
what a nonzero coverage_fallbacks count means, and it is the number to read
before attributing a pause profile to compaction.
Correctness is no longer the blocker: the heap corruption was
five codegen sites pushing an untagged object reference onto the JIT's
simulated operand stack. Throughput is the remaining work, now tracked as
an optimization program rather than as a precondition, because the memory-
footprint win decided the flip: on Binary-Trees-18 the moving path measures
2.1× the sweep (six interleaved rounds, min 2905 ms vs 1363 ms), down from
5.0× before the pre-cycle object-start walk stopped building a hash set of
every object in from-space — but bt18 at -Xmx512m does not complete at all
without compaction. bt18 is simultaneously the worst case for a copying
collector and the workload that justifies it. The open items are the pointer_map
FxHashMap, the disabled self-call spill elision, and the unpriced
jit_frame_record helper.
Object layout:
[ObjectHeader (32 bytes)] [field0] [field1] ...
Field cell width depends on the field's type and on the layout in force
(types/src/heap_types.rs, types/src/field_layout.rs):
| Field kind | Width | Notes |
|---|---|---|
| Reference | 8 B | Bare pointer, 0 = null; 4 B when compressed oops are explicitly enabled. |
boolean / byte | 1 B | Tagless physical field. |
char / short | 2 B | Tagless physical field. |
int / float | 4 B | Tagless physical field. |
long / double | 8 B | Tagless physical field. |
CompactLayout computes aligned offsets and the precise GC reference map
(ref_offsets). The interpreter/native Value enum is a boundary type, not
the physical instance-field representation.
The 32-byte header (ObjectHeader, types/src/heap_types.rs) is the current
compatibility contract. A 24- or 16-byte header would require a separate object
model: folding forwarding state into the mark word couples collector relocation
to thin/inflated monitor state, while removing the identity hash alone saves no
space after alignment. Header compression is therefore an experiment, not an
unfinished requirement of the compact field layout.
Arrays use compact element sizes (1/2/4/8 bytes per element depending on type;
element_byte_size), with reference elements at REF_ELEMENT_SIZE = 8 B.
Compressed oops (gc/src/compressed_oops.rs) are wired into the live
heap now — enable_for_live_heap is called from vm/src/vm/vm_init.rs:874 —
but behind an opt-in -XX:+UseCompressedOops / CRATONVM_COMPRESSED_OOPS=1
gate, and use_compressed_oops defaults to false (vm/src/config.rs:426), so
the default path does not use it. When the gate is on, reference instance fields
and reference array elements narrow from 8 bytes to 4; the klass pointer is
deliberately not narrowed (ObjectHeader::class_id is already a u32). The
reason the gate stays off is a throughput regression, not incompleteness: the
JIT's inline compact-field fast paths are disabled while compressed oops are
active, so getfield/putfield fall back to the always-correct helpers.
JIT Compiler (jit/ crate)
Custom x86-64 / AArch64 JIT compiler (~105,000 LoC; x64.rs alone is ~42,600).
Extracted into the cratonvm-jit crate, with shared API types in
cratonvm-jit-api.
lib.rs— JIT infrastructure: compiled code cache, OSR entry points.x64.rs— x86-64 machine code emitter with 26 optimization rounds.aarch64.rs— AArch64 machine code backend (partial coverage).ir.rs,ir_optimize.rs,ir_schedule.rs,ir_lower.rs— optional sea-of-nodes IR pipeline (build, optimize, schedule, lower to x64).
Compilation pipeline: the JIT has two paths. The default is a
single-pass emitter that lowers bytecode directly to x86-64 native code
(x64::compile). Methods that pass ir_compatible() instead go through
an optional sea-of-nodes IR pipeline
(bytecode -> IrBuilder -> Graph -> optimize -> schedule -> lower -> x64,
in ir.rs, ir_optimize.rs, ir_schedule.rs, ir_lower.rs), which
decouples optimization from instruction selection; others fall back to the
direct single-pass path.
The two front ends now share one runtime-sensitive lowering library.
jit/src/runtime_lowering.rs owns the x86-64 contracts for allocation,
megamorphic dispatch, and live monitor calls. The single-pass emitter and IR
lowerer may differ in optimization and scheduling, but both emit those
stateful operations through the same ABI.
ir_compatible() admits up to 64 invokes, 64 instance-field operations, 64
static-field operations, and 16 allocation sites; ir_compatible_sized uses
an 8,000-byte method cap and a 20,000-node graph cap. Static calls lower
directly, virtual/interface calls use the same MIC/four-entry PIC plus compact
eight-set/two-way hashed tail, and escaping Op::New nodes lower through the
class-initializing, TLAB-aware allocation helper. Live monitor bytecodes call
the thin-lock runtime path; only an exact per-PC scalar-replacement proof may
elide a lock.
Coverage gaps such as unsupported invokedynamic, multianewarray, or
exception-frame shapes remain fail-closed admission boundaries: the method
uses another tier or the interpreter rather than receiving different runtime
semantics.
Key optimizations: register allocation for locals, magic division, LICM, bounds check elimination, AVX2 SIMD, on-stack replacement (OSR). Inlining is tiered: trivial callees up to 6 bytes, cold callees up to 35 bytes, and profile-hot callees up to 325 bytes, with total budgets of 750 bytes (cold) or 2,000 bytes (hot). There is still no general cross-call register allocation and no bounded-depth recursion inlining; see BENCHMARK.md for what that costs on recursion-bound rows.
Methods are compiled after CRATONVM_JIT_THRESHOLD invocations (default 500;
see vm/src/runtime/env_cache.rs). Codegen runs off-thread by default — the
tiered manager enqueues a CompilationTask and a background worker publishes
into shared.jit_cache, while the mutator keeps interpreting until the entry
appears (CRATONVM_BG_COMPILE=0 restores inline compilation). Two calling
conventions: pure methods (direct call) and context methods (receive
SharedVm pointer as hidden first argument).
Native Methods (native-api/ and native-* crates)
A large compatibility surface of native implementations and application
bridges, split across domain-specific crates. The real-JDK default registers a
smaller bridge/intrinsic subset; the synthetic-jdk feature adds the synthetic
standard-library surface. native-api/ defines narrow heap, class, invoke,
thread, exception, system, and related capability traits composed by
NativeContext.
native-builtins/— java.lang.*, registration, and Java-object marshalling.native-builtins-crypto/— separately compiled cryptographic compatibility kernels.native-builtins-security/— separately compiled JDK security and SunEC implementation pack.native-collections/— java.util.* native methods.native-io/— java.io/nio native methods.
These stubs are the synthetic standard library — the fallback mode, not the
default one (see Key Design Decision 1). In the default real-JDK mode the same
crates register a much smaller essential-native surface underneath real
java.base bytecode, so a method you find implemented here may not be the one
executing. The NativeContext capability facade provides VM-agnostic,
loader-aware operations in both modes without making native crates depend on
the concrete VM. Common native argument decoding, forwarding, and pin-index
buffers retain up to eight slots inline and spill correctly for larger
descriptors.
Threading (vm/src/threading/)
jvm_thread.rs— Per-thread state (call stack, printed output).thread_registry.rs— Global thread tracking.monitor.rs— Object monitors (synchronized/wait/notify).gc_barrier.rs— Stop-the-world safepoint coordination.virtual_scheduler.rs— Virtual thread scheduler (Java 21+).
GPU Offload (opt-in, --features gpu-offload)
Entirely feature-gated: a default build links none of this and the
interpreter's hot path carries zero extra branches. Gating features:
gpu/gpu-driver on cratonvm-cli, gpu-offload on cratonvm-vm (which
forwards to cratonvm-gc and cratonvm-native-builtins). See
BUILD_GUIDE.md for the build
levels and docs/gpu/README.md for the full reference.
Crates. jit-cuda lowers Java bytecode to PTX: analyzer.rs decides
whether a static method is GPU-eligible (primitives only, no allocation,
calls, fields, or reference arrays), lowering.rs / loop_recog.rs /
emit.rs turn an eligible method's counted loop into PTX text. cuda-bridge
is a thin CUDA Driver API wrapper with a no-driver backend_stub.rs
default and a real backend_cuda.rs (behind cuda-bridge/cuda) built on
cudarc.
Pipeline. The interpreter's execute_invokestatic hook
(vm/src/runtime/interpreter.rs) calls runtime::offload::try_dispatch
(vm/src/runtime/offload.rs), which asks the per-VM OffloadCache to
analyze-and-lower a callee once and cache the resulting CompiledKernel
(a loaded PTX module) by (ClassId, method_index). On a cache hit,
dispatch_method_from_native / dispatch_async marshal the Java
primitive-array arguments — via vm/src/runtime/gpu_marshal.rs's
host_view_<T>/write_back_<T> packed-copy path, or a zero-copy DMA
straight against the JVM heap arena once an array's element type is
proven — then hand them to cuda-bridge's DeviceContext. The context
runs the upload, launch, and download on three CUDA streams (copy_h2d,
compute, copy_d2h) ordered by CUDA events rather than a blocking sync
per stage. Kernels signal failure (e.g. an out-of-bounds index) by writing
a device failure_flag word instead of throwing; finalize_submission
checks it once the event chain completes and deopts to the CPU
interpreter — leaving no partial GPU state in the heap — on a failure,
or copies results back and resumes the Java frame on success.
vm/src/runtime/gpu_residency.rs separately tracks longer-lived
GpuArray<T> host/device residency for the explicit async API, independent
of this transparent per-call path.
GC coordination. While kernel arguments are in flight, the calling
thread holds a SafepointToken from Heap::enter_gpu_critical() and pins
the argument ObjectRefs via Heap::pin_ref (both in gc/src/heap.rs).
The collector checks the resulting gpu_critical_count and yield-spins
rather than moving or reclaiming while any token is alive; the root walker
visits pinned refs so a kernel never reads through a stale or relocated
pointer.
vm-cli — Command-Line Interface
clap-based entry point (~6,000 LoC; main.rs is ~4,900). Parses arguments and
-XX: flags, constructs a Vm, calls main(String[]), and handles exit codes.
Key Design Decisions
-
Two standard libraries, and the default one is the real JDK's. This decision has been reversed since it was first written, and the old wording ("no
JAVA_HOME, nort.jar, no dependency on any JDK installation at runtime") no longer describes the default build. CratonVM carries two complete implementations of the Java standard library:- Real-JDK mode (the deterministic launcher default). Class bytecode is
loaded from an explicitly validated JDK —
$JAVA_HOME/jmods/java.base.jmodor thelib/modulesjimage, via the jimage reader — and Rust supplies only the essential native surface underneath it.VmConfig::for_launcher()always chooses this mode; detection validates availability but no longer selects the library. - Synthetic mode. A large Rust stub library stands in for
java.*, so the VM can run with no JDK on the machine. It must be selected explicitly with--synthetic-jdk, and the launcher rejects that selection when the binary was not built with thesynthetic-jdkCargo feature.
--real-jdkand--synthetic-jdkare symmetric and mutually exclusive.VmConfig::default()remains synthetic for hermetic embedding/tests, while the launcher default is the namedLAUNCHER_DEFAULT_JDK_MODE. Version and fatal-error output identify the selected mode.VmConfig::with_host_jdk_default()is kept only as a compatibility alias forVmConfig::for_launcher()— see the doc comment on either invm/src/config.rs.Practical consequence for anyone diagnosing a failure: establish which library the run used before anything else. A missing real-JDK native and a synthetic-stub gap produce indistinguishable stack traces, and a fix applied to the wrong one changes nothing on the path that actually failed.
- Real-JDK mode (the deterministic launcher default). Class bytecode is
loaded from an explicitly validated JDK —
-
Two-layer exception model.
MethodCallFailed::ExceptionThrownwraps Java-catchable exceptions;MethodCallFailed::InternalErrorwraps VM bugs. The interpreter catches the former at catch/finally blocks. -
Compact frame values. Frames and operand stacks normally encode tag and payload together in an 8-byte
CompactValue. Locals use a parallel one-byte kind only to disambiguate rawlong/doublepatterns and preserve JVM category-2 slot semantics. -
Direct bytecode-to-x86-64, with an optional IR. The JIT's default path compiles bytecode directly to machine code in a single pass, which keeps that path simple at the cost of limiting cross-instruction optimizations. Methods that qualify (
ir_compatible()) are instead routed through an optional sea-of-nodes IR that enables broader optimization before instruction selection. -
Dense, memoized native dispatch. Name resolution hashes the
(class, method, descriptor)triple and verifies the full strings on every digest hit. The result is a stable denseNativeMethodId; aNativeCallSitecaches the registry generation plus slot, so a warm call is an atomic load, generation comparison, and bounds-checked array access. Call sites that still useNativeMethodRegistry::findpay the full string hash and should migrate to the shared memo mechanism. -
Threading: one OS thread per Java thread.
thread_start(vm/src/vm/vm_exec.rs) spawns a realstd::thread::BuilderperThread.start(); virtual threads are multiplexed over carriers byvirtual_scheduler.rs. Beware stale comments elsewhere in the tree that describe Java execution as single-OS-threaded under cooperative scheduling — that has not been true since real thread spawning landed, and anything resting on it (notably theunsafe impl Send/Sync for ObjectRefargument intypes/src/value.rs) should be read with that in mind. -
Configuration is env-var-driven, and the surface is large. There are hundreds of distinct
CRATONVM_*identifiers across the workspace. There are roughly 530 directstd::env::var/var_ossource call sites in the core directories;vm/src/runtime/env_cache.rscontains 40 itself. The typed and grouped flag layers centralise many identifiers, but direct reads remain scattered. Most are debug or diagnostic gates, but a meaningful subset changes semantics (CRATONVM_COMPACT_REF_FIELDS,CRATONVM_REAL_NET_SOCKETS,CRATONVM_REAL_FORKJOINPOOL,CRATONVM_JIT_GETFIELD_HELPER,CRATONVM_BG_COMPILE, …). Most are cached in aOnceLockon first read, but not all — check before adding one to a hot path, and prefer extendingVmConfig/runtime::env_cacheover introducing a new barestd::env::varcall. -
Compatibility policy is a runtime token in
types, not a build feature.--jdk-only(CompatibilityMode::JdkOnly) says real class bytes are authoritative; the defaultCompatibilityMode::Compatibleis today's behaviour, byte for byte. It is orthogonal toJdkMode— that picks which class library boots, this picks which substitutions are permitted — andJdkModewas deliberately not overloaded with strictness. The default isCompatibleon both entry points (LAUNCHER_DEFAULT_COMPATIBILITY_MODEandEMBEDDED_DEFAULT_COMPATIBILITY_MODE,vm/src/config.rs), and you reach it by doing nothing. Contract:docs/feature-designs/jdk-only-mode.md.Four structural shapes were forced here, and each is worth understanding before touching this code, because each has an obvious alternative that does not work:
- The policy token lives in
types, alone. The policy has to be legible to both the native registry (native-api) and the class loader (classloading), and those two are peers in the dependency flow above — neither can see the other. The only crate below both istypes, sotypes/src/compat.rsholdsCompatibilityMode/ExecutionPolicyand nothing else about the feature.ExecutionPolicycarries a barereal_jdk: boolrather than aJdkModefor the same reason:JdkModelives invm, whichtypescannot see. - The predicates live next to their own types, not in one place.
NativeKind::allowed_in(mode)stays innative-apiandClassOrigin::allowed_in(mode)stays inclassloading. A singlefn allowed(kind, origin, mode)would need both enums in one scope, which means either a new edge between two crates that must not depend on each other, or dragging both enums (and their whole surfaces) down intotypes. The shared thing is only the mode; the judgements stay local to the type that can answer them. - One policy-aware dispatch resolver, which every path routes through.
resolve_dispatch/DispatchDecision(vm/src/vm/vm_exec.rs) is the single native-vs-bytecode decision point for the interpreter, JIT, reflection, JNI and method handles. This is not tidiness: the repository already had two independent override gates that disagreed aboutjava/lang/Stringand carried comments asking that they be kept in sync by hand. A policy evaluated at N call sites is N policies, and the divergence is invisible until something silently takes the wrong branch. The resolver exists and the main interpreter path goes through it; paths that still bypass it are marked// JDK-ONLY-WAVE2:so they can be found mechanically rather than by memory. - Policy state is per-VM, never a process global. It lives on
VmConfigand is pushed into the native registry and theClassManagerat VM init, before any registration pass. A process global would break multi-VM-in-one-process runs, and this repository has a documented history of exactly that failure with process-global native caches leaking across VMs. Two pre-existing globals reachable from this feature's paths are logged as violations to remove, not as precedent.
Two consequences that surprise people. First,
Classnow carries aClassOrigin(boot image, application classpath, user-defined, array, hidden, lambda, proxy, reflection accessor, VM-internal, compatibility stub), and the olderis_synthetic_stubbool is a derived mirror of it — write both throughClass::set_origin, never one alone. The bool survives because ~160 read sites across 17 files still use it; collapsing it is a later wave. Second, this capability deliberately violates the "ship opt-out" convention below, and is the one shape that legitimately can: an opt-out flag makes a capability run by default, and a policy whose whole job is to refuse work cannot be default-on without changing what every existing program does. So it is recorded here as what it is — an internal diagnostic, off by default (CompatibilityMode::Compatible,vm/src/config.rs), enforcing at class fabrication and synthetic-native registration only, with the remaining dispatch paths counted rather than blocked. See ROADMAP.md for the staged rollout that is meant to end that exception. - The policy token lives in
How to tell whether a feature actually runs
This repository has a chronic and well-evidenced failure mode: a capability
lands behind a CRATONVM_* environment variable, the variable defaults to
off, the work is recorded as "implemented", and the code never executes on
the default path. The docs then describe a VM that nobody is running. Several
of the corrections recorded in this document were instances of this.
Confirmed instances (all verifiable in the tree today):
| Capability | Where the default lives | Status |
|---|---|---|
moving_young | types/src/flags.rs::DEFAULT_MOVING_YOUNG | Resolved. Restructured as opt-out with a compatibility opt-in, and the compiled default is now true. The former independent allow_moving_young gate has been deleted. Retained here as the worked example of the fix shape, not as an open instance — but note the second gate: a moving cycle still needs its per-cycle coverage proof, so "flag on" and "compaction ran" remain distinct claims. |
use_compressed_oops | vm/src/config.rs | Opt-in, defaults false. Fully wired but never enabled on the default path — and the doc claimed for a while that it was not wired, which is the same failure mode in the opposite direction. |
safepoint_reg_spill | jit/src/x64.rs:2650 | Was the canonical case: several call sites' own comments claimed a register spill as their protection, but that spill "only ran when the SEPARATE CRATONVM_JIT_SAFEPOINT_REG_SPILL env var was ALSO set — off by default, so the documented protection never actually happened." Now folded into default-on precise_maps and inverted to opt-out (CRATONVM_NO_PRECISE_REG_SPILL). This is the shape the fix should take. |
precise_maps register-spill half | jit/src/x64.rs:2646-2662 | Same item; precise_maps was default-on while its register-spill branch was not. A flag being on does not mean all of its branches are. |
Checklist — apply this before recording anything as "implemented":
- Find the default. Look the flag up in
types/src/flags.rs(orvm/src/runtime/env_cache.rs/vm/src/config.rs).present(src, "…")means presence is truth — the flag is off unless the variable exists. Do not infer the default from the call site. - Check that the default path takes the new branch. Read the condition that guards it with the flag at its default value substituted in. If the branch is unreachable that way, the capability does not ship.
- Check for a second gate.
moving_youngformerly neededallow_moving_youngtoo, andprecise_mapsformerly neededsafepoint_reg_spill; both duplications have since been removed. They remain examples of why one flag being on is not evidence that a code path runs. - Check whether the old path was deleted or merely bypassed. If the previous implementation is still present and still reachable, the new one is an alternative, not a replacement — and the old one is what most runs use.
- State the default in the sentence. Not "X is implemented" but "X is
implemented and is on/off by default (
FLAG_NAME,path/to/file.rs:NNN)". A doc sentence that does not name the default is not a claim that can be checked.
Convention for new work: a new capability should ship opt-out, never
opt-in. Land it on by default with a named escape hatch for bisection
(CRATONVM_NO_<FEATURE>), the way precise_reg_spill_disabled was rewritten.
An opt-in flag is acceptable only as a temporary bring-up state, and it must be
recorded as not shipped until it is inverted. Correspondingly, any doc
sentence describing a capability must name its default explicitly — that is
the single cheapest defence against this whole class of drift.
Data Flow
.class bytes
|
v
reader::read_class() parse into ClassFile
|
v
ClassManager::load_class() link, verify, prepare
|
v
interpreter::execute() run bytecode
| ^
| | (hot method threshold)
v |
jit::compile() emit x86-64, cache
|
v
native call compiled code calls back into VM
Testing Strategy
- Unit tests live in
#[cfg(test)]modules alongside production code. - Integration tests in
vm/tests/run compiled.classfiles through the full VM pipeline. - Java test classes in
test_classes/andvm/tests/resources/are compiled bybuild.rsifjavacis available. - CI (
.github/workflows/ci.yml) runscargo fmt --check,cargo build,cargo clippy, andcargo testacross the workspace on Linux and Windows.