CratonVM
Phase 3 Java async API — `GpuExecutor`, `GpuFuture`, `GpuArray`, `GpuStream`
Phase 3 Java async API — GpuExecutor, GpuFuture, GpuArray, GpuStream
Reference for the explicit, async Java surface introduced in Phase 3.
Companion to the automatic invokestatic offload path already documented
in README.md.
The two paths coexist:
- Automatic offload (Phase 1/2, the "transparent" path elsewhere in
these docs) — a
gpu/gpu-driverbuild plus--gpuon the command line; a static method the analyzer accepts, called viainvokestatic, is routed to the GPU without any call-site change. The user writes no async code. Methods outside the accepted shape run on the CPU as usual. - Explicit async API (Phase 3, this document) —
GpuExecutorlets the application schedule offloaded work, chain kernels on the same device stream, and read results back asGpuFuture<T>.
Same lowering pipeline, same analyzer, same PTX cache. Different entry point.
--gpuis required for both. This document used to say the explicit path needs no flag because "the executor probes the driver itself". The executor does acquire its ownDeviceContext— and that is not the context the dispatch uses.OffloadCache::newsetsctx = Noneunlessconfig.gpu_offload_enabled, so without--gpueverylookup_or_compilereturnsSkipand every submission is recorded as failed. What made the wrong claim survive is thatGpuExecutor.open()still succeeds: the failure appears one call later, atget(), as "method not offloadable (Skip)".
Quick start
Before — automatic offload of a single eligible kernel via invokestatic:
// Run on a JVM started with --gpu. The call below is dispatched to the
// GPU by the interpreter hook with no application-visible asynchrony.
int[] out = new int[n];
Pipeline.vectorAdd(a, b, out);
useResult(out);
After — explicit submission through GpuExecutor:
try (GpuExecutor exec = GpuExecutor.open()) {
GpuFuture<int[]> f = exec.submit(() -> Pipeline.vectorAdd(a, b));
useResult(f.get());
}
This snippet shows the target shape of the API — a kernel that allocates and returns
int[]directly. As implemented today, the dispatch layer only ever surfacesSerializedResult::Void: a kernel that itself returns an array (rather than writing into a caller-supplied output array) is not something the current marshalling path can hand back throughGpuFuture<T>. Every example in this document that shows a return-value kernel signature should be read as the intended API shape; see Current limitations for exactly what dispatches today.
The explicit form is for code that wants to overlap host work with kernel execution, chain kernels on a single stream (one H2D, multiple launches, one D2H), or fail soft on machines that have no GPU.
GpuExecutor
GpuExecutor owns a DeviceContext, a default GpuStream, and a
reference to the per-VM OffloadCache. One executor binds to one CUDA
device; cross-device traffic requires a second executor.
Opening
| Constructor | What it does |
|---|---|
GpuExecutor.open() | Probe device 0. Throws GpuException if the driver is missing only on gpu-driver builds; on gpu (stub) builds it returns a stub executor whose every submission yields an immediately-failed future. See Stub-mode behavior. |
GpuExecutor.open(int deviceOrdinal) | Attach to a specific CUDA device. Same fall-back behavior. |
The executor is AutoCloseable. Always wrap it in try-with-resources.
Closing the executor:
- Synchronises the default stream.
- Cancels any not-yet-launched submissions.
- Releases the
DeviceContextand the underlying CUDA primary context handle.
Submitting work
| Method | Purpose |
|---|---|
<R> GpuFuture<R> submit(Supplier<R> kernel) | Schedule a kernel for execution. The Supplier body must be a static method reference (e.g. () -> Pipeline.vectorAdd(a, b)); see Limitations. |
<R> GpuFuture<R> launch(KernelHandle<R> handle, Object... args) | Lower-level form. Skips lambda inspection; handle is a CompiledKernel returned by a previous prepare(...) call. |
GpuStream newStream() | Create an additional CUDA stream bound to the executor's context. The default stream is used implicitly if you never call this. |
submit returns immediately. The analyzer runs synchronously on the
calling thread before the future is returned, so any
AnalyzerRejection is observable at submit time, not at get() time.
Ownership
try (GpuExecutor exec = GpuExecutor.open()) {
GpuFuture<int[]> f = exec.submit(() -> Pipeline.vectorAdd(a, b));
int[] result = f.get();
}
GpuExecutor is AutoCloseable and not thread-safe for close();
submissions are. Treat it like a java.util.concurrent.ExecutorService:
one owner thread closes it, many threads may submit while it is open.
GpuFuture<T>
Returned by every submit / launch. Modelled on
CompletableFuture but bound to a CUDA stream.
Completion model.
dispatch_async(vm/src/runtime/offload.rs) launches the kernel, records a CUDA event, and returns immediately with the submission in theRunningstate.GpuFuture.get()still works exactly as before: it calls the blockingNative.futureSynchronize→finalize_submission, which callsevent.synchronize()(a blocking host wait) and drains the writebacks. What's new is thatisDone()/futureStatusare no longer a dead read of stale state: they now callpoll_submission_status(vm/src/runtime/offload.rs), a genuinely non-blocking check that first looks at adevice_doneflag set by a best-effortcuLaunchHostFunchost callback registered at dispatch time, and falls back to a non-blockingEvent::query()(cuEventQuery) if the callback hasn't fired yet. If either signals the device is actually done,poll_submission_statusruns the same finalize workget()would have — inline, on the polling thread, with no further device wait (the event has already fired) — soisDone()returningtruenow means the submission really is finalized, not just "probably."Completion no longer requires a Java thread to call anything. The same
cuLaunchHostFunchost callback that setsdevice_donealso wakes a process-wide completion reaper thread (ensure_completion_reaper_started/completion_reaper_loopinvm/src/runtime/offload.rs) that runsfinalize_submissionitself, off the mutator entirely — draining writebacks, releasing the GC-critical guard, and flipping the submission toCompleted/Failedwhile the application does something else.isDone()/get()now frequently just read an already-terminal status. The reaper is a single background daemon thread for the whole process (modelled on the existing background-JIT-compiler worker), started lazily on first dispatch and captures aWeak<SharedVm>so it can never keep a torn-down VM alive.
| Method | Semantics |
|---|---|
boolean isDone() | A real non-blocking device probe (Native.futureStatus → poll_submission_status): checks a host-callback flag, falls back to non-blocking Event::query(), and finalizes inline if the device reports done. Safe to poll in a loop — each call does bounded work, never a blocking device wait. The submission is often already finalized by the background completion reaper by the time this is called at all. |
T get() | Block the calling thread until done. This is the call that actually finalizes the submission (waits on the CUDA event, drains writebacks) if isDone()/getNow() haven't already done so. Throws GpuException if the kernel failed or the analyzer rejected the lambda. |
T get(long timeout, TimeUnit unit) | Bounded wait. Throws TimeoutException on expiry. |
Optional<T> getNow() | Non-blocking peek, same underlying poll_submission_status probe as isDone(). Returns the result if the device reports the submission complete (finalizing it inline as a side effect, same as isDone()), Optional.empty() while still Running. |
<R> GpuFuture<R> thenApplyGpu(Function<T, R> next) | Spec'd stream-resident chaining — see Current limitations; today's dispatch layer has no mechanism to hand a kernel's device-side output directly to a second launch without a host round-trip. |
<R> CompletableFuture<R> thenApplyAsync(Function<T, R> next, Executor cpu) | Standard CPU continuation. Inserts a D2H copy. |
CompletableFuture<T> toCompletableFuture() | Bridge into JDK async land. Inserts a D2H copy on the first read. |
Stream-resident chaining
thenApplyGpu is specified to be the one continuation that stays on
the device: the function must itself be a @GpuKernel-eligible static
method reference, admitted the same way submit's lambda is. That
lambda-resolution half is real (gpu_resolve_lambda_target /
gpu_dispatch_method in vm/src/vm/vm_exec.rs) and is what
submit/launch/submitWithArg(s) already use.
What is not real yet: the actual "stays on the device" part. A
kernel's result is only ever surfaced to the host as
SerializedResult::Void — the write into the caller-supplied out
array is the result; there is no code path that keeps a kernel's
output as an opaque device-resident handle and feeds it as the next
kernel's input without a host round-trip (see Current
limitations). Until that lands, treat
thenApplyGpu chains as target-API documentation rather than a
working no-D2H fast path, and expect each stage to behave like a
fresh submit against host-visible arrays.
Mixing CPU continuations (thenApplyAsync) and GPU continuations
(thenApplyGpu) on the same future is fine. Each CPU continuation
forces a D2H copy at the boundary.
GpuArray<T>
A GpuArray<T> is a typed handle to a primitive Java array whose
device residency is managed by the executor. Wrapping a host array
does not copy it immediately; the upload happens on the first
kernel launch that consumes the array.
| Method | Purpose |
|---|---|
static GpuArray<int[]> wrap(int[] host) | Create a handle. Native.arrayWrap* takes an eager byte-copy snapshot of the array into a Rust-owned buffer (native-builtins/src/craton_gpu.rs::wrap_primitive_array) — there is no Heap::pin_ref or other GC-pinning call; the snapshot exists precisely so GC moving the original Java array afterward is a non-issue. Type parameter T is one of the supported primitive-array types. |
static GpuArray<int[]> allocate(GpuExecutor exec, int len) | Allocate a device-only array. toHost() materialises a fresh Java array on first call. Rust-side shim landed, Java jar binding still pending. native-builtins/src/craton_gpu.rs now registers arrayAllocateInt/arrayAllocateLong/arrayAllocateFloat/arrayAllocateDouble (builtin_array_allocate_*, minting a zero-filled device-only state::ArrayEntry the same way arrayWrap* does for a host-backed one) alongside the existing arrayWrapInt/Long/Float/Double. What's still missing is the craton-gpu-java side: the external craton/gpu/internal/Native class and the public GpuArray.allocate(...) factory method that would call it. Until that binding lands, the native entry points exist and are ready to call but nothing in the Java jar calls them yet. |
CompletableFuture<T> toHost() | Schedule a D2H copy and return a future for the host array. Always synchronises the stream. Treat it as the expensive read-back operation it is. |
int length() | Element count. Free; does not touch the device. |
void close() | Release the device buffer. Idempotent. |
Residency across kernels
The point of GpuArray<T> is that the array can sit on the device
across many kernels with no intermediate H2D / D2H traffic. The first
kernel that writes an array marks it dirty; subsequent kernels that
read the array see the dirty version without a host round-trip. A
toHost() call is the only operation that forces a D2H copy.
GpuArray<int[]> img = GpuArray.wrap(image);
GpuFuture<int[]> step1 = exec.submit(() -> Pipeline.convolve(img.toHost().get()));
GpuFuture<int[]> step2 = step1.thenApplyGpu(c -> Pipeline.threshold(c, t));
return step2.get();
The step1 lambda above takes the wrapped image and runs a convolve
kernel; step2 chains a threshold kernel on the same stream. With
thenApplyGpu the intermediate int[] never round-trips to the
host. The final step2.get() is the single D2H copy.
Type erasure caveat
GpuArray<int[]> and GpuArray<float[]> are distinct only at compile
time. The runtime element type is recovered from the Java class of the
wrapped array (image.getClass().getComponentType()). If you build a
GpuArray<Object> via raw types and feed it to a kernel expecting
int[], the analyzer rejects the launch at submit time with
Reason::TypeMismatch. There is no kernel-side casting.
GpuStream
GpuStream exposes one CUDA stream. It is rarely needed; most users
should rely on the executor's implicit default stream.
| Method | Purpose |
|---|---|
<R> GpuFuture<R> submit(Supplier<R> kernel) | Submit onto this specific stream instead of the executor's default. |
void synchronize() | Block until the stream is drained. |
void close() | Destroy the stream. Outstanding futures complete or fail before close returns. |
Executor default-stream affinity is real; explicit
GpuStreamrouting is still not. These used to be one limitation; they are now two different states.submit/launch/submitWithArg(s)/submitMethodall route throughresolve_or_create_default_stream(native-builtins/src/craton_gpu.rs), which lazily creates one real CUDA stream perGpuExecutorhandle and caches it (executor_default_stream: HashMap<u64, u64>) — every submission on the sameGpuExecutor, with no explicit stream involved, now serializes on that one real device stream, instead of each dispatch minting and tearing down its own one-shot stream.Native.newStreamalso mints a genuine CUDA stream now (ctx.gpu_stream_create(), not bookkeeping), but there is still no registeredNative.*entry point that lets Java code aim a dispatch at that explicit handle instead of the executor's default —GpuStreamin the implemented API surface ishandle()+close()only, nosubmit. So: one executor implicitly shares one real stream across its submissions today;newStream()mints a real stream you cannot yet route work onto.
Reasons the API intends to let you take a stream explicitly (once the routing above lands):
- Overlap. Two streams = concurrent H2D, kernel, and D2H across
stages of a pipeline. (
cuda-bridgedoes not yet expose a page-locked/pinned host-memory allocator — seestreams-events.md— so the overlap benefit here comes from stream concurrency, not pinned transfers.) - Isolation. Errors in one stream do not affect work queued on another.
- Deterministic ordering. Within a single stream, kernels execute in submission order. Across streams there is no ordering guarantee beyond what the application enforces.
If you don't have a specific reason to call newStream(), don't — the
handle it returns is real, but nothing in the registered Native.* surface
lets you route a dispatch onto it, so it still has no observable effect on
where work actually runs today. If your goal is "one executor's submissions
serialize predictably," you already have that from the default stream with no
newStream() call at all.
Three-stage pipeline example
A realistic image-processing pipeline: convolve → threshold → histogram.
Two host arrays (image, kernel) are wrapped, three kernels are
submitted, the final histogram comes back to the host. Note the device
residency: one H2D for image, one D2H for histogram.
import static cratonvm.gpu.GpuExecutor.open;
int[] runPipeline(int[] image, int[] kernel, int threshold) {
try (GpuExecutor exec = open()) {
GpuArray<int[]> img = GpuArray.wrap(image);
GpuArray<int[]> ker = GpuArray.wrap(kernel);
// Stage 1: convolve. Stays on device.
GpuFuture<int[]> convolved = exec.submit(
() -> Pipeline.convolve(img.toHost().get(), ker.toHost().get()));
// Stage 2: threshold. Same stream as stage 1; no D2H.
GpuFuture<int[]> thresholded = convolved.thenApplyGpu(
c -> Pipeline.threshold(c, threshold));
// Stage 3: histogram. Same stream; no D2H.
GpuFuture<int[]> hist = thresholded.thenApplyGpu(
Pipeline::histogram);
// The first and only D2H copy:
return hist.get();
}
}
The kernel methods themselves are ordinary static methods marked with
@GpuKernel; the executor consults the analyzer to confirm that each
one is offload-eligible before issuing the launch.
If any stage's lambda fails analyzer admission (e.g. Pipeline.histogram
contained an allocation), the thenApplyGpu call that referenced it
returns an already-failed future and the downstream stages never reach
the device. The earlier stages, already in flight, still drain on the
stream and their results are dropped on executor close.
Stub-mode behavior
On a cargo build --features gpu (stub) binary, or on a gpu-driver
build run on a machine with no CUDA driver, GpuExecutor.open() does
not throw. It returns a stub executor with three properties:
- Every
submit(...)andlaunch(...)returns an already-failedGpuFuture<T>whoseget()throwsGpuException("no CUDA device available"). thenApplyGpu(...)on a failed future propagates the same failure without invoking the function.close()is a no-op.
This lets unit tests exercise the control flow — try / catch, future
chaining, fallback paths — on hardware without a GPU. The stub futures
are immediately in the failed state, so isDone() returns true and
getNow() returns Optional.empty() instantly.
Code that wants to fall back to a CPU path on stub builds:
int[] result;
try (GpuExecutor exec = GpuExecutor.open()) {
result = exec.submit(() -> Pipeline.vectorAdd(a, b)).get();
} catch (GpuException e) {
result = Pipeline.vectorAddCpu(a, b);
}
The try-with-resources close on a stub executor does not raise.
Error handling
GpuException is unchecked. It is thrown by:
| Site | Cause |
|---|---|
GpuExecutor.open(...) (gpu-driver only) | Driver init failed for a reason other than "no driver" — e.g. CUDA returns ERROR_OUT_OF_MEMORY while creating the primary context. The stub path swallows NoDriver; everything else propagates. |
submit(...) | The lambda's referenced method failed analyzer admission. Wraps the analyzer's Reason (allocation, non-static, monitor, …). |
GpuFuture.get() | The kernel ran but wrote 1 to failure_flag (array bounds, future arithmetic). Or the kernel itself failed to launch (resource exhaustion). |
GpuFuture.thenApplyGpu(...) | The continuation's referenced method failed analyzer admission. Returned as a failed future, not thrown. get() on that future raises. |
GpuArray.toHost() | D2H copy failed mid-stream. Rare; usually indicates the kernel that wrote this array faulted. |
GpuException carries:
cause()— the underlying RustDeviceErrorrendered as a Java exception chain.DeviceError::NoDriversurfaces as the only exception that the stub executor produces.getKernel()— the simple class/method name of the offending kernel, ornullfor analyzer-time rejections that fire before a kernel was ever bound.getReason()— the analyzer'sReasonenum value when applicable;nullfor device-side failures.
GpuException is intentionally not a CompletionException. It does
not get wrapped twice when transiting toCompletableFuture().
Cleanup
try-with-resources is the only supported lifecycle. Specifically:
try (GpuExecutor exec = GpuExecutor.open();
GpuArray<int[]> img = GpuArray.wrap(image)) {
// ... submit kernels ...
}
The order matters: GpuArray closes first, releasing its device-side
buffer and host-bytes snapshot (see the correction under
GpuArray<T> — there is no GC pin to release, just a
Rust-owned copy and, once uploaded, a cached DeviceBuffer); then the
executor closes, synchronising the stream and releasing the context.
StreamCleaner daemon
If a GpuExecutor is abandoned without close(), the JVM will
eventually collect it. A daemon thread named gpu-stream-cleaner
holds PhantomReferences to all live executors and, on enqueue,
releases the device handles. This is a backup, not the contract:
- The cleaner runs at GC pace, which is unpredictable.
- It cannot synchronise the stream from within itself (no JVM context),
so in-flight kernels may have their output buffers freed before they
complete on the device, producing CUDA
ERROR_ILLEGAL_ADDRESSerrors surfaced on the next kernel launch in any executor. - It logs a
WARNline "executor closed by cleaner; prefer try-with-resources" so misuse is observable.
Always close() explicitly. The cleaner exists so that one forgotten
executor doesn't pin a CUDA context for the JVM's lifetime; it is not
a substitute for resource management.
Current limitations
Implementation-status gaps between this document's target API and what
vm/src/runtime/offload.rs / native-builtins/src/craton_gpu.rs
actually do (post f4311e5f3). These are distinct
from the by-design Limitations below.
- Push-driven completion model (closed).
isDone()/getNow()still work exactly as described underGpuFuture<T>: a non-blockingpoll_submission_statusprobe that finalizes a submission inline the moment it observes the device is done. What used to be missing — anything that drives completion without a Java call — now exists: thecuLaunchHostFunchost callback (Stream::add_host_callback) wakes a background completion reaper thread (ensure_completion_reaper_startedinvm/src/runtime/offload.rs) that finalizes the submission itself, off the mutator, while application code is doing something else entirely. Seeasync-completion-reaper.mdfor the reaper's design. - Only
Voidresults were surfaced at one point; scalar reduction results now reach Java too.SerializedResult'sScalarI32/I64/F32/F64variants (integer/long reduction kernels,)I/)Jdescriptors with ared.global.addepilogue) are wired intoNative.futureGetResult, which now tries the real submission registry first (NativeContext::gpu_future_take_result) and boxes a scalar result via the sameInteger/Long/Float/Doubleboxing path (box_scalar_resultinnative-builtins/src/craton_gpu.rs), falling back to the old pre-Phase-6 synthetic stub-future map only when the real registry has nothing for that handle.PrimitiveArrayI32/I64/F32/F64(a kernel returning a whole array by value, as opposed to writing into a caller-suppliedoutarray) are still never constructed — that part of the target API in this document remains aspirational. Every array-returning kernel signature shown in this document (Pipeline.vectorAdd(a, b),Pipeline.histogram, etc.) still describes the intended surface, not today's behavior; a scalar-returning reduction kernel now genuinely works end to end throughGpuExecutor. GpuStreamaffinity is partially wired up.resolve_or_create_default_streamgives every submission on a givenGpuExecutor— viasubmit/launch/submitWithArg(s)/submitMethod, with no explicit stream involved — one real, shared, lazily-created CUDA stream instead of a fresh private one per dispatch. What's still not wired:newStream()mints a genuine CUDA stream, but no registeredNative.*entry point lets a dispatch be routed onto that explicit handle instead of the executor's default — see the note underGpuStream.- JIT-caller bypass closed. The automatic path still enters through the
interpreter's
invokestatichook, sooffload_jit_gatekeeps a caller with an eligible offload site out of JIT and OSR compilation while--gpuis active. Callers without eligible sites remain compilable. Seejit-caller-gate.md.
Limitations
thenApplyGpurequires a static method reference. Java's type erasure prevents the executor from inspecting an arbitrary lambda's body for@GpuKerneleligibility; the runtime can only identify the target method when the lambda is a directMethodHandleInfowith a resolvableREF_invokeStatickind. A non-method-reference lambda body —c -> { int[] r = new int[c.length]; … }— falls back to a CPU continuation with a D2H copy. The executor logs aDEBUGline identifying which call site lost stream residency.- No graph-capture API.
cudaGraphand friends are not exposed. The three-stage example above is implemented as three discrete launches. When a graph API arrives it will be a new surface onGpuStream, not a retrofit ofthenApplyGpu. - No event timing.
GpuStreamdoes not expose CUDA events to the Java caller. Latency and throughput numbers are measured in Rust viatracingspans (seefirst-results.md) or by the application wrapping itssubmitcalls inSystem.nanoTime. A user-facing event timing API is deliberately deferred — the cost of exposing it cleanly across the stub / driver builds is not yet justified. - Single device per executor. Multi-GPU sharding is the application's job: open one executor per device, partition the work.
- No cancellation mid-launch.
GpuFuture.cancel(true)only succeeds if the kernel has not yet been issued to the stream. OncecuLaunchhas run, the kernel runs to completion.
What this is NOT
- Not a Java CUDA wrapper. You cannot
cuMemAllocfrom Java. The device-side surface is intentionally narrow — wrap an array, launch a kernel, read it back. Anything more is the Rust side's concern. - Not a replacement for the automatic
invokestaticoffload. Code that has been running fine under--gpushould keep running fine; the explicit API is for new code that wants async semantics. - Not a Project Babylon stand-in. Babylon's Code Reflection is a much broader Java-to-anywhere lowering effort. CratonVM's Phase 3 is narrow on purpose: static methods over primitive arrays, async scheduling, nothing more. If Babylon ships in mainline OpenJDK with a PTX target, we will reconsider the surface. Today, we don't depend on it.
See also
streams-events.md— the Rust-sideStream/Eventprimitives that backGpuStreamandGpuFuture. Required reading if you are extending the Java surface.annotations.md—@GpuKernel, the marker the analyzer looks for when admitting a method. Phase 3 lambdas resolve to methods bearing this annotation.README.md— top-level reference: build matrix (gpuvsgpu-driver), CLI surface, file index.reductions.mdandjit-caller-gate.md— scalar-reduction support and the closed JIT-caller bypass.