Java can already tell you a great deal about a method through reflection: its name, parameter types, annotations and modifiers. It cannot tell you what the method does. The body is compiled to bytecode, and bytecode is a stack-machine encoding tuned for the JVM, with the structure of the original expressions, loops and lambdas flattened away. That gap is why Java has struggled to target GPUs and machine-learning runtimes: to translate Java into a CUDA kernel or an ONNX graph you need the program's structure, and the only standard artefact that carries it is thrown away after compilation.
Project Babylon is the OpenJDK project that closes the gap with code reflection: javac keeps a structured, typed model of selected methods and lambdas, and a program can ask for that model at run time, inspect it, transform it and hand the result to something else. This article explains the model from first principles, walks through the API described in the incubator JEP draft, builds a small transformer, shows how the HAT toolkit in the Babylon repository uses the same machinery to produce GPU kernels, and ends with the risks of adopting an incubating API. At the time of writing the JEP is at Submitted status with no release target, so names can still change; treat every listing as a sketch to check against the current JEP text.
The problem Babylon solves
Consider a library author who wants Java developers to write GPU kernels in Java rather than in a separate language. They need three things from each kernel method: its control flow (loops and branches as structures, not jumps), its data flow (which value feeds which operation) and its types. Classic Java reflection gives none of these, because it stops at the method signature.
The workarounds all hurt. Bytecode parsing works but gives an unstructured stack program, from which loops must be rediscovered and lambda bodies chased into synthetic methods. Annotation processors see source trees at compile time only, and cannot look inside code compiled elsewhere. Embedding another language in strings, the approach many GPU and SQL libraries take, loses type checking and IDE support. Each existing Java-to-GPU project has had to build and maintain its own front end on top of one of these.
Code reflection makes the front end part of the platform. The JEP lists two goals: let developers use familiar constructs such as lambdas and static types to program non-Java models, and let libraries expose such models without asking users to embed foreign code. Its non-goals are just as useful: it does not change Java semantics or the JVM instruction set, it does not standardise javac's internal syntax tree, and it is not a general macro facility.
The architecture
There are four moving parts. The compiler translates each reflectable method or lambda into a code model and stores it in a class file related to the one carrying the bytecode. The reflection API in module jdk.incubator.code returns that model as a tree of Java objects. Transformers rewrite models into new models: lowering structured control flow, converting to pure SSA, replacing operations. Consumers are the libraries that turn a final model into something executable: GPU source, an ML graph, or new Java.
Opting in is explicit. Only code marked with @Reflect gets a model, and both compiling and running need the incubator module: javac --add-modules jdk.incubator.code Main.java and java --add-modules jdk.incubator.code Main. The bytecode is generated as usual, so the method still runs on the JVM exactly as before; the model is additional metadata.
The code model, from first principles
A code model is a tree of code elements with three kinds of node. An Op is an operation: an addition, a variable declaration, a method call, a return, or a whole function. An op can contain one or more Body objects, a body contains Block objects, and a block contains a sequence of ops. The root of a method's model is a CoreOp.FuncOp.
Data flow uses Value: every op's result is a value, and so is every block parameter. Operands refer to values, never to names, and each value is assigned exactly once. That is static single-assignment (SSA) form, the representation compilers use because it makes data flow explicit. Each value has a CodeType. Here is the documented textual form for a method that subtracts two doubles:
func @"sub" (%0 : java.type:"double", %1 : java.type:"double")java.type:"double" -> {
%2 : Var<java.type:"double"> = var %0 @"a";
%3 : Var<java.type:"double"> = var %1 @"b";
%4 : java.type:"double" = var.load %2;
%5 : java.type:"double" = var.load %3;
%6 : java.type:"double" = sub %4 %5;
return %6;
};Read it top to bottom. Parameters become block parameters %0 and %1. Java local variables are still visible as var ops holding boxes, so the model stays close to the source. Loads pull values out, sub consumes two values and produces one, and return ends the function. The real output also carries @loc attributes with source line and column, which is how a consumer can report errors against the user's Java code. The JEP describes toText() output as unspecified and intended for debugging, so never parse it.
Structured control flow is kept as structure. A Java for loop appears as a single op with bodies for its initialiser, condition, update and loop body, which is exactly what a GPU or graph translator wants to see.
Getting a model: methods and lambdas
@Reflect can go on a method declaration, on a variable or field declaration whose initialiser is a lambda, or on a cast applied to a lambda. Methods are fetched through the classic reflection object; lambdas are fetched from the lambda instance, and return a Quoted that pairs the model with the run-time values of captured variables.
import jdk.incubator.code.*; // sub-package names have moved during the prototype; check the JEP
public class Gray {
@Reflect
static int gray(int r, int g, int b) {
return (29 * r + 60 * g + 11 * b) / 100;
}
public static void main(String[] args) throws Exception {
var m = Gray.class.getDeclaredMethod("gray", int.class, int.class, int.class);
var model = Op.ofMethod(m).orElseThrow(); // a FuncOp; empty if not @Reflect
System.out.println(model.toText());
byte[] rgb = new byte[300], out = new byte[100];
java.util.function.IntConsumer toGray = (@Reflect java.util.function.IntConsumer) i ->
out[i] = (byte) gray(rgb[i * 3], rgb[i * 3 + 1], rgb[i * 3 + 2]);
var lambdaModel = Op.ofLambda(toGray).orElseThrow().op(); // captured rgb, out held in Quoted
lambdaModel.elements().forEach(e -> System.out.println(e.getClass().getSimpleName()));
}
}Two details matter in practice. First, Op.ofMethod returns an Optional, empty when the method was not compiled as reflectable, so a library should fail with a clear message naming the method rather than a bare NoSuchElementException. Second, elements() streams every code element in pre-order, which is the easiest way to write analysis passes such as "reject any method call that is not on the allow-list".
Transforming models
Models are immutable. A transformation walks an input model and builds a new one through a CodeTransformer, a functional interface invoked for each op with a builder, the input op and the already-transformed operands. Returning the op unchanged copies it; emitting something else rewrites it. The JEP's own example turns every integer add into a call to Integer.sum:
MethodRef SUM = MethodRef.method(Integer.class, "sum", int.class, int.class, int.class);
CodeTransformer addToCall = CodeTransformer.opTransformer(
(Function<Op, Op.Result> builder, Op inputOp, List<Value> operands) -> {
switch (inputOp) {
case AddOp _ -> builder.apply(invoke(SUM, operands)); // rewrite
default -> builder.apply(inputOp); // copy
}
});
FuncOp rewritten = model.transform(addToCall);Two built-in passes do most of the heavy lifting for consumers. CodeTransformer.LOWERING_TRANSFORMER replaces structured ops such as java.for with plain blocks connected by branch ops that pass values as block arguments, the shape a code generator for a flat instruction set wants. SSA.transform(model) then removes the var boxes, replacing each load with the value most recently stored, which leaves pure data flow. A typical consumer pipeline is: fetch, validate, apply domain rewrites while structure is still visible, lower, convert to SSA, then emit.
Models can also be built directly with builder methods such as func, constant, add and return_, which is how a transformation such as automatic differentiation emits a new function, for example the gradient of an input method. The Babylon project publishes an article on doing exactly that.
Worked example: from a Java lambda to a GPU kernel with HAT
HAT, the Heterogeneous Accelerator Toolkit, lives in the Babylon repository and is the project's main demonstration; it is not part of the JEP. A HAT kernel is an ordinary reflectable static method. It receives a KernelContext whose gix field is the global thread index, and data in buffer types such as F32Array that are backed by off-heap memory through the Foreign Function and Memory API, so the same bytes can be handed to a GPU driver without copying Java arrays.
@Reflect
public static void vectorMul(@RO KernelContext kc, @RO F32Array a, @RO F32Array b, @RW F32Array c) {
if (kc.gix < a.length()) { // guard: the grid may be larger than the data
float x = a.array(kc.gix);
float y = b.array(kc.gix);
c.array(kc.gix, x * y);
}
}
@Reflect
public static void compute(@RO ComputeContext cc, @RO F32Array a, @RO F32Array b, @RW F32Array c) {
NDRange range = NDRange.of(Global1D.of(a.length()));
cc.dispatchKernel(range, kc -> vectorMul(kc, a, b, c));
}This listing follows the shape of the HAT matrix-multiplication article (which reads the values into int locals; float is used here). The data flow at run time is:
- The host calls
computethrough a HAT accelerator object bound to a backend such as OpenCL or CUDA. - HAT fetches the code model of
computeand every reflectable method it reaches, so it discoversvectorMulfrom the lambda passed todispatchKernel. - Transformers rewrite the model into a GPU-friendly one: calls on
KernelContextbecome thread-index built-ins, buffer accessors become pointer arithmetic, and Java-only constructs are rejected. - A code generator walks the final model and emits OpenCL C or CUDA source, which the vendor driver compiles.
- The
@ROand@RWannotations tell the runtime which buffers must be copied to the device before the launch and back afterwards.
Compare this with the Vector API, which gives explicit SIMD on the CPU through ordinary library calls that the JIT recognises. The Vector API needs no code model because the operations themselves are the API. Babylon is for the other case, where the whole method must be re-expressed for a different machine.
Trade-offs and failure modes
| Approach | What the translator sees | Cost |
|---|---|---|
| Code reflection | typed, structured SSA model of opted-in code | incubating API; models add class-file size |
| Bytecode analysis | stack code, lambdas in synthetic methods | must rediscover loops; fragile across javac versions |
| Annotation processing | source trees at compile time | no access to code compiled elsewhere; no run-time data |
| Embedded DSL strings | text only | no type checking, no IDE support, injection risk |
- API churn. Incubator modules can change or disappear between releases, and package names have already moved during the prototype. Isolate every
jdk.incubator.codeimport behind one adapter class in your code. - Missing models. A method compiled without
@Reflect, or a call into a third-party method that is not reflectable, leaves a hole that a GPU translator cannot fill. Validate the whole call graph up front and name the offending method in the error. - Semantic mismatch. Java has exceptions, garbage-collected objects, virtual dispatch and defined integer overflow; GPUs have none of the first three. Consumers must reject what they cannot translate rather than silently change meaning. Test every kernel against its plain-Java execution on the same inputs.
- Forgotten module flags. Without
--add-modules jdk.incubator.codeat compile time the annotation does not resolve; at run time the API classes are not found. Put the flag in the build and the launcher, not in a README. - Parsing text output.
toText()is for humans. Tools that depend on its format will break; walk the object model instead.
What to do next
- Read the Code Reflection JEP draft (JDK-8361105) and note its current status and package names before writing any code.
- Build the Babylon repository's code-reflection branch, compile one
@Reflectmethod with--add-modules jdk.incubator.code, and print its model withtoText(). - Apply
LOWERING_TRANSFORMERandSSA.transformto a method containing a loop and compare the three models side by side. - Write one small
CodeTransformerof your own, for example a checker that rejects allocation inside a hot method, and test it on both passing and failing inputs. - If you have a GPU, run the HAT examples with the OpenCL or CUDA backend and check each kernel's output against plain Java.
- Keep all incubator imports behind one interface so that a future API change touches a single file, and revisit the JEP whenever you upgrade the JDK.