Java 25 gives you several performance changes worth testing: compact object headers, ahead-of-time method profiles, and a production-ready generational mode for Shenandoah. Each addresses a different part of application behavior.
There is no universal 30% CPU reduction from upgrading to Java 25. A result from one service or benchmark does not predict another application’s savings. This guide explains the features and provides a measurement plan; it does not present an original production benchmark.
What to Test First
| Feature | Question it helps answer | Main measurements |
|---|---|---|
| Compact object headers | Does object layout materially affect this workload? | Heap, allocation rate, GC activity, CPU |
| AOT profiling and caching | Does the service spend too long warming up? | Startup, first-request latency, time to steady throughput |
| Generational Shenandoah | Are GC pauses preventing the latency target? | Pause distribution, request latency, CPU, memory |
Start by comparing your current runtime with Java 25 using equivalent resource limits and workload. Then change one feature at a time. A combined result is useful only after you understand the individual tradeoffs.
Compact Object Headers
JEP 519 makes compact object headers a product feature in Java 25. They remain opt-in in that release, but no longer require unlocking experimental VM options:
java -XX:+UseCompactObjectHeaders -Xmx4g -jar application.jar
The underlying layout, described in JEP 450, reduces headers to 8 bytes from the traditional 12 or 16 bytes on 64-bit platforms. That is a header-size change, not a promise that every object shrinks by the same percentage. Field layout, reference compression, and alignment affect the final size.
An object count multiplied by a guessed size is not a heap benchmark. Nor is allocation volume per hour the same as live memory: many objects can be created and collected within that hour. Compare measured heap occupancy and allocation behavior, and include total process memory when evaluating container capacity.
AOT Profiling: Faster Warmup Without Rewriting the Application
JEP 515 lets HotSpot reuse method-execution profiles from a training run. Those profiles help the JIT compiler make optimization decisions sooner. This is not a promise that every method is already compiled into native code before the first request, and it does not require a custom annotation on application methods.
The training workload should exercise the paths that matter in production. A service trained only through its health endpoint may miss the routes whose warmup you are trying to improve.
Create and Use an AOT Cache
JEP 514 introduces a one-command workflow for training and cache creation. For a conventional application JAR containing com.example.App:
# Training run: exercise the workload, then allow the application to exit.
java -XX:AOTCacheOutput=application.aot \
-cp application.jar com.example.App
# Subsequent run using the cache.
java -XX:AOTCache=application.aot \
-cp application.jar com.example.App
These commands illustrate the JDK workflow. A Spring Boot executable JAR has different packaging; follow the framework’s documented extraction and cache procedure rather than substituting its archive into a classpath example.
Build the cache for the artifact and runtime configuration you intend to deploy, check the JVM’s cache diagnostics, and repeat the startup measurement across several fresh processes. Record both readiness and time to steady throughput: they measure different things.
Generational Shenandoah
JEP 521 removes the experimental status of generational Shenandoah in Java 25. It does not make that mode the default. On a JDK build that includes Shenandoah, enable it explicitly:
java -XX:+UseShenandoahGC \
-XX:ShenandoahGCMode=generational \
-Xmx4g \
-Xlog:gc*:file=gc-shenandoah.log:time,uptime,level,tags \
-jar application.jar
A collector experiment needs both GC data and application data. Shorter pauses do not guarantee lower CPU use, more throughput, or better request latency when the database is the bottleneck. Compare against the collector you actually operate, under the same heap limit and traffic pattern.
A Repeatable Java 25 Benchmark Plan
Use a test environment with recorded JDK builds, container limits, JVM options, application commit, dataset, and load-generator settings. Keep the load generator separate from the system being measured so its resource consumption does not distort the result.
Run this sequence:
- Current runtime baseline. Record the service’s existing behavior under representative traffic.
- Java 25 baseline. Keep application code and resource limits equivalent. Document any required compatibility changes.
- Compact headers only. Compare layout-related memory and CPU measurements.
- AOT cache only. Measure fresh-process startup and warmup separately from steady-state throughput.
- Collector comparison. Compare GC and request latency as well as throughput and resource use.
- Selected combination. Retest the settings that helped individually, including sustained and burst traffic.
Use repeated runs and report their spread. Do not pick the fastest run or compare a cold baseline with a warmed-up candidate.
Record Metrics That Explain the Result
| Metric | Why it matters |
|---|---|
| Successful requests per second | Establishes the actual work completed |
| CPU time per successful request | Makes CPU comparisons meaningful when throughput differs |
| p50, p95, and p99 latency | Shows typical behavior and slower requests |
| Error and timeout rates | Prevents failed work from looking like a speedup |
| Heap, allocation rate, and GC pauses | Explains memory-management behavior |
| Process memory and CPU throttling | Captures container-level constraints |
| Readiness and warmup time | Separates launch speed from steady performance |
For a CPU comparison, calculate:
CPU per request = process CPU seconds / successful requests
Change (%) = 100 × (candidate CPU per request / baseline CPU per request - 1)
A negative result means less CPU per successful request in that test. Before turning it into a capacity or cost claim, verify the latency target, failure rate, peak demand, and minimum redundancy requirements.
Roll Out the Measured Improvement
Keep the previous JDK image and application artifact available. Start with a canary, compare it with unchanged instances, and monitor under ordinary traffic and peak periods. Revert settings that fail the acceptance criteria even if they improved a synthetic benchmark.
For a wider upgrade plan, use the Java 25 enterprise migration guide. For help designing a benchmark around your application’s constraints, talk through the workload with Katyella.
Related Articles
- Java 25 Enterprise Migration Guide
- Java 8 to 17 Migration Guide
- Java Virtual Threads with Spring Boot
- OpenTelemetry with Spring Boot
Java Modernization Readiness Assessment
15 questions your team should answer before starting a migration. Takes 10 minutes. Could save you months.