Skip to content
Back to skills

Jvm Ml Inference

ASecurity

Engineering CPU and accelerator-backed ML inference from JVM applications: choosing in-process versus remote serving, bounding native sessions and predictors, coordinating engine and request parallelism, batching under a latency deadline, reusing direct buffers, warming deployments and diagnosing native memory outside NMT. Use when DJL, ONNX Runtime or another native inference engine loses throughput as concurrency rises, leaks RSS, overloads a model pool or needs graceful degradation. Traini...

  • 2 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 19, 2026
developmentgojavatestingapiperformancedocumentation

Works with

  • api

Security analysis

A100/100

Pro scans all 3 files and shows the line behind each finding

Scanned September 29, 2026

npx -y skills add robsonkades/agent-skills --skill jvm-ml-inference --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Jvm Ml Inference?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Jvm Ml Inference
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/robsonkades-jvm-ml-inference/badge)](https://www.skillsdirectory.com/skills/robsonkades-jvm-ml-inference)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: jvm-ml-inference
description: >
  Engineering CPU and accelerator-backed ML inference from JVM applications: choosing in-process
  versus remote serving, bounding native sessions and predictors, coordinating engine and request
  parallelism, batching under a latency deadline, reusing direct buffers, warming deployments and
  diagnosing native memory outside NMT. Use when DJL, ONNX Runtime or another native inference
  engine loses throughput as concurrency rises, leaks RSS, overloads a model pool or needs graceful
  degradation. Training and choosing model quality targets are outside scope.
---

# JVM ML Inference

## Purpose

Treat inference as a bounded native resource and queueing system. More request threads do not create
CPU, accelerator streams, native sessions or memory bandwidth; they can oversubscribe the engine and
worsen both throughput and tail latency.

## Workload contract

Use the request and existing artifacts to identify the affected inference path and decision. For
capacity or serving design, record model/version, engine/provider/version, target devices,
input shapes and batch distribution, pre/post-processing, in-process or remote boundary,
session/predictor ownership, engine thread
settings, admitted concurrency/queue, warm-up state, latency SLO, useful throughput, heap/RSS/native
memory and failure/fallback semantics. For a focused correction, collect the subset needed to
establish its contract; missing unrelated measurements need not block it.

Inspect build toolchains, compiler release, runtime image, resolved Java/native artifacts and
provider/device support. This skill prescribes no universal JDK baseline; engine requirements
and virtual-thread guidance are version-specific (virtual threads are final in JDK 21).
Do not upgrade Java or the engine merely to apply an example from current documentation.

For changes to tensor conversion, batching, precision or provider, establish the accepted input
and output contract before treating a performance gain as useful. Check representative outputs
with project-defined tolerances; model-quality targets belong to the model owner, but preserving
them is a serving obligation. Keep findings-only reviews and diagnosis read-only unless fixes
are authorized.

## Workflow

1. When choosing or reconsidering the serving boundary, compare in-process versus remote serving
   from latency, isolation, scaling, model cadence, accelerator sharing, failure domain and
   operational ownership. Preserve an adequate existing boundary.
2. For native lifecycle changes or memory diagnosis, inventory the affected resources and lifetimes.
   Bound sessions, predictors, arenas, tensors and direct buffers; close resources that expose
   ownership/close contracts and bound retained storage where reclamation is GC-managed.
   Reuse only after actual native completion.
3. For concurrency tuning, use the relevant axes of outer requests, session count,
   intra-op/inter-op threads and device streams. Start from current settings and a measured
   bottleneck; verify actual provider/operator placement and transfer costs before attributing
   low accelerator utilization to too few workers. Measure absolute goodput and tail latency,
   not speedup alone.
4. If batching is used, bound size/storage and maximum wait, and dispatch before the earliest
   member deadline minus the execution/remaining-work budget. Test sparse and burst traffic;
   full-batch throughput is not a latency policy. Equal shapes do not establish independent
   requests: preserve sequence state, ordering and affinity when the model requires them.
5. For deployment/readiness changes, warm the JVM code path and the model/engine separately,
   then gate readiness on representative inference with expected outputs rather than model-file
   load. During replacement, budget old/new overlap and drain actual native calls and consumers
   before releasing the retired generation; a caller timeout is not completion.
6. When changing admission or failure behavior, test overload in an isolated or authorized load
   environment. Use bounded admission, deadline-aware rejection/cancellation and an explicit
   fallback, if required by the contract, whose quality is also verified.

## Decision rules

- Pool only resources documented as non-thread-safe or expensive to create. Pool size must match a
  measured useful concurrency limit, not request concurrency. A documented shareable session
  with bounded admission may already suffice; non-thread-safe resources can also be confined
  or serialized without a pool.
- Reuse direct, native-order buffers when the API permits, with exclusive ownership across
  filling, native execution and result consumption. Moving allocation from heap to direct
  memory inside the hot path does not remove allocation or guarantee zero device copies.
- Native CPU work can retain a virtual-thread carrier and does not gain throughput from virtual
  threads. Isolate/admit it with a bounded executor when necessary.
- NMT excludes many third-party native allocations. Compare process/cgroup RSS with NMT categories
  and application counters for live sessions/tensors; use native profilers where required.
  RSS growth alone does not identify a leaking owner, and RSS minus NMT is not a leak-size metric.
- `jdk.VirtualThreadPinned` absence cannot clear CPU-bound time inside native code; the event needs a
  relevant park/block path to become visible.
- A cancelled Java future may not stop native computation. Define abandonment, late completion and
  resource reclamation explicitly.

## Evidence output

For a focused review, return the finding, supporting contract/evidence, consequence and smallest
correction or check. When measurements are missing, keep causal diagnoses and sizing conditional
and name the evidence that would distinguish the hypotheses. For experiments, separate measured
result from analytical ceiling, pin environment and raw output, and compare the same metrics
after a change, including output-regression checks. Report what was verified and what remains
unknown; an adequate existing design is a valid outcome. If output acceptance is unresolved,
pass the model owner the changed configuration and representative before/after outputs; keep
the accepted serving path while that consequential decision remains open.

## References

- [Native resources, parallelism and batching](references/native-resources-and-batching.md) — read
  when reviewing native ownership, sizing pools, engine threads, buffers, batching or model
  replacement; also read before changing tensor conversion, precision or provider. Includes
  output contracts, versioned library sources and diagnostic coverage limits.
- Use `jni-and-ffm`, `off-heap-memory`, `concurrency-limiting-and-bulkheads` and
  `load-testing-advanced` for their owning mechanisms.

Files in this skill

  • SKILL.md5.5 KB
  • references/native-resources-and-batching.md5.8 KB
  • skill.yaml1.5 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…