Skip to content
Back to skills

False Sharing And Contended

ASecurity

Proving and mitigating cache-line false sharing between independent locations with hot writes. Covers ownership and address/layout hypotheses, coherence/HITM evidence limits, JMH topology, arrays and object placement, `@Contended` module/restriction mechanics, grouping and padding, JOL/address validation, manual padding fragility, striping, compact headers, memory cost and cross-socket/NUMA validation. Use when updates to distinct fields scale poorly, when padding is proposed, or when `@Conte...

  • 2 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added September 19, 2026
developmentjavagitapiperformance

Works with

  • api

Security analysis

A100/100

Pro scans all 4 files and shows the line behind each finding

Scanned September 29, 2026

npx -y skills add robsonkades/agent-skills --skill false-sharing-and-contended --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of False Sharing And Contended?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for False Sharing And Contended
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/robsonkades-false-sharing-and-contended/badge)](https://www.skillsdirectory.com/skills/robsonkades-false-sharing-and-contended)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: false-sharing-and-contended
description: >
  Proving and mitigating cache-line false sharing between independent locations with hot writes.
  Covers ownership and address/layout hypotheses, coherence/HITM evidence limits, JMH topology,
  arrays and object placement, `@Contended` module/restriction mechanics, grouping and padding,
  JOL/address validation, manual padding fragility, striping, compact headers, memory cost and
  cross-socket/NUMA validation. Use when updates to distinct fields scale poorly, when padding
  is proposed, or when `@Contended` appears ignored. Cache fundamentals, logical contention
  and general object sizing have separate owners.
---

# False sharing and `@Contended`

## Purpose

Establish that writes invalidate a coherence granule/cache line used by another core for a
different logical variable (writing or reading),
then choose ownership/layout changes whose throughput/latency benefit exceeds memory and maintenance
cost. Cache misses or poor scaling alone do not prove false sharing.

Reuse the requested outcome, existing captures/layouts, ownership and deployment constraints.
Preserve the project's Java target; use Java 25 as the reference when none is specified.
Answer a narrow layout/flag question from its contract without requiring a performance study.
For an investigation, ask only for missing facts that change attribution or the next experiment;
retain an adequate design and work within the accepted capture/recovery budget.

## Ownership boundary

- This skill owns false-sharing hypothesis, layout/placement, `@Contended`, padding and validation.
- `cpu-cache-and-numa` owns cache/coherence/NUMA fundamentals.
- `lock-inflation` and `lock-free-patterns` own logical lock/CAS contention.
- `object-layout-and-footprint` owns general object sizing/header trade-offs.

## Proof contract

```text
independent logical variables/slots, writer ownership and readers:
write/read frequency and production thread/key topology:
actual address offsets/alignment and cache-line size(s):
JDK/layout/header/GC/allocation stability assumptions:
coherence/PMU evidence with support/multiplex/scope:
controlled ownership/layout perturbation:
memory/GC/locality cost and next bottleneck:
```

## Distinguish contention types

| Type                     | Shared meaning                                 | Typical evidence                            | Candidate direction             |
| ------------------------ | ---------------------------------------------- | ------------------------------------------- | ------------------------------- |
| lock contention          | one guarded invariant                          | monitor/park/owner wait                     | reduce/partition guarded work   |
| true data/CAS contention | same logical variable                          | retries/RMW/coherence                       | shard/batch/owner/semantics     |
| false sharing            | different variables on same line               | layout + writers + coherence + perturbation | separate/align ownership/layout |
| capacity/cache locality  | working set misses without writer invalidation | miss/working-set/topology                   | compact/block/localize/prefetch |

Padding true contention does not make the logical hotspot independent.
True contention and false sharing can coexist: threads may contend on one CAS word while its
writes invalidate another core's copy of a distinct neighboring field. Attribute interference
to named access pairs; separating neighbors does not remove contention on the shared word.
At least one participant must write; read-only sharing does not create this invalidation
mechanism. Padding and `@Contended` do not add happens-before, visibility or atomicity:
retain required volatile/atomic/locking semantics or explicit ownership and handoff.

## Evidence ladder

1. Localize the scaling/tail/CPU regression and compare load, locks, CAS hotspots, GC/JIT and I/O
   as competing or coexisting contributors.
2. Map writer/reader access pairs to logical variables and runtime layout/address relationships;
   distinguish same-variable contention from interference between distinct fields/array slots.
3. Collect supported coherence/cache events (for example HITM/snoop variants on some CPUs/tools),
   validating event semantics, multiplexing, skid, process/CPU scope and topology.
4. Apply a controlled separation/ownership perturbation without changing useful semantics/work.
5. Confirm the same production metric improves while memory, GC and locality remain acceptable.

Generic `cache-misses`/LLC misses are not specific and false sharing may manifest as coherence traffic
without the naive counter pattern. A cache-miss flame graph compared with CPU samples is not a
standalone proof.

If PMU access, addresses or placement cannot be verified, report a hypothesis with the
available perturbation evidence and its limits. Do not fabricate zero contention from an
unsupported counter or call a padded speedup alone proof. Return the affected variables,
evidence, proposed change, correctness constraints and measured or pending production check.

## Layout and placement

JOL reports class/instance field layout under its current VM model; it does not by itself prove the
absolute address/alignment of two separately allocated objects over time. Arrays provide predictable
element stride but array base alignment and hardware line size still matter. Derive padding/stride:

```text
stride elements >= ceil(cache-line bytes / element bytes)
```

This is a starting-distance bound, not a proof that the full accesses occupy disjoint lines.
Check each slot's entire accessed byte range against the actual base alignment and line size;
an access straddling a boundary can share a line with the next slot despite distinct starting
lines. “Stride 8 for long” assumes a 64-byte line and suitable alignment; it is not universal.
Account for adjacent-line/prefetch behavior separately where measured.

GC can move objects and allocation adjacency is not a stable API. Prefer layout within one object/
array or ownership partition that can be verified, and test the collector/JDK used.

## `@Contended`

`jdk.internal.vm.annotation.Contended` is internal JDK API. Application source normally needs an
appropriate compile-time module export, and HotSpot commonly restricts user-class padding unless
`-XX:-RestrictContended` is enabled. Runtime exports are needed only when application runtime code
must access the internal type; the VM can recognize annotation metadata without a blanket claim that
every run needs `--add-exports`.

Verify on the exact JDK:

- annotation is present in compiled class and applied in runtime layout;
- effective `EnableContended`, `RestrictContended` and padding settings/support;
- field/class contention group semantics;
- instance versus static fields, and reference slot versus referenced data (HotSpot 25 ignores
  static-field annotations; padding a reference does not pad its object's fields or array elements);
- actual gaps/offsets and object/array placement;
- memory footprint across number of instances and GC consequence.

Padding width is a spacing policy, not “bytes charged per field” exactly; headers, alignment, groups,
field layout and multiple annotations determine total size. Do not justify a default width with one
microarchitecture's prefetch story as a universal guarantee.

Because this is internal API/flag surface, prefer JDK-supplied striped abstractions or an explicit
layout type when feasible, and include upgrade tests.

## Mitigation framework

| Mechanism                          | Prefer when                                               | Cost/risk                                |
| ---------------------------------- | --------------------------------------------------------- | ---------------------------------------- |
| ownership/confinement then combine | exact combination point exists                            | delayed aggregation/semantics            |
| striping                           | commutative/associative approximate or partitioned update | memory, read aggregation, skew           |
| padded field/class                 | stable hot independent fields                             | footprint, internal API/layout drift     |
| array stride/struct-of-arrays      | indexed owners and stable layout                          | wasted space, alignment/index complexity |
| batch updates                      | delayed visibility acceptable                             | burst/tail/failure semantics             |
| compact layout instead             | read/locality dominates, not writer sharing               | can worsen writer density                |

`LongAdder` scales hot cumulative updates using striped cells, but its sum is not an atomic snapshot.
It is not a drop-in replacement for IDs, exact bounds or balances. After writers have stopped and
their completion is observed, `sum()` can provide an accurate total; do not call that result
inherently approximate. Concurrent `reset()`/`sumThenReset()` do not provide an exact interval
boundary. Require a quiescent handoff or a different aggregation protocol when that boundary matters.

Compact object headers can change density/offsets and therefore both footprint/locality and sharing
risk. Feature status/defaults vary across JDKs. Inspect exact JEP/build/layout and re-run evidence;
do not predict false sharing from header size alone.

## Benchmark design

Match JMH state and payload allocation to the ownership being tested, with deterministic mapping
from worker role to field/slot. JMH can pad generated `@State` instances, masking interference
between separate objects; a thread-scoped state with no slowdown does not rule out production
false sharing. Inspect the actual generated state and ordinary payload layouts. Sweep:

- one writer through expected concurrency/overload;
- core, SMT sibling, socket and NUMA placement;
- reads/writes and production work between updates;
- padded/unpadded/owner-local/striped alternatives;
- exact JDK, collector, header and container CPU configuration.

Preserve raw forks and report useful operations, CPU/op, tail, PMU coverage/events, memory footprint,
allocation/GC and placement. A fixed three-fork rule or expected “magical magnitude” is not validity.

## Failure modes

- annotation ignored/restricted or layout differs after JDK upgrade;
- padding separates fields inside an object but adjacent objects/array slots still share;
- manual dummy fields reordered/grouped or optimized around by layout rules;
- stride miscomputed for line/base alignment;
- striping removes false sharing but hot-key/cell collisions create true contention;
- footprint increase causes cache/GC regression larger than coherence benefit;
- benchmark thread-to-field mapping differs from production;
- PMU event unavailable/multiplexed/virtualized and interpreted as zero contention.

## Anti-patterns

| Anti-pattern                            | Failure                                     | Better approach                               | Narrow exception |
| --------------------------------------- | ------------------------------------------- | --------------------------------------------- | ---------------- |
| LLC misses prove false sharing          | many mechanisms cause misses                | layout + ownership + coherence + perturbation |
| `@Contended` costs exactly 128 B/field  | grouping/alignment/layout vary              | measure actual layout/footprint               |
| Always stride eight longs               | assumes 64-B line/alignment                 | derive and verify target layout               |
| JOL alone proves two objects share      | relative field layout != absolute adjacency | address/topology or controlled array layout   |
| Padding before excluding CAS contention | true shared line remains                    | change semantics/ownership/striping           |
| Production thread count alone           | placement/socket/SMT matters                | validate actual topology distribution         |

## Definition of done

Use the relevant checks for the requested result; a supported no-change decision or a bounded
hypothesis with the next discriminator can complete a review. Neither is a proven performance fix.

- [ ] Independence, writer/reader topology and layout assumptions support the stated diagnosis,
      or unresolved evidence is identified explicitly.
- [ ] Coherence/PMU coverage and any separation experiment support only the claim actually made.
- [ ] If using annotation/padding, its flag/module/group/layout behavior is verified on target.
- [ ] If claiming a fix, correctness and the intended production metric survive relevant load/
      placement, with acceptable footprint, locality, allocation/GC and maintenance costs.

## References

- [`@Contended` mechanics and layout](references/contended-mechanics.md) — read when applying the annotation, choosing groups, or checking ignored padding/module access.
- [Proving and fixing false sharing](references/proving-and-fixing.md) — read when designing the separation experiment or interpreting PMU/JMH results.
- [JEP 142: Reduce cache contention on specified fields](https://openjdk.org/jeps/142)
- [OpenJDK 25 `Contended`](https://github.com/openjdk/jdk/blob/jdk-25-ga/src/java.base/share/classes/jdk/internal/vm/annotation/Contended.java)
- [OpenJDK 25 Striped64](https://github.com/openjdk/jdk/blob/jdk-25-ga/src/java.base/share/classes/java/util/concurrent/atomic/Striped64.java)
- [Java 25 LongAdder contract](https://docs.oracle.com/en/java/javase/25/docs/api/java.base/java/util/concurrent/atomic/LongAdder.html) — quiescent totals and concurrent sum/reset limits.

Files in this skill

  • SKILL.md12.2 KB
  • references/contended-mechanics.md6.9 KB
  • references/proving-and-fixing.md4.9 KB
  • skill.yaml1.5 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…