Load, run, and test VST3 / AU / CLAP / LV2 plugins from Pulp code. Use
when working on `core/host/` (scanner, plugin_slot, signal_graph), when
adding a new format backend, when wiring a plugin into a SignalGraph, or
when writing an integration test that needs a real plug-in binary.
Installs into .claude/skills of the current project.
Are you the author of Hosting?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/danielraffel-hosting)
---
name: hosting
description: |
Load, run, and test VST3 / AU / CLAP / LV2 plugins from Pulp code. Use
when working on `core/host/` (scanner, plugin_slot, signal_graph), when
adding a new format backend, when wiring a plugin into a SignalGraph, or
when writing an integration test that needs a real plug-in binary.
---
# hosting
## When this skill applies
- Adding or modifying a format backend under `core/host/src/plugin_slot_<format>.cpp`.
- Extending `PluginSlot::load()` to handle a new format in `core/host/src/plugin_slot.cpp`.
- Building or routing nodes in `SignalGraph`.
- Writing tests that need to load a real plug-in binary.
## Mental model
`PluginSlot` is the uniform interface. Each format backend is a single
free function — `load_<format>_plugin(info)` — that returns a
`std::unique_ptr<PluginSlot>` or `nullptr`. `PluginSlot::load()` in
`plugin_slot.cpp` is a small compile-time dispatcher. There is no dynamic
registry and no plug-in-per-file hooks; adding a format means:
1. Write `core/host/src/plugin_slot_<fmt>.cpp` that defines
`std::unique_ptr<PluginSlot> load_<fmt>_plugin(const PluginInfo&)`.
2. Add the file to `core/host/CMakeLists.txt` under a `if(PULP_HAS_<FMT>)`
guard. Link the format's SDK. Define `PULP_HOST_HAS_<FMT>=1`.
3. Forward-declare the loader and add a `case PluginFormat::<FMT>:` to
the dispatcher in `plugin_slot.cpp`, guarded by the same macro.
Everything else — tests, scanner, graph wiring — is format-agnostic.
## CLAP reference backend
`plugin_slot_clap.cpp` is the simplest backend to study for dlopen,
factory lifetime, parameter metadata, automation, MIDI, and state patterns.
VST3 / AU / LV2 also have real loaders, so treat CLAP as a reference for
its ABI shape rather than as the only implemented backend. Patterns to mirror:
- `dlopen(RTLD_LAZY | RTLD_LOCAL)`; on macOS resolve
`<bundle>.clap/Contents/MacOS/<name>` before `dlopen`.
- `dlsym("clap_entry")`; call `entry->init(path)` exactly once before
`entry->get_factory(...)`, and `entry->deinit()` + `dlclose()` in the
slot's destructor.
- Pick factory descriptor by `info.unique_id` when set, else first
available. Then fill the returned `PluginInfo` with any missing
name / vendor / version / id fields from the descriptor.
- The slot must own the `clap_host_t` it exposes to the plug-in; the
plug-in stores the pointer and will deref it later.
- After a successful `CLAP_EXT_STATE` restore, clear any cached host
parameter edits in the slot. Otherwise `get_parameter()` can report a
stale host-side value even though the plug-in restored its own state.
### VST3: re-sync a separated edit controller on state restore
VST3 splits a plug-in into an `IComponent` (processor) and an
`IEditController`. When they are *separate objects*, restoring only the
component state (`IComponent::setState`) leaves the controller — and the
vendor UI — showing stale values. The host contract is: after
`component->setState`, push the same processor state into the controller with
`IEditController::setComponentState`, and separately save/restore the
controller's own `getState`/`setState`. Combined plug-ins (one object
implementing both interfaces) skip all of this — see the identity trap below
for how to detect the combined case.
A separated controller also needs `setComponentState` at **load**, not only on
restore: a factory-created controller comes up on its own defaults, so its
parameter cache — and the editor about to open on it — shows values the
processor will not render. `vst3_push_component_state` (in `vst3_state_sync.hpp`)
is that push and `Vst3Slot`'s constructor calls it.
### VST3: combined-vs-separated is an FUnknown QUERY, never a pointer cast
`static_cast<FUnknown*>(component) == static_cast<FUnknown*>(controller)` looks
like the identity test and is wrong. A combined plug-in inherits `IComponent`
and `IEditController` *separately*, so each interface carries its own `FUnknown`
base subobject at a **different address** — the cast-and-compare therefore calls
every combined plug-in "separated". Verified against the SDK's own
`SingleComponentEffect`. COM identity is defined by the query: ask both sides
for `FUnknown::iid` and compare the returned pointers (`vst3_same_object` /
`vst3_is_separated` in `vst3_connection.hpp`; release both +1s).
Getting it backwards is not cosmetic — a combined plug-in misread as separated
gets `IPluginBase::terminate()` called **twice** on unload, has its state stored
twice, and would be connected to itself.
### VST3: connect the separated halves, and give them a real IHostApplication
A separated plug-in's two halves only reach each other through
`IConnectionPoint` — the host queries both, connects each to the other, and
disconnects before either terminates (`vst3_connection.hpp`). Everything a
plug-in cannot express as a parameter (preset banks, meter feeds, editor
handshakes) travels as `IMessage` over that link, so an unconnected plug-in
opens an editor that can never talk to its own processor.
Those `IMessage` / `IAttributeList` objects are allocated by the host, through
`IHostApplication::createInstance`. A host-application stand-in that returns
`kNotImplemented` there silently makes the connection useless, so `Vst3Slot`'s
`HostApp` derives from the SDK's `Vst::HostApplication` (which also supplies
`IPlugInterfaceSupport`) and only overrides `getName` plus the refcount, because
the instance is a process-wide singleton a plug-in must not be able to delete.
That pulls `hostclasses.cpp` + `pluginterfacesupport.cpp` + `stringconvert.cpp`
+ `commonstringconvert.cpp` into the `vst3-sdk` CMake target — the last two are
transitive and only show up as a link error in an unrelated tool.
`Vst3Slot` lives in an anonymous namespace, so this logic lives in a
free-function seam, `pulp::host::detail::vst3_state_sync.hpp`
(`vst3_serialize_state` / `vst3_restore_state`) — that is also the only place a
host-slot test can reach it. It writes a versioned `PV3S` container (component
plus an optional controller section); a blob without the magic is treated as
legacy raw-component state, so sessions saved before the container existed
still load. Length fields are bounds-checked as `len > remaining` (never
`remaining - len`) so a malformed blob cannot underflow into an out-of-bounds
read.
### VST3: embed the editor like CLAP — a parent-consuming IPlugView
`IEditController::createView("editor")` hands back an `IPlugView` that, like
CLAP `set_parent`, CONSUMES a parent: `IPlugView::attached(container, "NSView")`
inserts the plug-in's view into a host-owned container rather than returning a
view. So the VST3 editor path reuses the same container as CLAP
(`create_editor_container`), and the ordering matters — query
`IPlugView::getSize` first to size the container, create the container (already
in the parent window), THEN `attached()`. Tear down in reverse: `removed()`
before `release()`, and close the editor before terminating the controller it
came from (`createView`'s view must not outlive its controller).
**The host MUST install an `IPlugFrame` before `attached()`.** `IPlugView`
documents `IPlugFrame::resizeView()` as callable from *inside* `attached()` —
it is how a plug-in reports the size it really wants — so a view with no frame
either mis-sizes itself or refuses to attach outright, and a refused `attached()`
is exactly the "editor returns null, nothing embeds" symptom. `setFrame` goes on
immediately after `createView`, before `isPlatformTypeSupported`; every release
path clears it first (`vst3_release_editor_view` does the `setFrame(nullptr)`)
so a plug-in can never call into a destroyed frame. Because the plug-in may
resize during `attached()`, the slot publishes `editor_view_`/`editor_container_`
*before* the attach call and reports the post-attach size, not the size queried
before the view existed.
The AppKit part is only the container; the VST3 negotiation (createView,
`setFrame`, `isPlatformTypeSupported`, `getSize`, `attached`,
`onSize`/`checkSizeConstraint`, `removed`) is a pure-interface seam,
`pulp::host::detail::vst3_editor.hpp`, which
is where a headless test drives it with a fake `IEditController` / `IPlugView`
(inherit the SDK `EditController` / `CPluginView` bases). The full container
attach is only reachable with a real native window, so it is proven by the
real-DAW smoke, not the unit test.
### AU: the Cocoa UI RETURNS a view — adopt it, don't offer a parent
An AUv2's editor is a Cocoa-view factory, the mirror image of CLAP/VST3: query
`kAudioUnitProperty_CocoaUI` for an `AudioUnitCocoaViewInfo` (a bundle URL + a
class conforming to `AUCocoaUIBase`), load the bundle, and call
`uiViewForAudioUnit:withSize:` to get an NSView the plug-in already laid out. So
the host does NOT hand the plug-in an empty container to fill — it creates the
container and ADOPTS the returned view into it (`editor_container_adopt_view`,
which sets an autoresize mask so the view tracks the container). Two traps: the
property hands back +1 CF references (the bundle URL and the class name) —
`CFRelease` both on every exit path; and destroying the container already
releases the adopted view (it is a subview), so the slot keeps no separate view
handle. The Cocoa-UI negotiation needs a real AU with a view, so it is proven by
the real-DAW smoke; only the container adoption is unit-tested.
### Defensive boundary for entry / factory calls
`scanner_clap.cpp` wraps `entry->init()` and `entry->get_factory()` in
`try/catch`. Throws across the dlopen boundary abort the whole scan
otherwise — observed in production with bundles whose static-init throws
C++ exceptions during `dlopen`. The fallback
emits a synthesized `PluginInfo` (filename-derived name, no metadata)
so the scan still surfaces the bundle. Static-init throws that fire
*before* `dlsym` returns can't be caught at this layer; that's the
case `pulp scan --no-load` exists for. When adding new entry-point
calls, wrap them too — the goal is "one bad bundle never crashes a
scan."
### dlerror() must be cached
**Never** call `dlerror()` more than once per failure log line. POSIX
clears dlerror's internal buffer after every call, so a ternary like:
```cpp
// WRONG — second dlerror() returns nullptr
runtime::log_warn("dlopen failed: {}",
dlerror() ? dlerror() : "unknown");
```
calls dlerror() twice and the SECOND call returns `nullptr`.
`std::format`'s `string_view(char const*)` ctor then runs
`strlen(nullptr)` and ASan flags a SEGV in `libsystem_platform`'s
`_platform_strlen`. **The bug is invisible on non-ASan builds** —
release builds happen to survive the near-null strlen because the
zero page is read-protected one access deep. The fixed behavior caches
`dlerror()` before formatting so the null-returning second call cannot
reach `std::format`.
Correct idiom:
```cpp
const char* err = dlerror();
runtime::log_warn("dlopen failed: {}", err ? err : "unknown");
```
The other format backends (`plugin_slot_vst3.cpp`, `plugin_slot_lv2.cpp`,
`plugin_slot_clap.cpp`, `core/runtime/src/dynamic_library.cpp`) all
already cache via a local. When adding a new format backend that calls
`dlopen`, mirror the cache-into-local pattern.
## Testing against a real plug-in
For repeatable black-box measurements of an installed Audio Unit instrument,
build or install `pulp-au-instrument-probe`. It renders offline without opening
an audio device, requires an explicit `--name`, lists vendor parameter IDs, can
apply plain-domain parameter values and timestamped MIDI hits, and writes a
local float WAV. By default an accidentally silent render is a failure;
`--allow-silent` is reserved for experiments where silence is itself expected.
```bash
pulp-au-instrument-probe --name "Reference Instrument" --list-params \
--note 60 --seconds 2 --hits "0:100,250:80" --out /tmp/reference.wav
```
The probe is a bench oracle, not an implementation oracle: keep commercial
renders out of version control, record the recipe and numeric measurements,
and derive DSP from published specifications or independently authored models.
Reference-specific names, parameter maps, and corpora belong in the private
validation project, not in the SDK tool.
For unattended, scriptable interrogation prefer the isolated CLI/MCP surfaces:
```bash
pulp audio plugin-inspect --plugin /path/to/plugin.component --format au
pulp audio render --plugin /path/to/plugin.component --format au \
--input-signal noise:7 --duration-ms 1000 --warmup-ms 1000 \
--initial-param 12=0.5 --settle-ms 250 --wav-format float32 --out /tmp/out.wav
```
`plugin-inspect` reports the complete host-visible parameter API. Both commands
instantiate vendor code in disposable child processes with timeouts. That is
crash/hang containment, not a security sandbox. The rich `pulp audio compare`
step is an optional Audio Quality Lab add-on; inspection, rendering, metrics, and
`pulp audio validate compare` are stock Pulp.
Integration tests gate on a compile-time path macro:
```cmake
if(PULP_BUILD_TESTS AND NOT ANDROID AND TARGET PulpGain_CLAP)
foreach(_pulp_clap_host_test IN ITEMS pulp-test-host pulp-test-host-regression)
if(TARGET ${_pulp_clap_host_test})
target_compile_definitions(${_pulp_clap_host_test} PRIVATE
PULP_TEST_CLAP_PATH="${CMAKE_BINARY_DIR}/CLAP/PulpGain.clap")
add_dependencies(${_pulp_clap_host_test} PulpGain_CLAP)
endif()
endforeach()
endif()
```
Keep this wiring after `add_subdirectory(examples)`: the top-level build
registers `test/` before `examples/`, so `test/CMakeLists.txt` cannot
reliably see `PulpGain_CLAP` at configure time.
Tests check `fs::exists(PULP_TEST_CLAP_PATH)` and `WARN` + return if the
plug-in isn't built, so the suite still passes on configurations that
skip the plug-in builds (Android, CI without GPU examples, etc.).
Pattern for a process test: load, `prepare(48000, 256)`, fill an input
buffer with non-zero samples, call `process`, assert the output buffer
has non-zero energy. A gain plug-in is the cheapest target — one param,
predictable output, no MIDI.
For a quick manual load/inspect of one installed plug-in, `examples/plugin-host-demo`
doubles as a format-neutral analyzer. `--path <bundle>` loads a CLAP / VST3 / LV2
by inferring the format from the bundle extension; an Audio Unit has no bundle
path, so select it with `--id TYPE:SUBT:MANU` (the OSType triplet printed by
`--list`). `--path` deliberately skips the full installed-plugin scan — a bulk
scan runs third-party discovery code across the machine, which is unrelated to a
trusted, user-named probe. Two extension gotchas when inferring format from a
path: shell tab-completion appends a trailing `/` to a bundle *directory*, which
empties `std::filesystem::path::extension()` (step up via `parent_path()` when
`filename()` is empty), and a plain `dlopen` of a relative bundle path triggers
the `@rpath` search dance — pass an absolute path.
`--editor` embeds the loaded plug-in's own editor in a window via the
hosted-editor path (`create_hosted_editor` → `EditorAttachment`), auto-closing
after `--editor-ms` (default 3000). It is the manual smoke for the whole
host-side editor chain (CLAP / VST3 / AU): the negotiation seams are unit-tested
headlessly, but the editor actually rendering needs a display and a real plug-in
GUI, so run `--editor` (or load a Pulp-hosted plug-in in a DAW) to confirm it by
eye. Heads-up: opening an editor can make the plug-in active and audible — the
bounded duration keeps that contained.
## Headless Audio Unit event-loop servicing
Licensed Audio Units may complete initialization asynchronously via XPC,
timers, or dispatch-main callbacks. GUI hosts service these naturally; an
offline analyzer may otherwise produce plausible but incorrect audio or ignore
early parameter writes without reporting an error.
Use `pulp::events::MessageLoopIntegration::pump_main_loop_for()` from the
process main/control thread to service bounded slices before parameter access
and between offline render blocks. Never call it from the audio callback. The
result reports event-loop progress only, not license or plug-in readiness, so a
tool must own a configurable warm-up and post-write/render-settle policy. Put
the entire instantiate/process probe in a child process; isolated scanning alone
does not contain crashes in deeper plug-in code.
## Catalog headers: the analog VCF is not part of the lo-fi catalog
`forge_analog_vcf_catalog.hpp` is a sibling of `forge_lofi_catalog.hpp`, not
something it re-exports. Both put their symbols in `pulp::host::forge_lofi`, so
a translation unit that includes only the lo-fi header and reaches for
`make_analog_vcf_node()` will fail to compile in a way that looks like a
namespace problem. **Include the analog header directly.**
The separation is deliberate: the lo-fi catalog is a grab-bag of small macro-knob
effects, while the analog VCF carries measured per-voicing calibration tables.
Letting lo-fi own that would make it the owner of analog DSP policy, and would
force every lo-fi consumer to compile the VCF.
Two things that bite when baking analog VCF nodes:
- **`cutoff` is a 0..1 panel knob, not Hz.** Every calibration law — the corner
table, resonance taper, headroom and cross-mod laws — is indexed by
front-panel position, so the knob domain is what keeps the node faithful to
the measurements. If a UI needs Hz, invert the table in the display layer.
`AnalogVcfT::cutoff_hz()` gives the *requested* corner for that purpose, and
is not necessarily the realised one (the pole is clamped to 0.45·fs).
- **Voicing and oversampling factor are construction/prepare-time, not baked
params.** There is one `type_id` per voicing (`vcf.juno`, `vcf.jupiter`,
`vcf.prophet5`, `vcf.minimoog`); you pick the voicing by picking the node
type, the same way the registry's `svf` row realizes per-mode type ids.
## Signal graph gotchas
### A hosted device's latency is discovered once, on the control thread
`PluginSlot::latency_query()` exists because `latency_samples()` returns an
`int` either way and so cannot separate "this device reports zero" from "this
backend has no way to ask". Ask the query first and fail closed on
`Unsupported` / `QueryFailed`; coercing an unanswered query to zero is what a
later alignment proof would then certify. The hosted LV2 backend really does
return `Unsupported`, so this is not hypothetical.
Read it on the CONTROL THREAD, at admission, exactly once, and cache it into the
prepared snapshot. The metadata accessors reach into the live plugin object
(VST3 `getLatencySamples()` and friends) and are unsafe from the audio thread —
`process()` must consume a prepared number, never a live slot. Range-check the
result against `CustomNodeType::kMaxLatencySamples`; do not declare a second
ceiling beside it.
Pulp as a host offers a hosted CLAP plugin only the GUI extension, so
`clap_host_latency::changed` is unreachable and a device that varies its latency
mid-stream is **not observable** until the next control-thread relatch. Two
consequences: a mid-stream change cannot be refused (there is nothing to refuse
on), and any test of change-handling needs a synthetic latency source rather
than a real backend.
The timeline binding builds on this to schedule compensated event streams; the
contract, including why an event-to-audio device contributes nothing to the
shift, is `docs/policies/event-stream-pdc.md`.
`TimelineGraphProcessCode::EventCompensationUnsupported` has exactly one cause:
a transport range that locates events by host beat, which gives a
document-sample shift no origin to land on. It once had a second — a read-ahead
that ran past an enabled loop's end — and that case no longer fails: the window
is folded to the post-wrap content and plays. So a refusal here IS evidence of a
host-beat range, and a loop-end crossing is not a reason to expect one.
### Timeline-owned built-in devices stay document-authoritative
`PluginFormat::BuiltIn` is the host-only identity for pathless Pulp devices. Its
wire spelling is `builtin`; Timeline's existing durable `DeviceKind` spelling
remains `built_in`. The only initial resolver key is
`pulp.instrument.basic`, with an empty path and an exact 0-input/2-output
shape. Do not add BuiltIn to external scan paths, bundle suffixes, the isolated
scanner/worker protocol, or CLI format selection.
Timeline lowering reads the immutable `Project` owner pinned inside the exact
`PlaybackProgram`; it never consults the latest project store. Resolver-owned
plugin NodeIds are derived graph state and must not enter Timeline JSON. Add and
prepare their slots inside `PreparedTopologyEdit`, publish graph plus binding
atomically, and remove only slots carrying that transaction's private ownership
marker. Teardown and replacement must let retired execution snapshots drain
before the slot's balanced release. An unchanged full `DevicePlacement` may
retain its node; any declaration drift requires a fresh prepare and must not be
accepted by content-only program adoption.
The initial contract accepts exactly one `PreFader` `EventToAudio` BuiltIn
placement, not bypassed, exact wet value 1.0, and no state reference. It creates
one MIDI edge and two stereo audio edges through the track mixer. Unsupported
positions, slot kinds, namespaces, keys, bypass/wet/state variants, multi-device
chains, and mixed caller/resolver ownership fail closed with distinct admission
codes. Keep a factory-failure atomicity test and a deterministic non-silent
save/reopen render proof beside the legacy caller-owned route controls.
- `SignalGraph` dispatches plugin nodes through the additive
`PluginSlot::process(format::ProcessBuffers&, ...)` overload. The default
implementation projects the active main input/output bus back to the legacy
`process(output, input, ...)` callback, and still calls legacy processing with
empty audio views for MIDI-only slots. Override the `ProcessBuffers` overload
when a hosted format or fixture needs direct bus metadata for sidechains,
auxes, surround, or multi-output products.
- **Canonical-executor routing (DEFAULT ON; `set_canonical_executor_routing_enabled`
toggles it).** The routed executor is the primary inter-node backend for every
eligible graph; it is bit-identical to the legacy walk for that subset AND reports
the same per-node `node_loads()` telemetry (the executor times each node's work via
a per-binding `AudioProcessLoadMeasurer` wired from the host's persistent node-load
map), so the default-ON flip is behaviour-preserving where it takes effect. Force it
OFF to render the walk — the routed-vs-walk parity oracles (`run_legacy`,
`signal_graph_block`) do that so the walk stays an independent reference. **EVERY
node kind SignalGraph produces is now eligible** — the only remaining walk triggers
are an unprepared graph or a routed snapshot/pool BUILD failure (e.g. a topology past
`GraphRuntimeLimits`); the walk is the deliberate reference/fallback for those, the
independent parity oracle, and is NOT slated for deletion. **`routed_walk_fallbacks()`**
counts blocks where a routed path was ELIGIBLE but its dispatch returned failure so
process() silently fell back to the walk — that degradation is invisible to the parity
test (the walk is both oracle and fallback), so this counter (plus a once-per-graph
debug warning) is the only signal an eligible graph stopped routing. It stays 0 in
healthy operation; a normal fallback (routing disabled / ineligible / walk-by-choice)
does NOT increment it. An eligible graph —
nodes AudioInput / AudioOutput / Gain / Plugin (a Plugin with NO live slot routes as
pass-through-or-zero via `custom_binding(nullptr)`, exactly matching the walk's
missing-plugin behavior) / MidiInput / MidiOutput / **Custom** (`CustomNodeType`,
stateless `process` or stateful `process_instance`; routed via `custom_binding`,
an unresolved/shape-mismatch custom node pass-through-or-zeros exactly as the
walk does; custom output regions are pinned `persistent_output` like plugins so
a partial writer keeps its stale tail), connections audio (feedforward,
one-block feedback, or sidechain — a sidechain edge routes as plain audio into
a higher input port of the destination plugin), MIDI (connect_midi event
edges), or parameter automation — sparse (connect_automation, two control
points) and dense (connect_audio_rate_modulation, per-sample) — can be driven
through the canonical `GraphRuntimeExecutor` instead of the legacy walk via
`set_canonical_executor_routing_enabled(true)`. Output is bit-identical to the
legacy walk (`signal_graph_executor_routing.{hpp,cpp}` translates the graph;
`test_signal_graph_executor_parity` is the guard). Plugin output slots are
pinned *persistent* in the buffer assignment (the `persistent_output` spec
flag), so a plugin that does not fully overwrite its output carries the same
stale tail across blocks that SignalGraph's per-node buffer does — the reason a
Plugin node needs a live slot to be eligible (a null-slot placeholder would
take the legacy pass-through-or-zero branch, which the executor does not
reproduce). A latency-reporting plugin IS eligible: the routed gather applies
the same per-connection plug-in delay compensation as the legacy walk
(per-node latency is propagated through the topology in the buffer assignment,
and each feedforward connection that needs it gets a delay ring sized in the
`GraphRuntimeBufferPool`), so fan-in paths of differing latency time-align
identically. Each pool ring's mutable samples + write cursor live in a
shareable state object while the RT lookup still returns raw pointers; this
is the seam `prepare_swap` uses off-RT to adopt identity-matched PDC history
without reading or copying live ring contents. MIDI edges route through
per-node MIDI scratch buffers owned by
the executor (`GraphRuntimeMidiScratch`); SignalGraph bridges its MIDI
mailboxes (inject_midi / extract_midi) around the routed call. External
per-block parameter events cross the same boundary through a per-node
`inject_parameter_events` mailbox. The routed call appends that publication
after executor-generated automation and commits its sequence only after the
whole dispatch succeeds, matching serial fallback and anticipation behavior.
Parameter automation routes through a `GraphRuntimeAutomationScratch`
(per-node parameter event queue + per-connection slew state + per-node dense
buffers): sparse edges sample the source at the block edges, map/slew/mix per
the connection's resolved bounds, and emit two control points; dense
audio-rate edges map every sample (through the same per-connection PDC delay
ring as audio), mix into a per-node buffer, and emit one event per sample —
both bit-identical to the walk and built into the same per-node event queue.
A node exceeding
`kMaxParamsPerNode` (64) distinct sparse OR dense params is kept on the legacy
walk.
- **Where the walk lives.** The legacy serial reference walk is no longer
inline in `process_impl`; it lives in
`core/host/src/signal_graph_reference_walk.cpp` and is entered via
`SignalGraph::run_reference_walk_` when no routed path takes a block. It is
kept deliberately INDEPENDENT of `signal_graph_executor_routing.{hpp,cpp}`
— do not share or merge its gather / PDC / feedback / MIDI / automation
execution with the executor. The MIDI-block helpers shared by both the
routed dispatch and the walk (`clear_midi_block`, `midi_block_has_drops`,
`copy_midi_block`) live in the shared header
`core/host/src/signal_graph_internal.hpp`. The dual-maintenance rule still
applies: any audio-output-affecting edit to the walk must be mirrored in the
executor (and vice versa), guarded by `test_graph_routing_differential_parity`,
`test_signal_graph_executor_parity`, and `test_signal_graph_offline_parity`.
- **Where the live-swap engine lives.** The no-silence topology-edit machinery —
the swap policy and scanned-plugin catalog, the staged-replacement pipeline,
the `begin_swap_edit` / `prepare_swap` / `abort_swap_edit` transaction with its
crossfade publish and rollback, `snapshot_is_plugin_reinit_free_locked_`, and
the load-admission gate — lives in `core/host/src/signal_graph_live_swap.cpp`,
so `signal_graph.cpp` keeps the topology/compile/process spine. The file
arrangement mirrors `signal_graph_reference_walk.cpp`, but the *reason* does
NOT: these are still ordinary `SignalGraph` members that name its private
nested types, only their definitions moved. **The reference walk's
independence rule does not transfer here** — there is no second
implementation to stay bit-exact against, and factoring shared code out of
live-swap into `signal_graph.cpp` (or `signal_graph_internal.hpp`, where
`prepare_midi_block_storage` already single-sources every graph MIDI block's
real-time capacities for both the compile path and the swap warm-up probe) is
a fix, not a violation. Everything in the live-swap TU runs on the CONTROL
thread.
- **Gap-free PDC carry is identity-based and conservative.** `connections_`
has a private parallel vector of monotonic identities; every insertion and
erasure must update both vectors. `CompiledGraph` snapshots those identities
beside `connections`. During `prepare_swap`, the old and candidate delayed
edge sets must form an identity-keyed bijection with equal delay sizes and
equal total graph latency. The candidate then shares, never copies, each old
domain's audio-thread-owned ring state: legacy `ConnectionDelay`,
`routed.serial.pool`, and `routed.parallel.pool`. Disconnect+reconnect mints a
new identity and is refused even when the public `Connection` values compare
equal. Because those domains keep independent histories, a PDC-active
`CompiledGraph` pins the execution domain chosen during `prepare()`; relaxed
routing toggles remain dynamic only for zero-PDC snapshots, and a live swap
that would change the pinned domain is refused. Feedback graphs,
routed-validity changes, and latency/delay-structure changes are also refused.
Tests: `test_signal_graph_pdc_swap_continuity.cpp`
uses D=97 with 64-frame blocks across all three execution domains and includes
the reconnect negative plus a concurrent swap hammer.
- **`_locked_` is a contract, and it is now asserted.** A `SignalGraph` helper
suffixed `_locked_` requires the caller to already hold
`graph_mutation_mutex_`; the convention now spans `signal_graph.cpp` and
`signal_graph_live_swap.cpp`, so it is easier to violate from the far side than
it was when one TU held every caller. Helpers whose call graph is entirely
internal open with `assert_graph_mutation_locked_()`, which reads a debug-only
owner record — so **take the mutex through `GraphMutationLock`, never a bare
`std::lock_guard`/`unique_lock`**. A bare guard locks correctly but leaves the
owner record empty, and the next `_locked_` helper you call asserts as though
you forgot the lock entirely (a debug-only false failure that reads like a real
one). Only `prepare_swap` calls `GraphMutationLock::unlock()` early, to drop
the lock before invoking user callbacks. Two `_locked_` helpers cannot assert —
`has_path_locked_` (reached via `would_create_cycle`) and
`total_declared_ports_locked_` (via `validate_generated_graph` /
`estimate_generated_graph_work_units`) — because those public entry points do
not lock; their suffix documents the internal contract only. (`node_load_mu_`
is a different mutex and is correctly taken with a plain `lock_guard`.)
- **`compile_()` is a mutator, so every caller holds the lock — including
the test hook.** Since sample-region quotienting made the executable
topology a private copy, `compile_()` writes each authored node's
`transport_sensitive` readback into `nodes_` through `node_mut_locked_`.
`prepare()`, `prepare_swap()`, and `PreparedTopologyEdit::prepare()` already
hold `GraphMutationLock` around it; `compile_snapshot_for_test` did not, and
the two tests built on it aborted in every Debug build for a week while the
Release gate — which compiles the assertion out — stayed green. Never hold
the mutex when calling that hook (non-recursive), and never call `compile_()`
bare. `test_signal_graph_prepared_swap.cpp` has a build-type-independent
probe: a custom type's latency query runs inside `compile_()` and a second
thread's `node_gain()` must be unable to complete until the hook returns.
When a lock assertion is the suspect, run the Debug build, not Release.
- **`CompiledGraph::routed` groups what is only ever valid together.** Each
`RoutedPath` (`routed.serial`, `routed.parallel`) owns its own `snapshot`,
`pool`, `plugin_ctx`, `custom_ctx`, and `valid` flag — driving one path's
snapshot against the other path's pool is the bug the grouping exists to make
hard, since the parallel path's assignment is reuse-free and the serial path's
is compact. The MIDI scratch, automation scratch, and MidiInput/MidiOutput node
lists sit on `routed` itself and are **deliberately SHARED** by both paths: the
plans are identical and only ONE path runs per block. That sharing is load-
bearing — anything that makes those structs carry per-path state (or that runs
both paths in a block) breaks it silently, because the MIDI mailbox bridge and
the automation queues would then interleave across paths.
- **`build_executor_snapshot` prefers the `ExecutorSnapshotBinders` struct.** The
positional overload still exists as a forwarder for unmigrated call sites, and
it is exactly where a resolver can go wrong quietly: several of its resolvers
are same-shaped `std::function<T*(NodeId)>`, so swapping two by argument
position compiles and mis-binds. Every binder field is optional and has a
documented fallback (`plugin_latency_for` / `plugin_params_for` empty means
"fall back to the live slot", which is fine for baked/anticipation callers but
NOT on the swap path, where the point of the cached accessors is that a
swap-time build makes no live `PluginSlot` metadata call).
- **Multi-path routing has an arithmetic guard.**
`test/test_signal_graph_audio_parity.cpp` (target
`pulp-test-signal-graph-audio-parity`) renders a fan-out/fan-in topology whose
paths carry distinct non-commutative transfer functions, and checks the output
bit-exactly against expectations derived by hand from the stimulus — on the
walk, the routed serial path, and the routed parallel path. Nothing in it is
captured from a build, so a routing change cannot move the bar with it: a
dropped, swapped, or reordered path fails. It also asserts
`routed_walk_fallbacks()` / `routing_executor_stats()`, since a routed case
that quietly degraded into the walk would otherwise pass vacuously (the walk is
both oracle and fallback).
- **Connection lane CLASSIFICATION is single-sourced** (distinct from the
execution-independence rule above). Which lane a host `Connection` carries —
audio / event(MIDI) / automation, plus the orthogonal feedback flag and the
dense-vs-sparse audio-rate split — is decided in ONE place: `classify()` in
`signal_graph_executor_routing.{hpp,cpp}`, returning a `ConnectionClass`. The
runtime structs carry this as a typed `graph::GraphRuntimeConnectionKind`
discriminator (`Audio`/`Event`/`Automation`, default `Audio`) instead of the
old independent `event`/`is_automation` bools; read it via the
`pulp::graph::is_event` / `is_automation_conn` / `carries_audio` accessors.
BOTH classification surfaces route through `classify()`: the executor-routing
gather (building `GraphRuntimeConnectionSpec`s) AND the compile-time
reference-walk edge bucketer in `SignalGraph::compile_` (audio / MIDI /
sparse-automation / dense-audio-rate / feedback buckets). The PDC/latency
passes share the matching `connection_affects_latency()` predicate. A
sidechain edge deliberately classifies as `Audio` (it is plain audio into a
higher dest port). This is single-sourced CLASSIFICATION only — the gather
math, PDC delay rings, and MIDI/automation evaluation stay independent and
dual-maintained per the rule above. New lane mappings are pinned by
`test_connection_classify`.
- **Transport-aware `process()`.** Alongside the no-transport
`process(out, in, n)` there is an additive
`process(out, in, n, const format::ProcessContext& transport)` overload. Both
delegate to one private `process_impl(..., const format::ProcessContext*)`; the
3-arg form passes `nullptr` and is bit-identical to its prior behaviour. When a
transport is supplied it populates the routed `ProcessBlock`
(`block.transport = &ctx`, `block.mode = ctx.process_mode`) so nodes that consume
it (e.g. a `ProcessorNode`, via `context_for_block`) see the host playhead, mode,
and render-speed hint. `block.render_speed` stays the numeric `1.0`: the
render-speed hint is categorical and travels through `*block.transport`, never the
multiplier. Nodes that ignore `block.transport` are bit-identical to the
no-transport path, so routed-vs-walk parity is unaffected. Under active
anticipation the transport stays LIVE: transport-sensitive nodes are excluded
from the ahead-rendered interior (see the per-node opt-in below and the
anticipation gotchas), so every ahead-rendered node is transport-insensitive by
construction and the forwarding is inert for it.
- **Per-node transport opt-in (`PluginSlot::wants_transport()` + transport
`process()` overload / a transport-aware custom callback).** A routed plugin or
custom node OPTS INTO the host transport: a `PluginSlot` overrides
`wants_transport()` to return `true` and overrides the appended
`process(ProcessBuffers&, midi_in, midi_out, param_events, n, const
format::ProcessContext&)` overload; a custom node's type sets
`process_transport` (stateless) or `process_instance_transport` (stateful).
`compile_` resolves the capability ONCE into the cached, prepare-stable
`GraphNode::transport_sensitive` (Plugin: from `slot->wants_transport()`;
Custom: from a non-empty transport callback registration), resolved BEFORE the
anticipation eligibility analysis. That ONE cached bit is read by BOTH the
routed binding (`PluginBindingContext::wants_transport` /
`CustomBindingContext::process_transport`, which forward the live transport when
the block carries one) AND the anticipation analyzer (which seeds
`AnticipationExclusion::TransportSensitive`), so the partition and the bindings
can never disagree. INVARIANT: never call a live `slot->wants_transport()` per
block on the audio thread — the bit is cached at compile and a capability change
requires a re-prepare. A node that does not opt in is byte-for-byte unchanged.
- **Parallel-executor routing (opt-in, default OFF, independent of the serial
opt-in).** `set_parallel_routing_enabled(true)` drives the SAME eligible subset
through `GraphRuntimeExecutor::process_parallel` — a levelized fork-join over a
persistent `GraphRuntimeWorkerPool` (the audio thread is participant 0). Output
is bit-identical to the serial executor and the legacy walk; the per-node body
(`run_routed_node`) is shared. Dispatch order in `process()`: parallel (if
enabled + valid + pool running + fits) → serial executor (if its toggle on) →
legacy walk. The two routed branches share one `dispatch_routed` bridge, and
every executor zeroes the output bus + the MIDI ingress is idempotent (consumed
mailbox sequences aren't committed until `run()` succeeds), so a failed parallel
attempt re-renders the block on a lower tier with no doubled output or
double-consumed MIDI. `SignalGraph::set_parallel_min_work_units(n)` forwards
to the executor's channel-sample break-even gate; default `0` preserves the
original "parallelize every eligible level" behavior, while a positive value
keeps low-cost levels serial to avoid fork/join overhead on small graphs. Use
`routing_executor_stats()` to verify the live path when testing the threshold.
GOTCHAS: (1) the parallel snapshot uses a REUSE-FREE
buffer assignment (`parallel_safe=true`) — concurrent same-level nodes must not
alias a recycled scratch slot; `process_parallel` refuses a non-parallel-safe
snapshot. (2) Levels containing an AudioOutput node run SERIALLY in topo order
(AudioOutput `+=` accumulates into the shared output bus; float add is
non-associative, so order is load-bearing for ≥3 sinks). (3) WORKER-POOL
LIFECYCLE is load-bearing: the pool is started ONCE (size = clamped hardware
concurrency, guarded by `worker_count() == 0`) and NEVER stopped/resized on a
re-prepare — `start()`/`stop()` join threads + reset the epoch/completion
counters, a UAF if run against an in-flight audio `run()`. The only legal stop
is `~GraphRuntimeWorkerPool` at SignalGraph destruction. Don't make the thread
count runtime-variable without a drain handshake. The pool's completion barrier
counts PARTICIPANTS finished (not tasks): an empty-range participant must still
register done, or it can race the next batch's published state.
(4) WORKGROUP CHANGES are generation-published, not applied from the caller:
`SignalGraph` implements `format::AudioWorkgroupClient`, and each persistent
worker leaves/joins on its own thread. `run()` executes inline while any worker
still advertises an older generation, so an AU `renderContextObserver` change
cannot dispatch a deadline into the previous workgroup. A failed non-null join
does not acknowledge the generation: the worker retries and `run()` stays
inline. For an explicitly owned device, publish null and call the
off-render-thread acknowledgment barrier before switching or closing it; only
then may the borrowed OS handle be invalidated. Re-query and publish the
replacement before rendering resumes. Do not cache a device-owned handle past
that drain point. Close must first disable new device-change notifications and
serialize with any switch already in flight, then publish null and drain again
under that serialization boundary; an external null publication alone can race
a switch which rebinds immediately before close. AU render-context teardown
remains publication-only.
- **Anticipative-rendering eligibility (`anticipation_eligibility.{hpp,cpp}`).**
`analyze_anticipation_eligibility(nodes, connections)` is the static SAFETY
contract for rendering a latent subgraph ahead of the audio deadline: it
classifies each node `None` (passed) or a hard-exclusion reason — seeds live
AudioInput/MidiInput nodes, both endpoints of every feedback edge, any node with
a sidechain inbound edge, and any node with `GraphNode::transport_sensitive` set
(the per-node host-transport opt-in), then propagates each exclusion forward
along feedforward (non-feedback) edges to a fixpoint so anything downstream of an
excluded node is excluded too. It's deliberately conservative: a false exclusion
only forfeits a speed-up, but a false inclusion would render an unsafe node
ahead. Host-clock sensitivity is handled by the `TransportSensitive` seed: a
host-clock-dependent node opts in via `wants_transport()` / a transport-aware
custom callback (resolved into `transport_sensitive` at compile), so it — and
its downstream cone — is kept out of the ahead-rendered interior and runs live.
`passes_static_exclusions(i)` true therefore IS sufficient for the partition to
treat node i as ahead-renderable. The `SignalGraph` anticipative splice gates on
this analysis when `set_anticipation_enabled(true)` is prepared.
- **Anticipation partition (`anticipation_partition.{hpp,cpp}`).**
`build_anticipation_partition(nodes, connections, eligibility)` carves the
renderable eligible INTERIOR (eligible nodes minus the live AudioOutput/MidiOutput
sinks, which are consumed at the real deadline and must never be written ahead)
and the BOUNDARY edges (interior-source -> outside-the-interior), which are the
splice points the renderer pre-computes and the live graph reads. `cost_weight`
(the same coarse max(in,out) proxy the parallel cost gate uses) +
`worth_anticipating()` gate out trivial/no-boundary partitions. Still pure static
analysis — no rendering, no RT path. Builds on the 6a eligibility result and is
rejected (ok=false) if that result isn't ok or doesn't match the node span.
- **Anticipation sub-graph (`anticipation_subgraph.{hpp,cpp}`).**
`build_anticipation_subgraph(nodes, connections, partition)` turns a partition
into a standalone renderable graph: it copies the interior nodes verbatim (plugin
slots/gain/ports preserved) and the internal edges, and synthesizes ONE
`AudioOutput` sink whose input ports correspond to the DISTINCT boundary output
ports (fresh id above every existing node id, so no collision), fed so boundary
output `i` lands on sink input/output-bus channel `i` — so the sub-graph renders
through the ordinary `build_executor_snapshot` + `process_routed` and its output
bus carries exactly the boundary signals without summing them together.
`outputs[]` maps each output-bus channel back to the `(source_node, source_port)`
it captures. GOTCHA: the interior plugin
GraphNodes are copied by value, so the SAME plugin instances render here — which
means a live splice (a later slice) must NOT also process those instances, or
their state double-advances. This slice does extraction only; it neither renders
nor changes any RT path.
- **Anticipation lane (`anticipation_lane.{hpp,cpp}`).** `AnticipationLane` renders
an eligible sub-graph AHEAD of the deadline into a `PlanarAudioRingBuffer`:
`prepare()` (off-RT, quiescent) builds the executor snapshot + sizes the ring for
a FIXED block size; `render_ahead()` (single background producer) advances the
interior's plugin state and pushes whole blocks; `consume()` (audio thread,
RT-safe, no-alloc) pops a pre-rendered block or reports underrun so the caller
falls back to a synchronous render. The block size is PINNED at prepare (before
any thread exists) so producer/consumer stay in lockstep and there's no
cross-thread block-size field — the consumed sequence is bit-identical to a
block-by-block synchronous render. GOTCHAS: (1) the interior plugins are advanced
ONLY by render_ahead — a live splice that uses a lane must not also process those
nodes or their state double-advances; (2) render_ahead is SINGLE-producer (all
calls, including priming, must be serialized — they share unsynchronized
executor/pool/scratch); only the ring mediates against the consumer.
- **Anticipation splice (`set_anticipation_enabled`, default OFF; runs on the
canonical executor path).** When enabled + the routed snapshot is eligible + the
graph has an eligible latent interior, `compile_` builds an `AnticipationLane` +
a `skip_mask` over the routed plan (the interior nodes) + a prefill map (each
lane output channel → the interior boundary-source's `exec_pool` output slot).
The host drives `pump_anticipation()` from ONE background thread (the producer);
`process()` consumes a pre-rendered block, copies it into the prefill slots (or
zeros them on underrun / block-size mismatch), and runs `process_routed` with the
interior masked — bit-identical to the canonical interior-live render. GOTCHAS:
(1) the branch is TERMINAL once entered — on a routed failure it zeros the output
and returns rather than falling through to a path that would re-run (double-
advance) the producer-owned interior. (2) `pump_anticipation` pins the live
snapshot (RCU object-lifetime only) and is single-producer-guarded, but the host
MUST stop/join the pump before any `prepare()`/mutation — prepare reinitializes
the SAME plugin instances the pump renders (a data race otherwise; same rule as
"no `process()` during prepare"). (3) Host-clock-sensitive nodes opt in via the
per-node transport capability (`wants_transport()` / a transport-aware custom
callback → cached `GraphNode::transport_sensitive`). The eligibility pass seeds
`AnticipationExclusion::TransportSensitive` on that bit, so a transport-sensitive
node — and its downstream cone — is EXCLUDED from the ahead-rendered interior and
always runs live/exterior, where it receives the live transport. The former
blanket transport suppression under anticipation is therefore RETIRED: the
transport stays populated on every block (inert for the transport-insensitive
interior). `transport_suppressed_for_anticipation()` is repurposed to count the
transport-sensitive nodes anticipation forced exterior (resolved at compile),
not per-block drops.
(A masked node must not be an `AudioOutput` or a feedback endpoint; the
partition guarantees this and `process_routed` debug-asserts it.) (4) The lane
uses a FIXED block size (the prepared max). A block of a different size — or a
ring underrun — silences the interior for that block (the interior is never
re-rendered live, so bit-identical-to-canonical holds only for fixed-size,
kept-up blocks); and an interior param/gain edit takes effect at render-ahead
time, a lead earlier than a live render. The anticipation branch is
STRUCTURALLY terminal once `anticipation_valid` — it never falls through to the
parallel/legacy paths (which would run the producer-owned interior live), even
if the pool can't fit the block (then: silence).
- `connect()` returns `false` on cycle — always check. `would_create_cycle`
lets you preview without mutating.
- `processing_order()` is recomputed each call; cache it in the audio
thread, don't recompute per block.
- Removing a node invalidates its `NodeId`. Connections referencing a
removed node are pruned automatically.
- For fail-closed structural publication while audio remains live, use
`begin_prepared_topology_edit()` instead of mutating the owner eagerly. Build
the complete candidate through its `PreparedTopologyEdit`, call
`prepare(sample_rate, max_block_size)`, verify
`routed_execution_ready(max_block_size)` when the caller holds a routed-only
lease, then `commit()`. A failed mutation poisons the one-shot edit, and
destroying it before commit rolls back topology, private connection IDs,
next IDs, the custom registry, routing flags, and the compiled snapshot.
Existing `PluginSlot`s are never re-prepared off-side; dimension changes are
accepted only without plugin nodes and when every retained custom type has
neither `prepare` nor `release`. New edit-owned custom instances may be
prepared. Unchanged PDC rings carry by private connection identity plus equal
shape in the same execution domain; removed rings retire, while new,
reshaped, or reconnected rings start fresh. Preserve connection identity for
unchanged owned routes and prune unused generated custom types in the same
edit so registry churn stays bounded. `MidiInput` ingress may carry through
its shared sequence mailbox, but `MidiOutput` egress is snapshot-local so an
old snapshot can expose its pending output exactly once. A prepared edit
returns `MidiOutputSnapshotLocalRequired` before callbacks whenever either
the live or candidate graph contains a `MidiOutput`; do not adopt or share
that output mailbox. Baseline plugin removal, and baseline custom removal
when its registered type has `release`, are also explicit fail-closed
results: ordinary graph release remains the only lifecycle-callback owner.
Commit's exception boundary is before authoring mutation: preflight the
generic live `runtime::Slot` retirement capacity and reserve destination
`node_load_` buckets first. After every failure gate, transfer load measurers
with C++17 unordered-map node handles (matching integral hash/equality and
allocator), move the authoring containers, and call the Slot's noexcept
prepared publication. Never put an allocating `emplace` or ordinary
`Slot::publish` after that boundary.
`set_live_dsp_telemetry_enabled()` is a control-thread operation serialized
by the graph mutation lock: a prepared commit re-seeds its snapshot from the
owner's authoritative desired toggle immediately before publication. Do not
move the toggle outside that lock or restore the candidate's creation-time
value; either change can update the retired snapshot while publishing stale
telemetry state.
A caller that has stopped audio processing and anticipation may instead use
`prepare_quiesced()` for a dimension change involving external plugins or
retained custom instances. Candidate preparation can touch those shared
lifecycle objects, so *every non-commit exit* — candidate failure, routed
rejection, commit rejection, exception, or simple edit destruction — must
restore each retained object whose candidate prepare callback was entered
before the old snapshot resumes. Track entry immediately before invoking user
code: candidate preflight can fail before every callback, and plugin/custom
preparation can stop midway. On an unprepared base, releasing an untouched
retained object is an unbalanced lifecycle call just as surely as failing to
release a touched one. A successful candidate prepare is not ownership
transfer; only a successful commit cancels that rollback obligation. If
restoration fails, the graph deliberately unpublishes and reports
`QuiescedRollbackFailed`; a coupled binding must unpublish too. Never resume a
partially restored graph. New custom instances created before a later prepare
failure still require their control-thread `release` callback.
- Per-node CPU load: `process()` wraps each node's work in a persistent
per-node `audio::AudioProcessLoadMeasurer` (keyed by `NodeId` in
`node_load_`), read via `node_loads()`. The measurers live on the
SignalGraph (not the snapshot) and `compile_()` only ever ADDS to the map —
never erase while a snapshot may be live, or the audio thread's raw
`NodeRuntime::load` pointer dangles. `begin()/end()` are relaxed-atomic and
RT-safe (proven under the no-alloc trap in test_signal_graph_rt_safety).
- Per-node live-DSP telemetry (`audio::LiveDspTelemetryStore`): richer than the
load measurer — fixed-slot p50/p95/p99 + jitter + over-budget attribution.
Disabled by default (`set_live_dsp_telemetry_enabled()`; the audio path is one
predicted-not-taken branch when off); drain + read a snapshot copy via
`poll_live_dsp_telemetry()` (single non-RT poller). Unlike `node_load_`, the
store is PER-`CompiledGraph` (not a SignalGraph member): it rides the RCU
snapshot lifetime, so telemetry resets on a topology recompile (a new topology
is a new timing baseline) and no re-prepare races the audio thread. Recording
is PATH-AGNOSTIC and lives at ONE site: a guard at the top of `process_impl`
destructs after the block (reverse-order vs the graph-load end guard) and reads
the values BOTH the routed serial executor and the legacy walk already stamped
into each node's persistent `AudioProcessLoadMeasurer` + `graph_load_`, then
pushes one fixed-slot record via `inject_block()` over a pre-sized scratch
(`external_record_scratch()`). This is why there is NO per-node hook in the
executor or the walk — both already time per node, so the store just harvests
those measurements once per block. `canonical_executor_routing_enabled_`
defaults TRUE, so the serial executor is the common path; harvesting the
measurers (not hooking the walk) is what makes telemetry work by default.
Per-node slot == the node's index in `ordered_runtime`, matching the store
metadata built at compile.
- `.pulpgraph` schema changes must go through the graph serializer migration
path. Bump the graph format version, add a deterministic migrator for older
fixtures, and keep future-version loads fail-closed instead of silently
accepting fields the current reader does not understand.
- Use `connect_automation()` for sparse two-point-per-block control events.
Use `connect_audio_rate_modulation()` only for continuous, automatable
`HostParamInfo::rate == AudioRate` params; do not route dense CV into
stepped/read-only/control-rate parameters.
- MIDI graph edges carry one block with three parallel payloads: short MIDI
events, SysEx, and optional UMP sidecars. When copying or clearing graph MIDI
scratch, handle all three together. If a `MidiBuffer` attaches a `UmpBuffer`
owned by `NodeRuntime`, attach it only after the runtime object is in its
final `CompiledGraph` storage; attaching before a move leaves a stale sidecar
pointer.
- `SignalGraph::inject_midi()` and `extract_midi()` cross the
control/audio-thread boundary through per-node mailboxes, not by mutating
audio-thread scratch directly. After `prepare()`, injection is `noexcept`,
fixed-capacity, lock-free, and allocation-free, so it may run on the audio
callback immediately before `process()`. Each MidiInput has exactly one
writer — audio or control, never concurrent. A `false` return means the live
node is unavailable or the source was truncated; any retained prefix is
still published. Publications are latest-wins and one-shot, and sequence
wrap skips zero. A gap-free `prepare_swap()` shares the ingress mailbox and
consumed sequence for a stable MidiInput NodeId, preserving unconsumed MIDI
across the swap. MidiOutput egress remains snapshot-local but uses an ordered,
fixed four-block SPSC queue: empty blocks cannot overwrite pending output;
overflow retains the earliest blocks and makes extraction incomplete. When
the destination lacks short-event, SysEx, or attached UMP-sidecar capacity,
extraction returns false and retains the undelivered suffix. Provide storage
and call `extract_midi()` again; it resumes without replaying the delivered
prefix. When
`prepare_swap()` returns `NeedsEagerPrepare`, the old live snapshot remains
valid, so drain it with `extract_midi()` before eager `prepare()` replaces it.
- `SignalGraph::inject_parameter_events()` uses a separate prepared per-node
mailbox with one control-side writer. Publications are latest-wins and
one-shot: the next successful serial, routed, parallel, or anticipation
block consumes the newest sequence once. Append injected events after graph
automation before the stable sample-offset sort; this preserves graph
automation when the fixed queue is full and lets injected events win a
same-offset tie. Reuse the mailbox across gap-free snapshots for the same
plugin node so an edit cannot discard a publication made before the swap.
- Keep plugin automation scratch preallocated by `SignalGraph::prepare()`.
The audio-thread `process()` path must not create per-block containers for
input pointer casts, sparse automation accumulation, or dense audio-rate
modulation accumulation.
- Signal-graph declarations are split by dependency weight:
`custom_node_type.hpp` owns the registration vocabulary,
`signal_graph_node.hpp` owns node identity and retained state,
`signal_graph_connection.hpp` owns edge metadata, and
`signal_graph_runtime.hpp` owns `SignalGraph` execution and mutation.
`signal_graph.hpp` is the compatibility umbrella. Include the narrowest
contract that supplies the names a header uses; do not move declarations
back into the umbrella.
- Custom graph nodes are registered per `SignalGraph` with `CustomNodeType`
(`type_id`, `version`, port counts, default name, optional process
callback), then instantiated with `add_custom_node(type_id)` or
`add_custom_node(type_id, version)`. `GraphSerializer` resolves exact
`(type_id, version)` matches with the saved port shape, preserves unresolved
custom identities as placeholder `NodeType::Custom` nodes, and reports them in
`LoadResult::missing_custom_node_types`, so do not coerce unknown node
strings to a built-in type. Runtime callbacks are attached only when the
registered version and shape match the node.
- **A duplicate designator in a binder initializer breaks the build — and BOTH
compilers told us, and it landed anyway.** The binders are designated-
initializer literals, so a merge that brings the same binder in from both
sides lists the member twice. That happened once with
`.custom_latency_for` in `signal_graph.cpp`.
Do NOT read this as "clang is permissive and only Linux catches it" — that is
the wrong lesson and it will send you looking for a diagnostic that was
already printed. What actually happened:
- **Clang WARNED, by default, twice.** Compiling the shape with no `-W`
flags at all yields `-Wreorder-init-list` ("field designators … in
declaration order") and `-Winitializer-overrides` ("initializer overrides
prior initialization of this subobject"). The root `CMakeLists.txt` already
adds `-Wall -Wextra -Wpedantic`, so the **required macOS gate printed both**
— and nobody reads a green job's log.
- **GCC ERRORED and the Linux job reported failure**
(`error: '.custom_latency_for' designator used multiple times in the same
initializer list`).
- It reached `main` regardless, because the Linux lane is **advisory**.
So this is a POLICY gap, not a coverage gap: a real compile failure was
detected, displayed, and merged past. Adding another GCC leg would be
building something that already exists. The durable fix is to make an
existing verdict binding — either promote those two clang warnings with
`-Werror=reorder-init-list -Werror=initializer-overrides` (free: clang
already emits them, so it can only fail on real defects), or make the Linux
COMPILE step blocking while leaving its tests advisory. Measure whether
Linux finishes inside the macOS leg's shadow before choosing; if it does,
blocking costs no latency.
There are **four** binder literals, not three — miss one and the grep is a
false all-clear:
```bash
grep -c '\.custom_latency_for' core/host/src/signal_graph.cpp # 1445
grep -c '\.custom_latency_for' core/host/src/signal_graph_executor_routing.cpp # 427 and 689
grep -c '\.custom_latency_for' core/host/src/baked_graph_processor.cpp # 623
```
If the duplicated bodies are identical, deleting either is behaviour-
preserving — clang was already using the last one. If they DIFFER, clang has
been running the second and GCC has never compiled the file at all, so decide
which is correct before deleting. Prefer keeping the copy that sits in
declaration order.
- **Custom-node intrinsic latency is a CALLBACK, not a count.**
`CustomNodeType::latency_samples` is `std::function<int(double sample_rate)>`,
returning a prepare-stable base-rate sample count regardless of internal
oversampling. It is a callback because intrinsic latency is routinely
rate-dependent — a lookahead declared in milliseconds, an oversampler's
half-band group delay — so a fixed count could only ever be correct at one
rate. A rate-independent latency is `[](double) { return n; }`. Empty means
zero, which preserves the historical zero-latency contract.
Three consequences that are easy to get wrong:
1. **Range is NOT checked at registration.** The registrars cannot range-check
a value that does not exist until a sample rate does. `is_valid_registration()`
checks structure only; the EVALUATED result is clamped into
`[0, kMaxLatencySamples]` where it is first produced — the live compile path
and `bake()`. Do not write a test asserting the registrar rejects an
out-of-range latency; assert the clamp against a prepared graph instead.
2. **A baked processor's latency is resolved by `prepare()`, not by its
constructor.** `BakedGraphProcessor::latency_samples()` reads
`prepared_latency_samples_`, which is 0 until prepare() supplies a rate.
Querying it straight after `bake()` legitimately returns 0 — that is the
honest answer to "how much latency at a rate I have not been told yet",
not a bug to work around by assuming 48 kHz.
3. **Latency attaches only when a callback will actually execute.** A
registered type with no live callback is transparent on the live graph and
must not add latency there; the same gate applies in `bake()` against the
node's plain OR baked-param callback. This is why a baked-only node can be
transparent live yet latent when baked.
`is_valid_registration()` is the ONE validation entry point — shape, the
transport-fallback rule, AND the bake-layer obligations (coherent
`baked_params`/`process_instance_baked_param` pairing, finite non-inverted
ranges, unique non-zero ids, complete lifecycle, `lowerable` runnability).
Every registrar routes through it: the direct one, the transactional prepared
edit, and the node-adding leaf. Do not add a parallel free-function validator
— an earlier split let a registrar validate shape only, which silently
accepted malformed baked-param declarations that then degraded to passthrough.
A transport-aware execution path that declares latency must also resolve a
plain fallback; PDC cannot vary with per-block transport presence. Changing
the value requires re-registration/re-prepare and a type version bump when
persisted graphs must distinguish the timing contract. Derive catalog values
from the DSP implementation (for example, Analog VCF's fixed-2x FIR), never a
duplicated host constant. Add dry/wet impulse parity coverage across walk,
routed, prepared-edit, and baked paths.
- **Stateful custom nodes.** `CustomNodeType` has an
optional lifecycle: set `create` and the graph owns one opaque instance per
node (RAII via `destroy`); `process_instance` runs instead of the stateless
`process`, and `prepare`/`release`/`reset`/`save_state`/`load_state` operate on
it. Empty callbacks = today's stateless node (no instance, no serialized
state). The instance is created/prepared on the UI thread inside
`SignalGraph::prepare()` (mirroring `PluginSlot`) and captured into each
`CompiledGraph` snapshot by `shared_ptr` — never allocate or create instances
on the audio thread, and never store a raw `GraphNode` pointer in the snapshot.
`process_instance` must be RT-safe; call `save_state`/`load_state` only on the
control path (graph not live, or after invalidate + re-prepare). Opaque state
is `std::vector<uint8_t>` via `SignalGraph::custom_node_state` /
`set_custom_node_state`; `GraphSerializer` persists it as `state_b64` and keeps
the blob even for **unresolved** nodes (save → load-missing-type → save keeps
state). Do not pull the `pulp_native_state_*` C ABI into `CustomNodeType`;
that belongs to the `pulp_node_v1` ABI.
- **Signed node-pack loader (`core/host/node_pack.{hpp,cpp}`).**
`load_node_pack(dir, manifest, trust)` derives its platform from the build
target and treats the pack as runtime-downloaded; pass the four-argument
overload with an explicit `NodePackHostPolicy` when the origin is known to be
app-bundled. A pack is a precompiled `pulp_node_v1` dynamic library exporting
`pulp_node_v1_entry` plus a JSON manifest. The generated-device policy runs
before trust verification or any binary file access: native packs are allowed
on desktop, allowed on Android only when `origin == AppBundled`, and denied on
iOS/web. Thus the three-argument overload fails closed for Android native
packs instead of treating downloaded code as bundled. After policy admission,
the signer key must be in the
`NodePackTrust` set, the Ed25519 signature over `node_pack_signed_message()`
(pack_id + abi_major + binary SHA-256 + declared nodes/resources/requirements)
must be authentic, the on-disk binary's SHA-256 must match the signed hash,
and the entry's `abi_major` must match — any failure returns a
`NodePackError` and loads nothing. Revocation = drop a key from the trust set.
`pulp-host` (and this loader) is compiled out on iOS, where native components
are static-bundled + signed with the app. The crypto
comes from `pulp::runtime` (`ed25519_verify`, `sha256_hex`); OS
codesign/notarization is a separate, additional distribution step on top of
the manifest signature. Registry/package discovery metadata still needs its
own signed canonical manifest; do not treat screenshots, validation reports,
licenses, or provenance as covered by the node-pack loader signature.
- **Routing a `SignalGraph` through the canonical executor
(`core/host/signal_graph_executor_routing.{hpp,cpp}`).** The eligible subset is
described above under "Canonical-executor routing" and enforced by
`signal_graph_topology_executor_eligible()` /
`signal_graph_executor_eligible()`. The builder fails closed for unsupported
Custom nodes, placeholder Plugin nodes, and per-node automation counts above
the fixed scratch caps. `build_signal_graph_executor_routing()` translates an
eligible prepared graph into a `format::GraphRuntimeSnapshot` + pre-sized
`GraphRuntimeBufferPool`; the live `process()` path embeds that snapshot and
its scratch pool per `CompiledGraph`, so a re-prepare rebuilds fresh routing
state without resizing buffers an in-flight audio reader holds. The routing
keeps the live compiled snapshot alive, reads live gain atomics, and invokes
the snapshot's live PluginSlots, so **rebuild routing after any re-prepare**
and keep this section aligned with `test_signal_graph_executor_parity`.
## Offline graph rendering (`OfflineSignalGraphHost`)
`core/host/offline_signal_graph_host.{hpp,cpp}` renders a prepared `SignalGraph`
offline by stepping a fixed block size across a frame range through the **public**
`SignalGraph::process()` — no live audio device, deterministic, allocation-free per
block (staging + output buffers are sized in `prepare()`). It is a control-thread
host, not a routing path: it adds no walk of its own and never touches graph
internals, so it stays clear of the in-flight routing/anticipation churn.
Gotchas:
- **Block-size silence clamp.** `SignalGraph::process()` zero-fills any block larger
than the prepared `max_block_size` (`prepared_max_block_size()`). `prepare()` refuses
if the configured `block_frames` exceeds the graph's prepared max — otherwise an
offline "one big block" render would silently drop to silence. To render one big
block, re-`prepare()` the graph at that block size first.
- **What "offline equals online" actually means here.** A `SignalGraph` carries no
`ProcessMode`/transport into its nodes, so an offline render is NOT distinguishable
from an online one by render mode — the only variable is the block partitioning. For
deterministic nodes, output is therefore block-size invariant: same input at any
block size → bit-exact for pure gain/sum, within ~1e-6 across re-partitioning. A node
whose output legitimately depends on block size (the exempt path) is declared
EXEMPT as harness-side metadata today (no per-node `ProcessMode` opt-out exists yet);
the equivalence harness flags and excludes it rather than failing.
- Keep the executor/parallel/anticipation opt-ins OFF for partition-invariance
fixtures — anticipation in particular is intentionally not block-size invariant.
## Baking a graph to a `Processor` (`BakedGraphProcessor`)
`core/host/baked_graph_processor.{hpp,cpp}` — `bake(const SignalGraph&)` freezes a
prepared, fully-lowerable graph into one `pulp::format::Processor` that runs a frozen
`GraphRuntimeSnapshot` through the SAME `GraphRuntimeExecutor::process_routed()` the
live graph uses, so baked output is bit-identical to the live graph for the lowerable
subset. The artifact is a *serialized fused plan* (data), not generated code — it
reuses the one backend, so the baked Processor only CALLS `process_routed`, never
defines a routing entry point.
Gotchas:
- **Lowerable subset is narrow by design.** Today: `AudioInput`/`AudioOutput`/`Gain`,
plus a `Custom` node whose registered type opts in (`lowerable = true`, shape
match, transport-independent). `bake()` REFUSES loudly (null processor + a
`LowerRejectReason`) for an unprepared or executor-ineligible graph, a hosted
`Plugin` node (opaque external state — not self-contained), or a `Custom` node that
does not meet the opt-in bar. The node-kind refusals are checked BEFORE the
eligibility predicate so a Plugin/Custom graph reports its specific reason instead
of a generic `NotExecutorEligible`.
- **The baked Processor owns its Gain values.** `bake()` copies each Gain's value into
the Processor; `prepare()` seeds one heap-stable `atomic<float>` per Gain (a
`unique_ptr` vector, never a value vector) and resolves the routed Gain bindings to
those owned atomics — so the baked Processor is independent of the source graph's
live snapshot lifetime. A second `prepare()` clears the old snapshot/pool/atomics
before rebuilding, so binding pointers never dangle.
- **Sizing mirrors live routing.** `prepare()` builds the snapshot via the same
`build_executor_snapshot()` the live routing uses and sizes the pool from
`buffer_slot_count()` × `max_buffer_size` plus the per-connection PDC rings, so
`process()` is allocation-free.
- **bake() captures topology + gain values, not hot runtime state.** The baked
Processor builds fresh feedback/delay/scratch in `prepare()` and starts from zero;
a source graph that has already processed blocks does not transfer its feedback
history. The parity proof covers both directions — baked output is bit-exact to the
live graph's legacy WALK and to its routed executor (the test asserts the walk case
explicitly by forcing routing OFF, since canonical-executor routing is now ON by
default).
- **Signed `.pulpbake` files carry authored Custom state, never sampled live
state.** `bake_to_plan()` copies each node's staged
`GraphNode::custom_state_blob` (set through `set_custom_node_state()`); it must
never call a live instance's `save_state()` because bake can run while DSP is
processing and that callback has no concurrent-process contract. Temporal DSP
history is intentionally excluded. Stage the intended publish state before
`prepare()`.
- **Disk load verifies first and restores fail-closed.** `load_baked()` verifies the
Ed25519 signature before bounded parsing, rejects duplicate plan node ids and
duplicate host registry identities, then resolves every Custom record by exact
type/version/shape. A stateful record requires a valid `create` +
`load_state` lifecycle (an empty byte span may still be meaningful); null instance
creation or rejected state names the exact offending node and aborts the whole
load. The authenticated blob is restored after the Custom lifecycle's
`prepare`/`reset` on every baked `prepare()`, so those hooks cannot silently erase
authored state. The in-memory `bake()` path remains a fresh-state stream and does
not perform this disk restore.
- **In-place hosts alias input over output — never let the executor's output-zero
destroy the input.** `process_routed()` zeroes the main output bus BEFORE its
`AudioInput` gather reads the input bus (AudioOutput nodes accumulate, so N sinks
mix). Logic-style hosts (AUv2, some AUv3) hand `process()` input and output views
over the SAME memory, so that zero used to wipe the input → total silence.
`BakedGraphProcessor::process()` now detects any input-channel/output-channel
overlap and reads the input from a scratch copy sized in `prepare()` (audio thread
does pointer compares + `copy_n` only — no allocation). Any OTHER direct caller of
`process_routed()` that bridges host buffers must apply the same guard; the
executor itself deliberately keeps the zero-then-gather order.
- **Stateful Custom instances need their lifecycle re-run at baked `prepare()`.**
The captured `CustomNodeProcessFn` is an opaque closure over the instance; the
instance's `prepare`/`reset` hooks are NOT inside it. `bake()` therefore also
captures per-node `CustomNodeLifecycle` closures (type `prepare` + `reset` bound to
the instance shared_ptr) and `BakedGraphProcessor::prepare()` runs them — re-prepare
at the HOST's real rate/block (load_baked only prepares at a nominal 48k/512), then
reset so stale DSP state (a delay line's contents) never survives a re-prepare.
Note the instance is SHARED with the source graph — the baked path does not clone
it — so a baked re-prepare also resets that node in a still-live source graph.
### Bake-layer parameter injection (control-thread writes into a baked node)
A baked custom node's parameters can be changed at runtime — sample-accurately,
RT-safely, without re-baking — via the bake-layer injection primitive:
- **Declaring:** a `CustomNodeType` opts in by filling `baked_params` (id + range +
default per param) and providing `process_instance_baked_param`, a param-aware
process callback that reads values through a `BakedParamView`
(`value_at(id, sample_offset)` — offsets must be non-decreasing within a block).
That baked-param DSP runs ONLY in the baked Processor, never on the live graph.
- **Injecting:** `BakedGraphProcessor::claim_param_injection(node)` hands back a
move-only `ParamInjector` — an EXCLUSIVE per-node claim (a second claim fails until
the first is released; the handle survives re-prepare). `inject()` publishes into a
single-writer per-node mailbox the baked `process()` drains next block; events are
`pulp::state::ParameterEvent`s (immediate or ramped), and a ramp longer than one
block carries across blocks to completion.
- **One-param-per-block-or-batch contract:** `inject(ParameterEventQueue)` is the
batch path — the whole queue lands as ONE sample-accurate batch, and the latest
published queue REPLACES a still-pending one. `inject(ParameterEvent)` (single)
ACCUMULATES: it merges into the still-unconsumed pending batch, superseding only
that param's pending entries, so N single injects to different params between
blocks all land. (Pre-fix this was latest-snapshot-wins — two single injects with
no intervening `process()` collapsed to the last one, silently dropping a param.
If you need many events for one param in one block, use the queue path; single
`inject` returns `PartialOverflow` only when the pending batch is already full of
other params' events.)
`test/test_baked_graph_param_injection.cpp` is the executable spec (claims, ramps,
sample accuracy, RT-allocation-free drain, the accumulate regression).
### Forge DSP catalog families
The Forge-facing baked-node adapters are grouped by behavior, not kept in one
registry header. Use the family header that owns the DSP you are exposing:
- `forge_saturator_catalog.hpp`, `forge_distortion_catalog.hpp`, and
`forge_tape_catalog.hpp` for nonlinear/color processors.
- `forge_dynamics_catalog.hpp` for feed-forward, VCA, FET, and diode-bridge
compressors.
- `forge_multiband_catalog.hpp` for the two-band Linkwitz-Riley plus compressor
composition, and `forge_sidechain_catalog.hpp` for the two-input compressor
whose port 0 is signal and port 1 is the external detector/key.
- `forge_wavetable_catalog.hpp` for the zero-input, one-output fixed-bank
wavetable source. Source-node consumers must preserve that 0 -> 1 shape; do
not invent a dummy audio input in an application registry.
- `forge_effect_modulation_catalog.hpp`, `forge_pitch_catalog.hpp`,
`forge_space_catalog.hpp`, `forge_synthesis_catalog.hpp`, and
`forge_sequencing_catalog.hpp` for the remaining Round-2 families.
The space catalog's GPU convolution route is an opt-in host realization, not a
replacement for the CPU node. Enable it with
`PULP_HOST_ENABLE_GPU_CONVOLUTION=ON` only in a GPU/provider-configured SDK
build; the exported family keeps `space.convolution_reverb` as the default and
adds the `gpu` realization with type ID `space.convolution_reverb_gpu`.
Consumers must preserve the seven convolution controls (IR gain, pre-delay,
wet/dry, width, low/high cut) and the route's fixed three-quantum PDC: a host
capacity is rounded up to the next power-of-two transport quantum (for example,
192 becomes 256) before reporting `3 * quantum` latency. The GPU node is
intentionally non-lowerable and accepts
only one- or two-channel dual-mono IRs; four-channel true-stereo assets remain
on the CPU realization until a channel-matrix GPU route exists. The exact
provider probe is `pulp-gpu-convolution-reverb-probe`, and a Forge consumer
acceptance must use an installed SDK whose catalog export contains both
realizations and whose source/build provenance is bound to the same Pulp SHA.
The sidechain HPF setter resets its biquad when the cutoff changes. Cache the
last applied cutoff in any baked adapter and call the setter only on an actual
change; calling it unconditionally per block manufactures a fresh detector
transient and prevents the HPF state from ever settling. Keep a regression that
runs enough unchanged blocks for a DC key to disappear through the HPF.
Each header owns its stable type/parameter IDs, declared ranges, and `make_*_node`
factory. Consumers should use those constants and factories rather than duplicate
IDs, limits, or construction policy in an application registry. A catalog range
must agree with the wrapped DSP's canonical accepted range; otherwise automation
develops a dead travel region where the host moves but the DSP silently clamps.
Keep construction choices separate from baked parameters. Circuit lineage, delay
character/tier, channel topology, latency-changing modes, and anything that
redesigns storage or filters belong in a distinct type ID or factory argument.
Only controls that are safe to apply at any sample belong in `baked_params`.
When adding or changing a family, extend its `test/test_forge_*_catalog.cpp` suite
with ID/default/range parity, finite-output, determinism, gain-bound, parameter
reachability, and RT-allocation coverage.
#### Not every knob belongs in `baked_params`
`baked_params` is for values a node can accept at ANY sample. A value that
changes the node's TOPOLOGY — which stages exist, how a buffer is laid out, or
anything whose setter designs a filter — must be a registration/construction
choice with its own `type_id` instead, following the per-mode `"svf"` pattern.
Two reasons, both learned the hard way:
- **The artifact's identity.** A baked build authored as one thing must stay
that thing for the artifact's life; a control-thread write should not be able
to turn a tape delay into a BBD mid-render.
- **The setter runs on the AUDIO thread.** Anything reachable from
`process_instance_baked_param` inherits the RT contract transitively. In
`forge_character_delay_catalog.hpp` the tape TIER and tape SPEED are
construction config precisely because changing the speed redesigns a bank of
FIRs; the age macro next to them is a baked param only because its filter
banks are pre-designed at `prepare()` and the audio thread merely interpolates
between two of them.
A related ordering trap: a node's `prepare()` typically configures construction
options and only THEN calls the DSP's `set_sample_rate`, so every config setter
runs while the instance is still unsized. Those setters must tolerate being
called before allocation — store the value and let `prepare()` design against it
— or they walk buffers that do not exist yet. That is an out-of-bounds write
that only fires for nodes constructed away from their defaults, so it survives
casual testing; `test_character_delay.cpp` has an explicit regression case for
it ("configuring the tape speed before the sample rate is safe").
#### Per-sample params on a block-oriented DSP
When the wrapped block processes buffers rather than single samples, the
faithful wrapping is to call it one sample at a time and apply every param each
sample. That is only cheap if the DSP's setters are stores rather than work —
smoothing and coefficient recomputation belong INSIDE the block, on its own
control-rate cadence. A wrapper that instead calls setters which recompute
filters is doing a filter design per sample per param. Check what a setter costs
before you put it in the per-sample loop.
#### Catalog nodes with an internal control cadence read params per CHUNK
`pulp/host/forge_fdn_reverb_catalog.hpp` (the multirate FDN reverb) is the
pattern to copy for any wrapped engine that runs its own control rate. Its
`process_instance_baked_param` walks the block in 32-sample chunks and re-reads
the `BakedParamView` at each one, rather than sampling once per block.
Two things follow from that, and both bite if you copy only half of it:
- **Read at the engine's cadence, not the block's.** Reading once per block
makes a knob sweep step audibly at large buffer sizes; reading per sample
would be discarded by an engine that re-derives on a 32-sample tick anyway.
Match the engine.
- **A param that reconfigures the engine must land on a CHUNK BOUNDARY.** The
reverb's `tank_rate` re-derives every delay length, filter coefficient and
resampler ratio. Applying that between the two halves of a resampler — after
the input leg produced its samples, before the output leg consumed them —
desynchronizes them permanently: a switch UP in rate silenced the wet output
for good, and it never recovered, because the deficit was re-created every
block. The engine now applies a pending rate change before either leg runs.
If you wrap something with a similar "reconfigure everything" param, apply it
at a boundary and add a test that switches in BOTH directions and compares
against a cold render — a one-directional test passes over this bug.
Node shape matters to the host too: this node is **true stereo** (2 in / 2 out
as one logical wire, not two mono halves) and **wet only**, so a graph that
wants dry needs a `make_drywet_node` after it. `test/test_fdn_reverb_catalog.cpp`
covers the injection path, the true-stereo claim and the RT probe across a live
rate change.
## Common tripwires
- **Not every timeline automation lane addresses a device — skip, do not
refuse.** `AutomationTarget` also names a track's own mixer controls, so a
track can carry lanes that reference no `DevicePlacement` at all. The route
admission scan in `timeline_automation_delivery.cpp` walks
`TrackAutomationProgram::programs()` and must consult
`AutomationProgram::device_target()` first: a null one is **skipped**, not
reported as `MissingDevicePlacement`. Refusing it would fail admission for the
whole track — every device lane on it stops being delivered — because one lane
was never meant to reach a plugin. Those lanes are applied where the track's
audio is accumulated instead. The same rule holds in
`TrackAutomationRenderer`, which never builds a device batch for them.
- **Instruments have no input bus — never address input element 0 blind.**
An AU instrument (`aumu`) and a generator (`augn`) expose **zero input
elements**; so does a MIDI processor (`aumi`, which despite the name is
`kAudioUnitType_MIDIProcessor`, not an instrument — and it may expose no
*output* element either). A MIDI effect (`aumf`) *does* have audio input and
is unaffected. Setting a per-element input property on an input-less AU —
`kAudioUnitProperty_StreamFormat`, `kAudioUnitProperty_SetRenderCallback` —
returns `kAudioUnitErr_InvalidElement` (**-10877**), *not* a format error.
Treating that as fatal rejects **every instrument on the system** while every
effect keeps working, so the failure is invisible to effect-only tests. Ask
`kAudioUnitProperty_ElementCount` on the scope first and skip the input-side
setup when it is 0 (`scope_element_count()` in
`core/host/src/plugin_slot_au.mm`). When the AU does not answer the query,
assume an input bus **exists**: a non-answering effect then behaves as it
always did, and a non-answering instrument fails loudly at the property set
with a named scope+status. Assuming *none* would skip input setup on that
effect and leave every `AudioUnitRender` failing while the caller's buffer
keeps stale contents — **silent wrong audio, the worst of the four outcomes.**
The general rule for any backend: derive the bus layout from what the plug-in
**reports**, never from the assumption that an input side exists. The other
slots already do this and are the pattern to copy — VST3 loops
`component_->getBusCount(kAudio, kInput)` and ignores `setBusArrangements`'
status ("missing buses degrade gracefully"); CLAP and VST3 both size
`ProcessData` from the caller's view. AU was the outlier.
- **The LV2 slot cannot safely host an instrument — and it fails as UB, not
cleanly.** Port discovery keeps only `lv2:AudioPort`/`lv2:ControlPort` stanzas
(`plugin_slot_lv2.cpp`, `if (!is_audio && !is_control) continue;`), so atom /
event / CV ports are never seen and never `connect_port`'d — yet `run()` is
called anyway. The LV2 spec requires **every** port be connected before
`run()` unless it is `lv2:connectionOptional`; running with unconnected ports
is undefined behavior and commonly segfaults. Every LV2 instrument has an atom
MIDI input port, and many effects carry atom ports for transport. There is
also no MIDI delivery path for LV2 at all. Unlike the AU trap above this does
not fail loudly — so **do not claim "Pulp hosts instruments" unqualified**:
AU yes, VST3/CLAP plausibly, LV2 no. Fixing it means an atom-sequence input
buffer + MIDI mapping; the minimum stopgap is to detect non-audio/non-control
input ports at discovery and refuse `prepare()` loudly unless
`connectionOptional`.
- **Test hosts against a real *instrument*, not just a real effect.** The
effect-only integration test in `test/test_plugin_slot_au.mm` passed happily
through the bug above. Apple's bundled `DLSMusicDevice`
(`kAudioUnitType_MusicDevice` + `kAudioUnitManufacturer_Apple`) ships on every
Mac, so an instrument fixture costs nothing —
`first_apple_instrument_unique_id()` is there for this. Any host change that
touches bus/format negotiation needs both shapes, or half the plug-in universe
goes untested. (These tests WARN-and-return when no system AU is registered;
a headless VM may not surface Apple's AUs, so treat a green run in CI as
"not disproven" rather than "covered" — a **skip is never a pass**.)
- Building `pulp-host` without adding a new `.cpp` to `target_sources` —
the file sits on disk but isn't compiled; link errors fire only in the
dispatcher's `case`. **Always** update `core/host/CMakeLists.txt`
alongside adding a backend.
- Missing `PULP_HOST_HAS_<FMT>` define — dispatcher silently returns
`nullptr`. Verify `grep PULP_HOST_HAS_ build/CMakeCache.txt` after
configure.
- CLAP bundles on macOS: don't `dlopen` the `.clap` directory; resolve to
the executable inside `Contents/MacOS/` first.
- LV2 manifest URI extraction must only use subject-position `<URI>` tokens.
A manifest stanza like `<plugin> rdfs:seeAlso <plugin.ttl> ; a lv2:Plugin`
should identify `<plugin>`, not the `seeAlso` object. Keep parser coverage
in `test/test_plugin_info_metadata.cpp` or `test/test_lv2_host_discovery.cpp`
when changing `core/host/src/scanner.cpp`.
- LV2 invalid-bundle tests deliberately use placeholder `.so` / `.dylib` files.
Keep the loader's magic-byte preflight before `dlopen` / `LoadLibrary` so
invalid modules fail quickly and consistently on Windows instead of waiting
on the platform loader.
- A slot's per-block scratch must be reserved in `prepare()`, not grown in
`process()`. The CLAP slot fills `in_ptrs_`/`out_ptrs_` and emplaces into
`in_event_storage_` each block; on a default-constructed vector the first
`resize`/`emplace_back` allocates on the audio thread. Reserve the channel
vectors from `PluginInfo::num_inputs/num_outputs` (the graph sizes node
buffers from these, floored at stereo) and the event scratch for
`params_.size() + ParameterEventQueue::kCapacity +` the realtime MIDI cap.
Guard with a `PULP_DBG_ASSERT(capacity >= needed)` tripwire (debug-only).
This holds for the graph-driven path; a direct caller passing more channels
or an un-capacity-limited `MidiBuffer` is outside the contract. No-alloc
coverage lives in `test_host.cpp` ("ClapSlot::process is allocation-free
after prepare() reserves"), gated on `PULP_TEST_CLAP_PATH`.
- The AU slot (`plugin_slot_au.mm`) has the same rule for its output
`AudioBufferList`: `AuSlot::process` builds an ABL pointing at the caller's
channels every block. Size the backing `abl_storage_` once in `prepare()`
(`num_channels_`) via `au_internal::reserve_audio_buffer_list` and only
*refill* it per block (`fill_output_audio_buffer_list`) — never allocate a
fresh `std::vector` in `process()`. The ABL build lives in
`plugin_slot_au_internal.hpp` so its no-alloc invariant is unit-tested
(pointer-stable across thousands of refills) without a live AU;
`test_plugin_slot_au.mm` additionally drives a real system Apple effect AU
through `process()` (skips honestly when none is registered — headless CI
may surface no AUs). Do NOT assert `allocs==0` over `AudioUnitRender` itself
(Apple allocates internally); assert the reuse invariant on our buffer.
- The VST3 slot has the same channel-vector issue *plus* extra per-block
allocation inside the Steinberg helper containers it builds each block
(`Vst::ParameterChanges` / `EventList` from `public.sdk/.../hosting`), so
reserving `in_ptrs_`/`out_ptrs_` alone does NOT make `Vst3Slot::process`
allocation-free — a `PULP_TEST_VST3_PATH` no-alloc test against
`PulpGain.vst3` still trips. Making the VST3 slot RT-safe needs those SDK
containers pre-sized too; tracked as a follow-up, not yet done.
- Fixture wiring (`PULP_TEST_CLAP_PATH`, future `PULP_TEST_VST3_PATH`) lives in
the ROOT `CMakeLists.txt` block *after* `add_subdirectory(examples)` — NOT in
`test/CMakeLists.txt`, which is registered before `examples/` so it cannot
see the `PulpGain_*` targets at configure time. A guard placed in `test/`
silently never runs (its define just appears stale in an incremental build).
## Audio-thread snapshot contracts
The host exposes a reader-pinned audio-thread snapshot, not direct member
reads. Anything you write that touches the audio thread (a
graph editor, an MCP bridge, a preset loader) must account for these
rules:
- **The snapshot lives in `runtime::Slot<CompiledGraph>`** (`live_slot_`), the
shared reader-pinned RCU primitive in `core/runtime/include/pulp/runtime/slot.hpp`.
It owns the atomic pointer the audio thread loads, the seq_cst reader count,
and the retire list. `SignalGraph` no longer hand-rolls any of that; the old
`live_` / `live_raw_` / `retired_snapshots_` / `active_process_readers_` /
`ProcessReadGuard` / `retire_snapshot_` / `prune_retired_snapshots_` /
`wait_for_retired_snapshots_` are gone. Don't reintroduce them.
- **A pin guarantees LIFETIME, not constness.** This is the single most
misread part of the contract. `CompiledGraph` is *not* immutable — the audio
thread writes every node's scratch buffer through the pin on every block,
`inject_midi` writes mailboxes, `drain()` consumes telemetry, `set_node_gain`
writes a gain. What is immutable is the *topology*. `Slot::ReadGuard::get()`
therefore hands back a mutable `T*`; a genuinely read-only publication says so
in the type (`Slot<const T>`).
- **Pin the exact committed generation when publications are coupled.**
`ExecutionSnapshot` is a strong handle to one specific compiled graph, and its
MIDI, parameter-event, and `process()` methods never redirect to a newer live
graph. Generic `inject_parameter_events` writes a node's live mailbox; timeline
device automation instead uses `inject_exact_parameter_events` (passkey-gated)
into a separate owner-claimed exact-generation mailbox, so a claimed node's
timeline stream and ordinary live injection never share one mailbox. A ramp
event is delivered at its start offset with its ramp duration preserved, so a
hosted adapter that consumes `ParameterEvent::ramp_duration_sample_frames`
glides across the block instead of stepping at the endpoint.
`TimelineGraphBinding` publishes that handle together with its immutable
playback program and bound track renderers as one `runtime::Slot` generation.
Topology and content adoption must replace that one generation; independently
latching the program store or looking up the graph's current live snapshot can
produce a mixed old/new audio block. As with ordinary graph processing, only
one audio thread may process these mutable execution snapshots at a time.
Keep implementation changes on the focused ownership seams:
`timeline_graph_binding.cpp` owns lifecycle and realtime processing,
`timeline_graph_binding_candidate.cpp` owns candidate assembly and atomic
publication, and `timeline_graph_binding_routes.cpp` owns route validation and
mixer-edge reconciliation. Do not grow the lifecycle shell with new lowering
or routing policy.
- **Read it, don't reach for it.** Anything that dereferences the snapshot off
the prepare/release thread must hold a pin for the whole dereference:
`auto pin = live_slot_.read(); if (auto* cg = pin.get()) { ... }`. That
includes control-thread readers (`inject_midi`, `extract_midi`,
`node_latency_samples`, `set_node_gain`, `pump_anticipation`) — without the
pin a concurrent `prepare()`/`release()` can retire and free the snapshot
mid-dereference.
- **The `live_*()` getters are NOT pinned.** `is_prepared()`,
`prepared_max_block_size()`, and friends read `live_slot_.live()` directly.
They are control-thread-only by contract, and `Slot` will not save a caller
who uses them concurrently with `prepare()`.
- **`unpublish()`, not `publish(nullptr)`.** The latter is ambiguous between
Slot's `shared_ptr` and `unique_ptr` overloads.
- **Mutation protocol.** Every UI-thread `SignalGraph` mutator
(`add_*`, `connect*`, `disconnect`, `remove_node`, `clear`)
invalidates the live snapshot. `process()` returns silence until
the next `prepare()` call republishes. Batch edits: mutate, then
`prepare()`, not the other way around.
- **Plugin ownership.** `GraphNode::plugin` is a `std::shared_ptr<PluginSlot>`.
The published snapshot copies the shared_ptr, so a plugin survives
past the removal of its GraphNode until the audio thread's stale
snapshot reference drops. Do not stash raw plugin pointers.
- **Release ordering.** `SignalGraph::release()` must unpublish the live
snapshot and wait for in-flight snapshot readers before calling
`PluginSlot::release()` or custom-node release callbacks. Do not move
release callbacks ahead of snapshot retirement.
- **Live control scalars.** If a control-path setter updates state inside the
already-published `CompiledGraph`, the audio-thread field must be RT-safe.
`set_node_gain()` stores into a per-runtime `std::atomic<float>`; do not
reintroduce plain mutable snapshot fields for values read by `process()`.
- **Parameter domain.** `HostParamInfo::min_value` / `max_value` /
`default_value` are the **plain** parameter domain. VST3-internal
normalization is hidden behind the loader.
- **Parameter flags.** Consumers must honor
`HostParamInfo::flags.{automatable, read_only, stepped, is_bypass}`
before writing. Automation routing refuses non-automatable edges.
- **ParameterEventQueue.** `PluginSlot::process()` takes a
`const ParameterEventQueue&`. The queue type now lives in
`pulp::state` and `pulp::host` re-exports it for compatibility, so
format and graph code can share the event ABI without depending on
`core/host`. Current loaders consume it for per-block automation where
the format supports sample offsets. Use it — not `set_parameter` — for
per-block automation.
- **External parameter-event mailbox.** Hosts publish per-block events with
`SignalGraph::inject_parameter_events()` before `process()`. The API is
additive on `SignalGraph`; do not add a `Processor` or `PluginSlot` virtual.
Sample offsets are block-relative. A `false` return reports an invalid or
unavailable node, or a source queue that already overflowed; a retained
source prefix can still be published and consumed. Destination overflow is
observed later when the audio-thread merge fills the fixed queue. The live
API rejects a node while a timeline binding owns its exact-generation writer
claim; claims are exclusive per node and expire with their binding state.
- **Node ABI surface.** `PluginSlot` includes
`pulp/runtime/node_abi.hpp` and participates in the node ABI
virtual-order gate. Existing virtual methods may not be inserted,
removed, or reordered; add new virtual methods only after the current
tail and let `tools/scripts/node_abi_gate.py --mode=report` verify
the diff against the PR base.
- **Thread rules doc.** `docs/reference/host-thread-rules.md` is the
canonical reference.
## Hosted plugin editors
`EditorAttachment::create(slot, window)` embeds a hosted plugin's own GUI.
Wired for **CLAP, VST3, and AU v2 on macOS**; LV2 slots still report no editor.
The non-obvious parts:
- **The attachment owns the slot's resize-request channel while attached.** It
installs a `set_editor_resize_request_handler` that re-bounds the child view at
the same origin, and clears it on release — otherwise a plugin-driven resize
moves only the slot's private container and the child view the host placed
keeps its old bounds (clipped or overflowing editor). The handler captures
`this` and the slot outlives the attachment, so the move constructor /
assignment must re-install it and `release()` must clear it; a stale handler
there is a use-after-free the moment the plugin asks. An app that wants to
observe or veto should wrap the attachment, not install its own handler.
- **The two formats' default answer to a resize request differs on purpose.**
With no handler installed, VST3 ACCEPTS (resizes its container and calls
`onSize`) and CLAP DENIES. VST3's `resizeView` is UI-thread-only and is how a
plug-in reports its real size from inside `attached()`, so denying it is how an
editor ends up mis-sized. CLAP's `request_resize` is `[thread-safe]` and can
arrive from a render thread, where touching a native view is illegal — denying
is the honest answer the spec allows. Do not "harmonize" these.
- **The APIs are inverted from `HostedEditor`.** CLAP `set_parent` and VST3
`IPlugView::attached` CONSUME a parent view — the plugin inserts its own view
into what you hand it and never hands one back. `HostedEditor` is the other
way round (the slot returns a handle the host embeds). The slots reconcile
this with a host-owned container `NSView` (`hosted_editor_container.hpp`)
reported as `native_handle`. AU v2's CocoaUI does return a view, but wraps it
in the same container so teardown is uniform. Do not "simplify" this away.
- **The container is inserted into the parent window at CREATION**, not at
attach. `EditorAttachment::create` calls `create_hosted_editor()` BEFORE
`attach_native_child_view()`, so a container that deferred insertion would
have the plugin run `set_parent`/`attached` against a windowless view —
editors that bring up Metal/OpenGL layers misbehave there. The later attach is
an `addSubview:` move within the same window.
- **`core/host` is compiled WITHOUT ARC.** Container views are retained and
released by hand. Check `flags.make` before assuming ARC in any `core/host`
`.mm`.
- **`clap_host_gui` must be served from `host_get_extension`.** It returned
nullptr for every extension before editors existed, so a plugin could not call
back at all. `request_show`/`request_hide` are denied — Pulp only creates
embedded editors. Deny rather than pretend.
- **The thread rules are ASYMMETRIC, and getting this backwards is the easy
mistake.** Every `clap_plugin_gui` call (`create`, `hide`, `destroy`,
`can_resize`, `set_size`…) is `[main-thread]`. But the `clap_host_gui`
callbacks a plugin invokes are `[thread-safe]` — `request_resize` and
`resize_hints_changed` are `[thread-safe & !floating]`, and
`request_show`/`request_hide`/`closed` are `[thread-safe]`. A plugin may call
them from a render thread. So a host callback must NEVER walk straight into a
`clap_plugin_gui` call or a native view — that runs a main-thread-only API off
the editor's thread. The slot records the thread that opened the editor and
refuses (`request_resize` → false, which the spec allows) or defers (`closed`
→ pay the destroy at teardown) off-thread. Any `bool` a `[thread-safe]`
callback touches must be atomic: two callers can otherwise both pass a
test-then-set and double-destroy the plugin's gui.
- **Do not call `set_scale()` on cocoa/uikit.** `clap/ext/gui.h` documents them
as logical-size APIs that must not receive it; win32/x11 are physical-size and
do want it.
- **The gui and the container have independent lifetimes.** A plugin can report
its gui destroyed (`clap_host_gui::closed(was_destroyed=true)`) while the
caller still holds a `HostedEditor` pointing at the container. Freeing the
container there dangles that handle, because `EditorAttachment` detaches by
pointer on release. Acknowledge with `destroy()` only; let the caller's
teardown release the container.
- **A native child ALWAYS composites above Pulp's GPU layer** — you cannot paint
Pulp chrome over an embedded editor. This is a fixed OS behavior, not a bug to
fix, and it constrains node-editor designs.
- **Pulp's OWN plugins withdraw `clap.gui` under `CI=1` / `PULP_HEADLESS=1` /
`PULP_TEST_MODE=1`** (`format/detail/editor_environment.hpp`). A dogfood test
asserting a fixed `has_editor()` will pass locally and fail in CI; read the
same flag instead.
- **Testing without a bundle:** `make_clap_slot(info, creator)` in
`core/host/src/plugin_slot_clap_internal.hpp` builds a slot around a
caller-created `clap_plugin_t`, skipping dlopen. The creator receives the
slot's real `clap_host_t`, so a fake can call back into `clap_host_gui`. See
`test/test_clap_hosted_editor.mm`.
## Per-format depth
Each format loader has parameter / state / automation handling on top of the
audio-thread snapshot contracts:
- **CLAP**: real `clap_input_events_t` (param_value + midi events
sorted by time), `clap_output_events_t` harvests MIDI to
`midi_out`, `CLAP_EXT_STATE` save/load via vector-backed
`clap_ostream` / `clap_istream`.
- **VST3**: `IEditController` queryInterface (combined or separate
with controller initialize), full parameter enumeration with
ParameterInfo flags mapped onto HostParamInfo, plain-domain
get/set via `normalizedParamToPlain` / `plainParamToNormalized`,
state save/load via a `VectorStream` IBStream implementation.
- **AU**: `AudioUnitScheduleParameters` per block from
ParameterEventQueue — sample-accurate AUv2 automation.
- **LV2**: control-port discovery extended into the regex TTL parser
(lv2:ControlPort + name/default/min/max), per-port float scratch in
`control_values_`, `connect_port` wired at process() block start,
param_events apply last-write-wins.
- LV2 bundle discovery has a private test seam in
`core/host/src/lv2_discovery.hpp`; keep TTL port/binary parsing tests in
`test/test_lv2_host_discovery.cpp` rather than reaching through real
plug-in binaries for deterministic coverage.
Param domain: **plain values** at the PluginSlot boundary (not
normalized). Loaders convert internally if they natively normalize
(VST3). Don't normalize host-side.
`connect_automation(src, port, dest, param, lo, hi, ...)` delivers
two control points per block (sample 0 + N-1) via the queue. Loaders
that interpolate sample-accurately (CLAP, VST3, AU via
ScheduleParameters) get smooth automation; LV2 control ports are
sample-at-block-start so the offset-(N-1) value wins.
MixMode::Replace is the default; second Replace edge to the same
(node, param) is rejected. MixMode::Add sums then clamps.
## Bypass is a parameter in VST3/CLAP and a unit property in AU
`ParamFlags::is_bypass` can only describe a bypass that IS a parameter. VST3
(`kIsBypass`) and CLAP (`CLAP_PARAM_IS_BYPASS`) both are, so the flagged
`parameters()` entry is the whole answer there. AU is not: bypass is
`kAudioUnitProperty_BypassEffect` (Global scope, `UInt32`, read/write), and
`AudioUnitParameterOptions` has no bypass bit anywhere in it — so every AU
parameter reports `is_bypass == false`, and reading the flag alone to decide
whether an AU has a bypass always answers no.
Ask `PluginSlot::bypass_surface()` instead. It returns `BypassSurface::parameter`
(look for the flagged parameter), `unit_property` (AU: drive `set_bypass()`, which
mirrors onto the property), or `none`.
**Do not recover the flag from a parameter's name and range.** It reads like the
obvious fix — a boolean parameter named "Bypass" is exactly the shape
`state::is_bypass_param` falls back to plugin-side — and it is wrong on a stock
macOS AU. Apple's AUNBandEQ publishes **eight** parameters named exactly
"Bypass", boolean over [0, 1], one per band; none of them bypasses the unit, and
the unit's real bypass is the property. A name match therefore reports eight
bypass parameters for a plugin that has none.
Two traps when probing the property:
- **`AudioUnitGetPropertyInfo` leaves its out-params untouched when it fails.**
A unit that does not implement bypass answers `kAudioUnitErr_InvalidProperty`
(-10879) and `writable` keeps whatever was in the variable — units in the wild
have been observed leaving a non-boolean `186` behind next to a 7-byte `size`.
Reading `writable` without first checking the status invents a bypass.
Require `st == noErr`, `size == sizeof(UInt32)`, and `writable` together.
- **Support splits by component type, not by vendor.** `aufx` and `aumf` get the
property from `AUEffectBase`; most `aumu` instruments do not implement it at
all (a few do). Probe the instance; do not infer from the type.
`set_bypass()` stays a host-side control with the same output guarantee for every
format — the slot passes input through while bypassed — and additionally mirrors
onto the AU property so a hosted plugin is not left believing it is active while
the host wires around it. The pass-through is still performed host-side, so the
guarantee never depends on a plugin honoring the property it accepted.
## Review-found host graph invariants
Keep these host graph invariants covered by tests:
- `connect_automation` rejects cycles via `would_create_cycle` (automation
edges contribute to topo order so back-edges are invalid).
- `Vst3Slot` dtor only calls `terminate()` once on combined
IComponent + IEditController objects (FUnknown-pointer equality check).
- `SignalGraph::process()` returns immediately on `num_samples <= 0`
rather than memset'ing with a wrapped size_t.
- MidiInput nodes' `midi_out` is drained at the END of `process()`, not
the start. Hosts call `inject_midi()` before each `process()` to refill.
## PluginManagerPanel
`pulp::view::PluginManagerPanel` sits on top of the scanner backend and
gives host apps a ready-made "manage plugins" UI. The widget is
header-only (`core/view/include/pulp/view/plugin_manager_panel.hpp`)
and drives everything through `PluginManagerModel`:
- Tests use `InMemoryPluginManagerModel` — pre-populate `scanned_rows`,
`failed_rows`, and `paths_by_format`, then assert on `visible_count`,
`rows`, and context-menu activations. The model exposes
`rescan_count`, `single_rescan_count`, `last_reveal_path` counters
for verifying the widget wired through.
- Real hosts subclass `PluginManagerModel` and back `start_rescan()`
with either `PluginScanner::scan()` on a worker thread or the
out-of-process `pulp-scan-worker` binary. `examples/plugin-host-demo
--manage` shows the threaded-scanner pattern end-to-end.
- Blacklist persistence goes through `pulp::host::ScanBlacklist::save_to
/load_from`; the widget itself is stateless beyond the filter string.
`set_blacklisted(path, true)` must save to disk so the row stays
blacklisted across sessions.
- The widget does not render a native popup for right-click; it exposes
`context_menu_path()`, `context_menu_items()`, `context_menu_label()`,
and `activate_context_item()` so hosts can wire their own popup
(or tests can drive menu activation directly).
When adding new context-menu items or bucket semantics, remember to
extend `test_plugin_manager_panel.cpp` in the same commit — the
`[issue-494]` Catch2 tag on those cases is the canary for regression.
## `.pulpgraph` save/load
`pulp::host::GraphSerializer::to_json(graph, layout)` /
`from_json(graph, json)` round-trips topology + per-node plugin state
+ editor layout. Plugin entries store identity (format, unique_id,
manufacturer, name, version, last_path) plus a base64 state blob from
`PluginSlot::save_state()`. **Plugin binaries are never embedded.**
Two-pass deserialize: instantiate every node (mapping old → new
NodeId), then walk connections and replay `connect / connect_midi /
connect_feedback / connect_automation`. Plugin re-resolution is
scanner-identity-first; missing plugins surface in
`LoadResult::missing_plugins` and the corresponding nodes are still
created with null slots so connection ids stay stable. `GraphNode`
gained a `plugin_info` member that survives a failed slot load so
re-saving an unresolved-plugin node preserves its identity.
## Crash-isolated scanning — and the un-isolated LOAD
Isolation covers **discovery**. `PluginSlot::load()` is in-process, and a
plug-in can fault inside its own factory or `initialize()` — before `load()`
returns and long before any audio is processed — taking the host with it.
Observed 2026-07-21: a shipping commercial VST3 (Roland Cloud TB-303) segfaults
four frames deep inside its own binary during `load_vst3_plugin` on a machine
without its licensing prerequisites, with `PluginSlot::load` the only Pulp frame
on the stack. Reproduces regardless of Pulp version, so do not go looking for a
host-side bug when a `.ips` shows only vendor frames. Anything that loads
plug-ins it did not choose should load them in a child process too.
`pulp::host::IsolatedPluginScanner` (in
`core/host/include/pulp/host/isolated_scanner.hpp`) is the high-level
wrapper around the long-standing `pulp-scan-worker` binary. Construct
it with a path to the worker, then call `scan(bundle_path, timeout_ms)`
to scan a single bundle in a child process via
`pulp::platform::ChildProcess::run()`. A crash, hang, or malformed
descriptor is reported as a `ScanResult { status, descriptor,
exit_code, error_message }` instead of taking down the host.
`ScanStatus` classifications:
| Status | Trigger | Caller action |
|-----------------|--------------------------------------------------|------------------------------|
| `Ok` | exit 0 + parseable JSON descriptor on stdout | use `result.descriptor` |
| `Crash` | exit ≠ 0,2,3 OR exit 0 with unparseable stdout | `ScanBlacklist::blacklist()` |
| `Timeout` | worker exceeded `timeout_ms` | blacklist as soft crash |
| `FormatError` | worker exit 3 (unsupported bundle extension) | skip (not a plugin format) |
| `NotPlugin` | worker exit 0 with empty stdout | skip |
| `WorkerMissing` | configured `worker_path` doesn't exist | operational error — surface |
Gotchas:
- The worker's exit-code surface is frozen: 0 = success, 2 = usage
error, 3 = unsupported extension. Anything else is treated as a
crash. If you grow the worker's exit semantics, update the parent's
classifier in `core/host/src/isolated_scanner.cpp` AND the test
matrix in `test/test_isolated_scanner.cpp` in the same commit.
- The descriptor parser is a flat string-search, not a full JSON
parser — it relies on the worker's `write_json_descriptor()` schema
being stable. If you add nested objects to the worker output, swap
to `choc::json` here instead of extending the string search.
- `ChildProcess::exec_code` is `-1` for any signal-kill on POSIX (the
Crash branch covers this) and an OS exception code on Windows
(also Crash). Don't try to disambiguate further; the only signal
the parent has is "did the worker exit 0 with valid JSON or not".
- Tests use a small `fixtures/isolated_scanner_crash_helper.cpp`
binary that segfaults / hangs / emits garbage on demand. Pattern is
reusable for any future ChildProcess-based isolation work — give
the helper a mode argv and exec it from the parent.
## PluginManagerPanel drag-add
Users can drag a row out of `PluginManagerPanel` and drop it onto a
graph editor surface to add a plugin node. The panel itself stays
surface-agnostic — it emits `on_row_drag_start` / `on_row_drag_end`
callbacks with the row payload, and hosts wire the drop into whichever
graph the cursor landed on.
Panel contract:
- `PluginManagerRow` gained the identity fields needed to round-trip
into a `PluginInfo` (`manufacturer`, `version`, `unique_id`,
`num_inputs`, `num_outputs`, `is_instrument`, `is_effect`) plus a
`to_plugin_info()` helper. All defaulted so older models keep
working.
- The panel tracks press → drag-threshold → drag start, emits drag
callbacks only for `scanned` bucket rows (failed and blacklisted
rows are suppressed because they cannot load anyway), and exposes
`simulate_row_drag()` so tests and hosts can drive the callback
without synthesising motion events.
Host-side: `pulp::host::add_plugin_node_from_drop(graph, info,
&loaded)` attempts a live `PluginSlot::load` via `add_plugin_node`
and falls back to `add_unresolved_plugin_node` when the bundle can't
load — preserving user intent across save/reload. Don't bypass this
helper and call `add_plugin_node` directly; the unresolved-fallback
path is what keeps `.pulpgraph` round-trips honest when a plugin
binary disappears between sessions.
## Host-backed view link boundary
`pulp::view::widgets::GraphEditorView` is a desktop host-backed widget,
not part of the view-only link surface. Any target that instantiates it
must link both dependencies directly:
```cmake
target_link_libraries(my_graph_editor PRIVATE pulp::view pulp::host)
```
The header sets `PULP_VIEW_HAS_GRAPH_EDITOR_VIEW` to zero when host
headers are unavailable. `pulp::view::PluginManagerPanel` has the same
link requirement because its model uses the host scanner, and exposes
`PULP_VIEW_HAS_PLUGIN_MANAGER_PANEL` for conditional consumers. Keep
uses behind the matching macro; do not restore a transitive
`pulp::view -> pulp::host` dependency to make an unguarded consumer
compile.
## NativeHandleVisitor — typed access to a plugin's native handle
Hosts that need to reach a format-specific handle (CLAP note ports,
VST3 IMidiMapping, AU AudioUnit property, LV2 instance) subclass
`NativeHandleVisitor` (`<pulp/host/native_handle_visitor.hpp>`),
override the `visit_*` methods they care about, and call
`slot.accept(visitor)`. Default `visit_*` fall through to
`visit_unknown` so a visitor that only cares about one format ignores
the rest automatically. The base `PluginSlot` dispatches to
`visit_unknown` for placeholder / unresolved slots so they degrade
gracefully.
Format-specific `*NativeHandle` structs deliberately expose handles as
`void*` so they don't pull SDK headers into client code. Callers that
*do* link the SDK `static_cast` back to the concrete type. Wired in
`ClapSlot`, `Vst3Slot`, `AuSlot`, `Lv2Slot`. Tests in
`pulp-test-native-handle-visitor` pin the visit dispatch + the
`NativeHandleFormat` enumerator values (the latter is a
reorder-detector).
The pre-rename spellings (`ExtensionsVisitor`, `*Extension`,
`ExtensionFormat`, and the `<pulp/host/extensions_visitor.hpp>` include
path) still resolve via deprecated aliases, so existing host code keeps
compiling — but new code should use the `NativeHandle*` names.
## Scanner identity rules
`PluginScanner` produces `PluginInfo::unique_id` values that
`graph_serializer.cpp` keys against on rehydration. The identity
contract is:
- **VST3**: the first audio-effect class's CID from
`Contents/Resources/moduleinfo.json`, normalized to a 32-char
lowercase hex string. Read via `scanner_vst3.cpp` — **no dlopen at
scan time**. This is deliberate: opening random VST3 bundles during
a bulk scan used to crash on plugins with duplicate ObjC classes and on
plugins whose `bundleEntry()` requires a real `CFBundleRef`. moduleinfo.json
is Steinberg's declarative discovery format (VST 3.7+) and lets us read
identity without running any plugin code.
- **LV2**: the plugin URI from `manifest.ttl` (the same URI
`plugin_slot_lv2.cpp` uses at load time to pick a descriptor).
Parsed via a tiny regex — we deliberately don't pull in
lilv/serd/sord for a single-field read.
- **CLAP**: `desc->id` from `clap_plugin_descriptor_t`, extracted by
briefly loading the bundle in `scanner_clap.cpp`. CLAP bundles are
the only format where scan-time dlopen is safe — the CLAP ABI is
designed for cheap metadata reads and the bundles don't ship ObjC.
- **AU**: scanned by `AudioComponent` API, not file-system walk. The
AU component's type/subtype/manu four-char codes serve as identity.
### An AU identity needs no scan at all
AU is the one format whose identity is also its complete loader
descriptor — there is no filesystem path to resolve. So when you already
know which component you want, **do not scan for it**:
```cpp
PluginInfo info;
if (pulp::host::plugin_info_from_au_identity("aumu:Vas7:RoCl", info))
auto slot = PluginSlot::load(info); // straight to AudioComponentFindNext
```
Scanning to reach a known AU walks every installed plugin — slow, and it
drags the caller through unrelated third-party bundles that may
instantiate badly. `plugin_info_from_au_identity()` (in `scanner.hpp`) is
pure string logic with no platform dependency, so it is testable and
tested everywhere, not just where AUs exist.
It accepts only the four component types Pulp can host, and the type
code decides the bus shape: `aumu` is a MIDI instrument (0 in, MIDI in),
`augn` is a no-input generator (0 in, **no** MIDI), and `aufx`/`aumf`
are effects (stereo in). Getting `augn` wrong is the easy mistake — it
is a source like an instrument but is not MIDI-driven.
It returns false rather than throwing so callers can fall through to a
scan for other formats. That fallback is why rejection is strict: a
parser loose enough to accept a CLAP-style id would silently skip
discovery and then report "no suitable plugin found".
Bundles that don't expose their identity through the safe path (e.g.
VST3 without moduleinfo.json) fall back to the directory stem. The
graph_serializer rehydration handles stem IDs the same way it always
did — best-effort — so the scanner stays safe across a user's entire
plugin folder.
**Placeholder plugin node**: when graph_serializer can't resolve a
saved plugin at load time, it creates a Plugin node with a null
PluginSlot so topology survives. `SignalGraph::process()` treats
null-slot nodes as deterministic input→output pass-through (or
zero-fill on channel-count mismatch). Don't assume a Plugin node
always has a live slot — always null-check `plugins[id]` before
dereferencing.
## `latency_samples()` alone is not a fact — ask `latency_query()`
`PluginSlot::latency_samples()` returns an `int` whether or not the backend can
actually ask the plugin. On its own it cannot distinguish **"this plugin reports
zero latency"** from **"this backend has no way to ask"** — and collapsing those
turns an unanswered question into a confident zero, which anything downstream
will happily treat as verified.
`PluginSlot::latency_query()` is the fact:
| | |
|---|---|
| `Available` | The value is the plugin's own report. Zero means zero. |
| `Unsupported` | This backend cannot ask. The value is meaningless. |
| `QueryFailed` | The backend asked and the plugin errored. |
It defaults to `Available` (CLAP, VST3, and AU all read a real value). **The LV2
slot overrides it to `Unsupported`**: it does not read the plugin's
`lv2:reportsLatency` control port, so its `latency_samples() == 0` is a
placeholder, not a claim. Wiring that port is the fix; until then, anything that
reports or gates on hosted latency must branch on `latency_query()` first.
`pulp audio render --latency-report` does exactly that, which is why an LV2 plugin
comes back `unsupported` rather than being falsely certified as zero-latency.
## A `CustomNodeType` pack is one header, and its consumer needs to be told
Bake-layer DSP packs live as header-only `CustomNodeType` factories under
`core/host/include/pulp/host/forge_*_catalog.hpp` — one header per pack
(`forge_lofi_catalog.hpp`, `forge_character_delay_catalog.hpp`,
`forge_modulation_catalog.hpp`). Each node declares its stable `k…TypeId`, its
`baked_params` (the injectable macro contract), and its port arity, and gets
registered on a `SignalGraph` before `bake()`.
Treat that consumer wiring as part of the DSP change, not a follow-up. A
musician-facing audio, MIDI, instrument, drum, or timeline capability anywhere
in Pulp must ship with one of these explicit outcomes:
1. an effect or graph utility exposed through a Pulp
`forge_*_catalog.hpp` adapter plus the matching Forge lane registry,
generation vocabulary, canonical display/range metadata, and bidirectional
capability-contract coverage;
2. an instrument/drum or MIDI-transform exposure through that lane's typed
catalog, rather than disguising it as an effect node; or
3. an explicit internal-primitive classification naming the higher-level
catalog capability that composes it. An internal refactor behind that stable
adapter does not require registry churn, but behavior changes still require
adapter and contract-test review.
Only advertise controls the capability intentionally exposes and can accept
safely at runtime. Do not turn internal members, coefficient-design switches,
allocation choices, port shape, latency, or topology into knobs merely because
they exist. Audible timeline or arrangement controls route through the
appropriate effect, instrument, or MIDI lane; structural timeline APIs and
generation grammar are not baked parameters.
The same rule runs backward for an exposed capability. Removing or renaming an
advertised DSP realization or injectable parameter requires a coordinated sweep
of the Pulp catalog header and the Forge registry, generation vocabulary,
canonical display/range metadata, and bidirectional contract tests. Preserve
published type IDs and node-local parameter IDs unless the change includes a
tested migration; never leave a registry row pointing at deleted SDK code, and
never leave stale Forge vocabulary teaching the generator an unavailable
capability. Removing an internal primitive behind an unchanged catalog adapter
does not remove that public identity.
`<pulp/host/forge_catalog_index.hpp>` is the canonical Pulp-owned catalog-pack
index and umbrella include. Every `forge_*_catalog.hpp` leaf must appear there.
`pulp-test-forge-catalog-index` scans the source include directory and compares
it with the index, so adding an unindexed pack or retaining an entry for a
removed pack fails closed. Downstream consumers should derive their Pulp catalog
header set from this index instead of maintaining another manual list.
Three things bite when adding a pack:
- **A control signal is an ordinary audio port.** The convention across every
pack is a unipolar `[0, 1]` signal on a normal port — a source declares
`num_input_ports = 0`, a consumer takes the signal on port 0 and the control
on port 1. Nothing in the graph distinguishes CV from audio, which is exactly
why modulation needs no new graph concept and bakes like anything else.
- **A downstream capability gate may enumerate the SDK by parsing one header
path.** Forge's `test_capability_contract` derives "no catalog node is
orphaned" from the catalog header the build points it at. A pack in a *new*
header is invisible to that check — the nodes exist, nothing reaches them, and
every test stays green — until the consumer's header list learns about it.
When you add a pack, add its header to that list in the same change.
- **Gain metadata belongs to the catalog that owns the DSP.** A downstream
registry must call a catalog helper, not reconstruct a bound from engine
constants. When the runtime helper depends on sample rate but the registry
stores one fixed number, expose a second catalog helper for the all-rate
ceiling (for example, `vocoder_all_sample_rates_worst_case_gain()`) and test
that it bounds the rate-specific helper. Otherwise DSP math leaks into the
importer/generator layer and the two formulas drift independently.
- **The index array is length-typed — bump the size, not just the entry.**
`kHeaderNames` is a `std::array<std::string_view, N>`, so adding a pack means the
`#include`, the name in the list, **and** `N`. Forgetting `N` is a compile error,
so it fails closed; it is only worth knowing so the error reads as expected rather
than mysterious.
**A CV pack works in VOLTS, and that is a deliberate exception to the `[0,1]`
convention above.** `forge_eurorack_utility_catalog.hpp` holds the modular
primitives with no DAW-plugin analogue — nobody ships a "Multiple" as a VST — so
they live in their own pack instead of being retrofitted into the musical-DSP packs
that stay shared across targets. Those nodes follow the published modular voltage
standards, not a normalized range: audio ±5 V, unipolar CV 0–10 V, bipolar CV ±5 V,
gates and triggers 10 V, and Schmitt detection with a 0.1 V low / 1.0 V high
hysteresis band. The constants live once in `pulp::host::eurorack`
(`kAudioPeak`, `kCvUnipolarMax`, `kCvBipolarPeak`, `kGateHigh`, `kSchmittLow`,
`kSchmittHigh`) specifically so nodes cannot drift apart — a baked param range is
written as `{kOffset, -kCvBipolarPeak, kCvBipolarPeak, 0.0f}`, in volts.
So there are now two control conventions in the catalog, and which one applies is a
property of the pack, not of the port: the shared DSP packs carry unipolar `[0, 1]`
on an ordinary port, and the eurorack pack carries volts. The graph still cannot
tell CV from audio, which is why neither convention needs a new graph concept — but
patching a eurorack node's output into a shared-pack modulation input without
scaling is a 10× error that will not fail to compile and will not fail a contract
test. Reuse these constants rather than writing `5.0f`, and convert explicitly at
the boundary between packs.
Reconfiguring state (a filter's cutoff range, an envelope's stage lengths) is a
per-block-on-change read of `BakedParamView::value_at(id, 0)`, not a per-sample
one: those setters recompute coefficients rather than scale a value, and a
lifecycle whose boundaries move every sample has no boundaries. Values that
genuinely scale (gain, depth, cutoff) stay per-sample so a knob sweep is
sample-accurate.
## A baked ParamID is a serialization contract, not a storage offset
`CustomNodeType::baked_params` IDs are chosen for persistence: a semantic keeps
its ID wherever it appears, and engine-specific banks get disjoint reserved
ranges so a later addition never renumbers an existing preset. That makes them
deliberately **sparse**. The drum catalog declares a few dozen controls per node
while its FM operator banks start at 200 and its wave bank reaches 287.
So never index a per-instance cache with a ParamID. Sizing an array by the
largest ID wastes most of it and — the part that actually bites — turns the next
reserved range into an out-of-bounds write on the audio thread, with no
compile-time or test signal until the range is used. Build a dense table once
when the node type is created (off the audio thread), map each declared ID to a
slot, and size instance caches by parameter *count*. Return an explicit
"no slot" for anything undeclared so a stray ID is inert instead of memory-unsafe.
Carry each declared `min`/`max`/`default` in that same table, because it is also
what lets injected automation be sanitized correctly. A generic "replace
non-finite with zero" wrapper is wrong: zero sits below the minimum of every
tuning, decay, cutoff, and component-value control, so it lands further outside
the contract than the value it replaced. Mirror `state::ParamCursor` — a
non-finite value becomes that parameter's **declared default**, then clamp.
## Declare only what the engine actually consumes
A node that registers a control its DSP ignores passes every registration test:
the parameter exists, has a range, clamps, and round-trips through state. It
still ships a dead knob — saved into presets, given a place in a UI, and worth
nothing. Two variants of one engine family are where this hides, because the
shared registration path is written once and the bodies diverge.
Registration tests cannot catch it. The check that can is a rendering one: set
each declared control to its default, render; set it elsewhere in its declared
range, render; require the two to differ. Controls that are conditional by
design (a decay time for a layer at zero level, an LFO rate with no depth) are
not dead — give each one the companion settings that make it observable and
take both renders under them, so the gate separates "needs a partner" from
"reaches nothing".
Confirm such a gate can fail before trusting it green: reintroduce one dead knob
and check it is reported. A parameter-efficacy gate that silently passes
everything looks identical to a clean contract.
## A voice's `reset()` must clear every filter it owns, not most of them
A drum voice's `reset()` is the contract the bake layer leans on: the node's
`type.reset` calls it, a kit calls it when transport stops, and a deterministic
corpus depends on it. "Clears all state" has to mean all of it.
The failure mode is quiet. `TomVoice::on_reset()` cleared its envelopes, click,
noise source, output stage and phase — and left the skin's 4-pole ladder holding
four stages plus a saturated feedback sample. Nothing sounds obviously broken:
the voice still starts, still decays, still chokes. What breaks is *repeatability* —
the same hit rendered twice from a reset voice produces different samples, so a
regenerated corpus never reproduces and a golden-file test fails on the second
run rather than the first. The sibling snare had always reset its filter, which
is why the gap survived review: the pattern looked established.
When adding or auditing a voice, list its stateful members and check each appears
in `on_reset()`. Filters and delay lines are the ones that get missed, because
they have no `is_active()` to make their leftover state visible.
The check that catches it is one assertion, and it is worth writing for every
voice: reset, render a hit, reset, render the same hit, require the two buffers
equal. Run it with the voice's *noisiest* path at full level — a path that is
silent in the test cannot carry stale filter state into the output, which is how
a partial reset passes a determinism test that looks thorough.
## A repeated designated initializer compiles on macOS and breaks Linux
`ExecutorSnapshotBinders` and the other binder structs are filled with designated
initializers, which makes them pleasant to extend and makes one mistake invisible
locally: naming the same field twice.
Clang accepts a repeated designator as an extension and silently keeps the LAST
one. GCC rejects it — `error: '.field' designator used multiple times in the same
initializer list`. So the whole macOS lane, including the required gate, stays
green while every Linux build in the repository fails. On a merge queue that
groups on all-green, that blocks PRs which never touched the host.
It arrives through merges rather than through typing: a hunk that adds a binder
gets applied to both sides of a merge and lands twice. `git diff` on the merge
looks unremarkable because each copy is individually correct.
When a binder initializer grows, or after resolving a merge that touched one,
count the designators:
```bash
grep -c '\.custom_latency_for' core/host/src/signal_graph.cpp # expect 1
```
If the duplicated bodies are identical, removing either is behaviour-preserving —
clang was already using the last. If they differ, clang has been running the
second and GCC has never compiled it at all, so decide which is correct before
deleting.
## A custom node declares its latency; it does not dodge it
`CustomNodeType::latency_samples` is a `std::function<int(double sample_rate)>`,
evaluated once off the audio thread at compile/lower time and fed into both the
legacy and routed PDC plans. It is rate-taking rather than a fixed count because
the two things that usually need it — a lookahead quoted in milliseconds, and an
oversampler's half-band group delay — resolve differently per rate.
Before it existed, a node with intrinsic delay had one option: remove the delay.
The Forge drum catalog did exactly that, forcing every voice's output stage to
`bypass` so it could keep the zero-latency contract. That workaround outlived
its reason by several releases, because the comment explaining it was never
revisited — it still claimed no latency surface existed.
Worth knowing how that ended, because "the constraint is gone, so lift the
workaround" turned out to be wrong: measured, the oversampled path was a net
loss at the shipped defaults. Every nonlinear stage is inactive by default
(drive 0, fold 0, quantiser at 24 bits), so the signal made a pure half-band
round trip and paid the FIR's linear-phase pre-ringing for a stage doing
nothing to it — about -0.1 dB above 8 kHz gained, against a material attack
smear on the voices whose character is dense transients. The node still
bypasses, but now it says so on evidence and derives its declared latency from
that one choice, so flipping it cannot leave a stale number behind.
The general lesson: removing a workaround is a behaviour change and wants the
same measurement any other behaviour change would. `tools/audio/quality-lab`
(`compare a.wav b.wav --profile transient-integrity --align latency`) is what
caught it; `--align latency` matters, because the delay you just introduced
will otherwise read as a difference all by itself.
Two things to know when wiring it up:
- **Where the value is captured depends on how the node runs.** The live compile
only records latency for a node that has a live callback
(`custom_processors.contains(id) && type->latency_samples`) — a node with no
live callback is transparent on the live graph and must not add delay there.
A **baked-only** node (`create` + `process_instance_baked_param`, no `process`)
therefore reports 0 from `SignalGraph::live_custom_latency_samples` and carries
its latency through lowering into `BakedGraphProcessor` instead. Asserting on
the wrong one of those two paths looks like the feature is broken.
- **Derive the number, never restate it.** The latency the host compensates and
the latency the DSP introduces have to be the same value. If a node names its
quality/lookahead in one place and its latency in another, they drift the
moment either default moves, and nothing fails — the drum sits a few dozen
samples off every undelayed path beside it. Name the choice once
(`OutputStage::kDefaultOversampling`) and compute the latency from it.
## Refusing a control is the last resort, not the first
A catalog node that cannot honour a control has two honest options: make the
control mean something, or do not declare it. Refusing a *declared* control is
the worst of the three, and it is the one that looks most rigorous.
The case that made this concrete: the bridged-T kick has no frequency
parameter — its pitch is a property of the network — so a `tune_hz` macro on
that body was refused with an accurate, actionable message. The message was
correct and the behaviour was wrong. A kick that cannot be tuned is not what
anyone asking for an 808 kit means, and "an 808 style kit" failed to build
because of it: the model generated the natural graph, was refused, regenerated,
and the attempt budget ran out. The refusal taught the user our implementation
detail instead of serving the request.
The control was implementable the whole time. The network retunes by
substituting capacitors, exactly as the hardware does; scaling both arms moves
the centre frequency and leaves Q invariant. The header had even said so —
"scaling both arms therefore retunes the drum while leaving Q exactly
invariant, which is the signature a tune control has to reproduce" — describing
a control nobody had built.
So when a node cannot honour a control, ask in this order:
1. **Can the underlying model express it another way?** Physical models usually
can: a frequency is some function of the components, a time is some function
of a decay coefficient. Solve the inverse. That is a real control, and the
parameter-efficacy gate will then pass it honestly.
2. **If genuinely impossible, do not declare it.** An emergent behaviour — the
circuit kick's pitch droop — is not a control, and declaring one would be a
dead knob.
3. **Only refuse when the pairing is expressible but wrong**, and then teach the
refusal wherever the capability is advertised. An untaught refusal is a retry
loop that costs a generation budget, not a guardrail.
## `ExecutorSnapshotBinders` is a designated-initializer aggregate
`build_executor_snapshot()` takes its collaborators as an
`ExecutorSnapshotBinders` aggregate, filled in `SignalGraph` with designated
initializers. Two rules bind when you add or move a binder:
* **Designators must stay in member-declaration order**, and no member may be
named twice. Clang accepts violations of both; **MSVC rejects them** with
`C7560`, so the mistake is invisible on the required macOS gate and takes out
`pulp-host` — and therefore every Windows plug-in — instead.
* A duplicated binder is what a merge conflict resolution produces here, because
the blocks are long lambdas that diff poorly. `main` carried a byte-identical
duplicate of `.custom_latency_for` for exactly that reason.
`tools/scripts/designated_initializer_lint.py` guards the duplicate case; adding
a binder out of order is still only caught by an MSVC build.
## A Forge-exposed parameter needs a descriptor, not just a baked range
`CustomNodeBakedParam` carries id, min, max and default — everything the audio
graph needs, and nothing an agent choosing the node can read. The rest of the
vocabulary (stable key, label, unit, stepped-vs-continuous, named states,
description, and the finite construction axes) lives in
`ForgeNodeDescriptor` (`pulp/host/forge_param_descriptor.hpp`), authored as an
`inline ForgeNodeDescriptor descriptor()` next to the node it describes.
`forge_fuzz_catalog.hpp` is the worked example.
Two rules that are easy to get wrong:
- **Descriptors carry no range or default.** Those are joined from the baked
node at export time so a number exists in exactly one place. Adding them back
reintroduces the drift the descriptor was built to end.
- **A parameter absent on some realizations is expressed with
`realization_modes`, not by dropping it.** Some families deliberately omit a
control that would be inert (an Ampex deck has no selectable EQ standard), and
an inert control presented as live is a worse answer than an absent one.
- **A named choice that exists only on some realizations carries its own
`realization_modes`.** Tape EQ is the worked example: Studer exposes NAB/CCIR
while cassette exposes Type I/II. A single unscoped four-choice list lies
about both baked ranges.
- **Each concrete realization explicitly selects every finite axis.** Do not
make a consumer parse an opaque mode like `silicon_x4`; its `settings` pairs
machine keys (`device=silicon`, `oversampling=x4`) with the exact `type_id`.
`audit_forge_descriptor()` joins descriptors against the node in both
directions, so a node that gains a control with no descriptor fails as loudly
as a descriptor for a control that was removed — the two ways a catalog goes
quietly stale. `tools/scripts/forge_descriptor_coverage.py` holds the other
half: its independent 81-key semantic manifest must agree with the 20-pack
index, descriptor sources, expected-node inventory, and explicit export
registrations. There is no `PENDING` escape hatch, so a newly indexed pack
cannot land undescribed.
`forge_catalog_export_nodes()` is the runtime projection used by
`pulp forge catalog export`. It constructs each registered semantic node and
copies `baked_params` beside its descriptor; JSON serialization joins by
`ParamID` and realization, so each `min`/`max`/`default` always comes from that
concrete DSP declaration rather than one representative preset.
`audit_forge_catalog_export()` checks both descriptor parity and an independent
expected-node inventory. Keep the missing-node negative control: removing a
registry entry must fail instead of emitting a shorter, superficially valid
document. SDK installs carry the checked projection at
`share/pulp/forge-catalog.json`; consumers read it from the selected SDK.
With `PULP_HOST_ENABLE_GPU_CONVOLUTION=ON`, the SDK install builds that JSON
from its own native `pulp-cli forge catalog export --json`; the default CPU
snapshot cannot describe this option. The generated file stays in the build
tree, depends on the exporter, and fails installation if generation fails or
omits the GPU realization. Default-OFF builds still install the checked snapshot.
A raw `cmake --install` does not build missing outputs; use the build's `install`
target to generate the enabled catalog before copying it.
The `cli-forge-catalog-check` test compares the complete runtime export against
that generated artifact when enabled, and against the committed snapshot in
default builds. A GPU mode's presence alone does not prove metadata equality.
The export's `add(...)` registrations live in one translation unit per catalog
family (`core/host/src/forge_catalog_export_<family>.cpp`, behind the private
`forge_catalog_export_detail.hpp`), not in `forge_catalog_export.cpp`. Every
pack's node factory is header-inline, so whichever TU constructs a realization
also instantiates that node's DSP; one file constructing all 20 packs compiled
the whole Forge DSP library into a single 14.5 s object (10.6 s of it backend),
the slowest TU in the tree. Add a new pack's registrations to the family file
whose packs it belongs with (or a new family file listed in
`core/host/CMakeLists.txt`); `forge_descriptor_coverage.py` reads the main file
plus every `forge_catalog_export_*.cpp` when it counts registrations, and the
audit suite pins the family append order, so a family appended out of order or
left out of `forge_catalog_export_nodes()` fails there rather than in a
consumer of the JSON.
Treat that JSON as a published v1 contract, not a regenerable implementation
detail:
- Preserve existing realization mode keys, axis tokens, axis numeric values,
and type IDs. Additive realizations may introduce new identities, but changing
an existing one requires an explicit schema migration.
- Axis declarations name construction dimensions; the concrete `realizations`
list is the authoritative supported finite subset, not an implied Cartesian
product. Keep it aligned with the registrations the consumer can instantiate.
- Computed realization modes and type IDs must be owned strings. A descriptor
built from factory-generated IDs cannot retain `string_view`s into temporary
nodes.
- A realization-specific baked contract must describe effective DSP behavior
after construction-time constraints. For example, a through-zero flanger
whose fixed offset clamps modulation depth exports that offset as its maximum
depth instead of advertising values the DSP collapses.
## A live-swap test must wait for the swap, not budget for it
`begin_swap_edit()` / `prepare_swap()` are control-thread calls, so a test that
proves the no-silence contract runs them on a worker while the main thread
renders. Nothing orders that worker's first iteration before a fixed render
budget runs out. On a loaded machine the whole budget can be spent before the
worker is ever scheduled, leaving every swap counter at zero — which reads as
"`prepare_swap` never published", a `SignalGraph` regression, when the swap
path was never entered at all.
Assert on the worker's progress, not on the render budget: spend the budget,
then keep rendering until a swap outcome is observed, under a deadline so a
graph that genuinely never swaps fails the check instead of hanging. The same
shape applies to any `SignalGraph` test whose assertion depends on a second
thread having reached a specific call.
## Offline rendering a timeline region: use the transport-aware renderer
`OfflineSignalGraphHost` will not render a timeline slice. It drives the
no-transport `SignalGraph::process()` overload by design, so timeline nodes
never receive the transport that note scheduling, automation and PDC read.
**It does not fail — it renders silence**, which is the expensive way to learn
this.
Use `render_timeline_offline()` (`pulp/host/timeline_offline_renderer.hpp`).
Its loop is `MasterTransport::begin_block()` →
`TimelineGraphPlaybackBinding::process()`, and it takes a freshly built or
quiesced graph because it prepares the binding and drives its own transport — a
graph already bound to a live transport would be double-driven.
Two gotchas worth knowing before you debug an output that looks *almost* right:
- **A region that does not start on a block boundary is pre-rolled from tick
origin with output discarded.** If you "optimise" that away by seeking
straight to the start, stateful nodes enter with an empty delay line and a
cold PDC pipeline: the region will still render, still be non-silent, and
still be wrong — it will merely resemble the same slice of a full bounce
instead of equalling it sample-for-sample.
- **Bounds are ticks and are derived only through the program's compiled tempo
map.** `CompiledTempoMap::ticks_to_samples()` saturates rather than trapping,
which is correct for a real-time reader but would silently turn an absurd
region into a plausible bounce here, so the renderer rejects a saturated
derivation instead of using it.
## Standing up a bounce graph: the device-free shape, and its two traps
`TimelineGraphPlaybackBinding` is expressed entirely in `NodeId`s, so a caller
that wants to render a compiled `PlaybackProgram` has to build a graph and a
route table before it can prepare anything. Use
`build_device_free_timeline_graph()`
(`pulp/host/timeline_offline_graph.hpp`) rather than assembling that by hand.
It adds one output node and returns one `TimelineTrackGraphRoute` per program
track, in program order.
The binding still creates the per-track arrangement-audio and mixer nodes
itself, so this topology is at exact parity with
`playback::ArrangementAudioRenderer` over the same program. Device chains are
additive on top of this shape, not a different one.
Two things about `TimelineTrackGraphRoute` are easy to get wrong because the
struct looks under-specified when it is actually complete:
- **All-zero post-device fields are the device-free contract, not a stub.**
With `post_device_audio_source` and `post_mixer_audio_destination` both zero
the binding inserts the track mixer directly between arrangement audio and
`audio_destination`. Populating them "for symmetry" brackets a hosted device
chain that does not exist.
- **`midi_destination` stays zero, and that is not a gap to fill.** A
device-free graph instantiates no instrument, so there is nothing for note
events to reach. Routing them at a node that does not exist renders the same
audio and hides the missing instrument.
An empty route set is a valid *build* result — a trackless arrangement renders
silence — but it must not reach `prepare()`, which rejects it as an
under-specified request (`TimelineOfflineRenderCode::InvalidProgram` documents
the same rule). Skip the graph and write the zero-filled buffer instead.
## Who owns a device chain is decided by authorship, never by its length
`resolve_timeline_device_route()` builds every node in an admitted chain and
generates the routes for all of them, so the route validation in
`timeline_graph_binding_routes.cpp` has to decide whether a chain is
resolver-owned before it demands routes from the caller. That test reads one
thing: whether every placement is `DeviceKind::BuiltIn`. It must not also bound
the chain length.
Bounding it by length looks harmless and produces a wrong *diagnosis*. An
over-long all-built-in chain then falls out of resolver ownership back into
caller-owned validation, which reports the routes the caller never supplied —
`MissingDevicePlacement` — instead of `UnsupportedDeviceChain`, the refusal the
resolver actually raises for that shape. The chain is still refused, so every
test asserting "this is rejected" stays green while the code the caller reads
to fix it is the wrong one. The length refusal lives in the resolver's own
`validate_declaration`, and it fires before the fixed-size descriptor array is
built, so letting an over-long chain reach it cannot overflow anything.
## An event-to-event device reports latency because it may move an attack later
`pulp.device.event.humanise` wraps the existing `pulp::midi::Humanize` kernel
and reports `kEventHumaniserWindowSamples`. That is not a cost of the
implementation: a device permitted to move an attack *later* must hold that
attack for the full width of the window it may move it into, and the reported
figure is exactly what the event-stream scheduler compensates. Compensation
shifts the scheduling window; it never rewrites an event's `sample`.
The admitted shape is deliberately one event-to-event device ahead of one
instrument. Every other declaration is a typed `TimelineGraphAdmissionCode` at
admission rather than a truncation, and the two bounds a caller needs to predict
one — `kAdmittedDeviceChainLength` and `event_device_latency_ceiling_samples()`
— are on the public resolver header so nobody re-declares the ceiling and drifts
from the constant the graph binding enforces.
## Hosting sample-region processors
Sample-region processors use the ordinary `Processor`/`SignalGraph` lifecycle.
Freeze and bind promoted parameters before exposure, prepare the candidate, and
publish only after the region proof accepts. The callback uses prepared storage;
live edits adopt immutable snapshots and fail-fast admission prevents retained
state from executing under two bindings. Signed baked artifacts are re-parsed,
reconstructed, and re-proved before use.
## An AU negotiates its channel width at initialize time — and a mismatch is SILENT
`PluginSlot` for AU fixes the stream format when the unit is initialized and cannot adapt to a
differently shaped buffer afterwards. If the width the host will actually render is not supplied
*before* `prepare()`, the unit initializes at the descriptor's width (falling back to 2) and every
render at any other width is rejected by `AUEffectBase` as a malformed shape: it zeroes the output,
sets `kAudioUnitRenderAction_OutputIsSilence`, and returns `noErr`.
**That failure is silent in both senses — no error status, no audio.** A mono render of a
stereo-negotiated AU produces a valid-looking WAV full of zeros and a success code. Nothing in the
logs says the audio was discarded.
So when hosting an AU at a width that is not the descriptor default, call
`slot->set_preferred_channel_layout(inputs, outputs)` before `prepare()`. It is a default no-op on
`PluginSlot`, overridden only by the AU slot; VST3, CLAP and LV2 take their width from the
per-block buffers and need nothing. The AU slot sets the input and output scopes independently, so
an asymmetric unit negotiates correctly too.
Diagnostic: the AU slot logs `AU v2: initialized with N channels`. If that N disagrees with the
width you are rendering, the render is silence regardless of what the status says — check it before
trusting any AU measurement.
## Querying the instance that actually runs
Do not use `nodes()` and an opaque pointer as a live diagnostics shortcut. A
shared owner keeps an object alive but cannot stop prepare/release from replacing
its engine. Use the graph-owned custom-node diagnostic query: it tries the same
mutation lock as lifecycle operations and retains the control-thread `Slot::live()`
owner. Never take an RCU reader pin under that mutation lock; release may wait for
those pins. The query is non-RT and its type-key lookup may allocate.
Keep the diagnostic descriptor separate from the positional `CustomNodeType`
aggregate. A live parameter wrapper receives its own opaque wrapper instance,
not the original GPU instance: its inspector must unwrap the private inner
pointer before calling the original inspector. Register against the exact alias
and preserve the original producer identity separately. Otherwise plausible
metadata can describe a different instance than the one producing audio.
A live counter snapshot is approximate. `AudioCallerStopped` is an explicit
caller promise, not a request to stop processing or a proof of worker drain.
Read selected-delivery totals after joining the sole process caller and before
release resets them; keep worker output distinct from authenticated GPU-selected
output. See [the query contract](../../../docs/guides/custom-node-diagnostics.md).