15 KiB
sound — the sound orchestrator
EngineSound governs every aspect of sound in the framework (redesign
2026-07-31, PLAN-sound-redesign.md):
the substrate (one AudioContext per document and its unlock lifecycle),
the mix (master → buses), the synthesis machinery with registered
voices, the sample path, and media citizenship (the playback transport
the service PROVIDES, audio focus, the UI↔content relation, MediaSession).
Nobody else creates an AudioContext or touches Web Audio directly, the way
every timer goes through uix.timers and every animation through
uix.motion. Consumers adapt to the service, not the other way around.
const sound = createEngineSound({ dom, timers }); // dom: a SoundDom port
sound.prime(); // SYNCHRONOUS — call it inside the user gesture
await sound.play(sig, { bus: 'ui', voice: 'sema', gainScale }); // an earcon; never rejects
await sound.preload(urls); // pre-decode samples into the URL-keyed cache
await sound.decode(bytes); // raw bytes → AudioBuffer, on THIS context
sound.master.setGain(0.4); // document-wide policy, the root of the bus tree
sound.bus('ui').setGain(0.5); // per-bus policy: gain / mute / duck holds
sound.bus('ui').setMuted(true);
const hold = sound.bus('ui').duck(0.2); // refcounted; strongest factor wins
hold.release();
sound.registerVoice('sema', voiceSpec); // calibration is the CONSUMER's data
sound.suspend() / sound.resume(); // explicit; wins over the autoSuspend policy
sound.dispose(); // idempotent
sound.state; // 'absent' | 'suspended' | 'running'
sound.context; // AudioContext | null — the raw output
The mix — master → ui + content
"Sound" means two different things in a UI framework, so the graph has two
buses under the master: ui carries the earcons (sema's channel) and
content carries the work itself (media playback, voice notes). Each bus
is a policy surface — gain, muted, refcounted duck() holds (the strongest
live factor wins; they don't multiply) — over a GainNode you can also reach
raw (bus.node) to hang an analyser or a source.
The structural invariant this exists for: UI-sound policy must never touch
the content's volume. Honest about the mechanism (audit AU-4): today
prefs.sound acts through sema's channel — BK-REDUCTIONS resolves the level
and hands it over as gainScale, always on the ui bus — and nothing wires
prefs to a bus directly. The ui bus gain/mute is the graph-side LEVER an
app may additionally bind (e.g. prefs.sound → bus('ui').setGain); what is
guaranteed by construction is the negative: no UI-side policy ever writes the
content bus or a source's own volume. And the three levers never write the
same value: the master carries document policy, the bus carries per-bus
policy, and gainScale rides one playback's own envelope.
Media citizenship — the service provides the transport
const media = sound.media(el, {
focus: 'exclusive', // per-source override of `contentFocus`
metadata: { title, artist, album, artwork } // MediaSession opt-in
});
// media = the player's MediaProvider shape: play/pause/seek/setVolume/
// setMuted/setPlaybackRate/snapshot/subscribe/destroy — plus attach().
Registering a content source is what buys the citizenship (D-SR.4…D-SR.6):
- Audio focus toward other content sources —
contentFocus(service-wide, default'mixed') or per-source:exclusivepauses the others,duckattenuates them (contentDuckFactor, default0.25≈ −12 dB) until the starter stops. The AVAudioSession / AudioFocus model. duckUiWhileContent(opt-in): while ANY content plays, a duck hold on theuibus — the "UI must not compete with the work" rule with an owner.- MediaSession projection of the active source when
metadatais given: the OS lock screen / media keys drive the handle (play/pause/seekto, plus position state); released when the source is destroyed. Players supply metadata; nobody touchesnavigator.mediaSessiondirectly. - Content-aware visibility:
autoSuspendnever suspends the context while a content source is playing.
Focus semantics, written down: ducking is a state derived from the playing
set — while any source with focus 'duck' plays, every other playing source
is ducked (a late arrival ducks on arrival); among simultaneous duckers the
most recently started one holds the floor. exclusive is an action of the
starter (the displaced source does not re-claim later, and is not restarted
when the winner ends — the AudioFocus model). Registering the same element
again replaces the stale registration; a remount never stacks listeners.
Honest limits, by construction: an unattached element does not pass through
the audio graph, so focus-ducking acts on el.volume (the handle's
snapshot().volume reports the OWNER's volume so a slider doesn't jerk during
a duck; a foreign el.volume write during a duck is clobbered on the next
policy write — set volume through the handle). attach() routes the element
into the content bus and is irreversible (createMediaElementSource),
and a cross-origin source without CORS goes mute — which is why attaching is
opt-in and never the default. MediaSession projects play / pause /
seekto + position state only; the other actions (stop, seekforward /
seekbackward, track skipping) are deliberately not projected — disposition:
add them with the player re-plan if its UX asks for them.
Voices — the calibration is the consumer's
The synthesis graph (two oscillators → biquad lowpass → optional tremolo →
ADSR-lite) is machinery and lives here. The tremolo sits IN SERIES before the
envelope — a gain node at 1 - depth modulated on its own .gain — so the
factor peaks at 1 and the trill scales with the note. Never patch a modulator
onto envelope.gain: a connection to an AudioParam ADDS to its automation
instead of scaling it, which leaves the depth constant while the envelope
moves (measured 2026-08-06: an audible tick closing every signal earcon, and
a signature declared at gain 0.03 coming out at RMS 0.143). Its perceptual CALIBRATION — the interval
ratio and mix, the envelope clamps, the contour sweep, the roughness→AM
mapping — is a SoundVoice, registered by whoever owns the vocabulary it
was tuned against (registerVoice(name, spec), the same shape as motion's
preset registry). Sema registers 'sema'; the engine ships a default voice
with the historical values so an unregistered caller still sounds and the
redesign changed nothing audible by itself. An unknown voice name falls back to
the default and logs once — audio is ornamental, it never throws.
This dissolves the earlier doctrine that baked the calibration into the art as
constants (PLAN-sound-engine.md §16.1, revoked): the machinery/doctrine
split now holds at the knowledge boundary too.
The split with sema
This art imports nothing from
$uix/sema. If it ever needs to, the cut was drawn wrong.
| Here (machinery + mix) | sema (doctrine) |
|---|---|
AudioContext lifecycle · unlock on gesture · bus tree + policy · synthesis + voice execution · sample cache · contour |
SOUNDS (the named catalogue) · SEMA_SOUND_VOICE (its calibration, as data) · the gesture resolvers · the cascade · which signature for which family × intent · the reduction policy |
A SoundSignature here is nine numeric knobs plus a contour — no family, no
intent, no evaluative loading. Sema's identically shaped type is structurally
assignable to it.
Two doors for consumers, and both lead to the same context
A consumer that needs its own graph (an AnalyserNode for a visualiser, gain
above 1, a BufferSource, sample-accurate scheduling) hangs it off
sound.context — or off a bus node (sound.bus('content').node) when it
should obey that bus's policy. That is the whole reason the art exists: one
owner, many consumers.
A consumer that needs the samples — a waveform reading peaks, an analyser
measuring an imported file — calls decode(bytes). Without that door it
would have to call decodeAudioData, and decodeAudioData needs a context, so
it would open its own: the exact second context this art prevents. That is not
hypothetical — the sema studio did it until 2026-07-30
(PLAN-sound-engine.md §16, A-1), and the art's own warning never fired,
because it only counts contexts it created.
decode returns null instead of throwing, decodes on a suspended context
(it never awaits resume(), which outside a user gesture can stay pending
forever), does not cache — raw bytes have no key — and lets the platform detach
the input buffer.
The ports
The art imports no other art. Both dependencies arrive injected as structural
ports — the MotionDom / SceneDom pattern:
SoundDom(getDocument+listen) — only to register the global unlock listener.ActiveDomsatisfies it. Without it the engine still plays; a context suspended by the autoplay policy is then resumed on the nextprime()/play()instead of on the next gesture.SoundTimers(thescheduleslice ofTimerScheduler) — the earcon's completion timer.ActiveTimers/uix.timerssatisfies it; without it the engine falls back tosetTimeout(direct unit tests only).
autoSuspend is opt-in — and that is the whole point
autoSuspend: true suspends the context while the tab is hidden and resumes it
when it returns. It defaults to false, deliberately:
Suspending on a hidden tab is the right policy for UI earcons and the wrong one for content. People listen to podcasts with the tab hidden. Which of the two you are is a decision of whoever composes — never of the engine.
And it is content-aware: a playing content source (sound.media(...))
vetoes the suspension — people listen to podcasts with the tab hidden.
suspend() / resume() called explicitly always win: the policy only
auto-resumes what it auto-suspended.
One live context, and it says so
Constructing a second engine is harmless — an engine with no context costs
nothing, which is exactly why ActiveUix can create one unconditionally. What is
almost always a wiring mistake is a second live context: browsers cap
concurrent AudioContexts and the unlock gesture is per-context, so one of the
two ends up mute.
So the art counts live contexts per realm and warns through its diagnostics
when a second one appears, pointing at the fix (take uix.sound, or pass
soundEngine to EngineSemantic). It is a warn, not an error: two contexts are
degraded, not broken.
Read the counter for what it is. It only sees the contexts this art
creates, so it catches double-engine wiring — not a bare new AudioContext()
somewhere else, which is the anti-pattern the message itself describes. Closing
that gap would mean monkey-patching the global constructor, which the art will
not do. It is written down instead, and the consumers who used to need their own
context now have doors that don't require one.
No side effects until it sounds
The constructor creates no context and registers no listener. An app that never plays a sound pays for nothing — and the sound path only enters the bundle when something imports this art.
Diagnostics
The full arts contract (diagnostics.ts, typed catalog, event names in
consts.ts) since the redesign — D-SR.9 reverted the earlier "mirror $scene,
no catalog" stance. Every failure is still a documented degradation, never a
programmer error (levels stay at debug / warn, no errors.ts): audio is
ornamental and must never abort its caller.
Behaviour notes
- Autoplay policy. The context starts suspended in Safari / iOS.
prime()creates + resumes it synchronously from inside the gesture; the unlock listener is registered afterwards, on the owner document, capture phase. - Synthesis. Two oscillators (sine + the voice's interval) → biquad lowpass
(
centroid) → ADSR-lite, clamped by the voice's envelope fractions so short notes keep an audible envelope instead of collapsing to a square. Roughness above the voice's threshold adds a fast AM modulator. Contour ridesosc.detuneover the voice's sweep. - Samples. A signature carrying
sampleUrlplays the decoded buffer (cached by URL) and falls back to synthesis when the bytes cannot be had, so the perceptual signal is never silently lost. Adata:URI is decoded in place rather than fetched:connect-src 'self'blocksfetch('data:…')(the directive matches by scheme) while letting a same-origin path through, so an inline pack routed throughfetchdegraded to synthesis under exactly the CSP that made someone inline it — and it is ~6x slower besides. - Errors are absorbed. Audio is ornamental: a failed earcon must never abort the caller's operation.
Composition
// active-app (attach path) — `defineUixServices()` already declares it
const App = createActiveApp({
services: { ...defineUixServices({ langs }), cache, session }
});
// active-uix (standalone) creates it directly and exposes uix.sound.
In both boot modes sema's sound channel receives this engine rather than making
its own: standalone passes soundEngine when it builds EngineSemantic, and in
attach defineEngineSemantic() takes sound as a service dependency and
forwards it.
Configuring the policy (masterGain / autoSuspend / contentFocus /
contentDuckFactor / duckUiWhileContent): in attach mode through
defineEngineSound({ … }) (the EngineSoundPolicyOptions slice). Standalone
createActiveUix does not expose it yet — disposition: plumb a sound
options slot when the first standalone consumer needs one, rather than growing
unused surface now. A hand-rolled service schema that declares events without sound
is the one case left where the channel opens its own context — pass
soundEngine: App.sound there.
Ownership rule for consumers: whoever creates the engine disposes it. A
consumer that receives a shared engine (injected by the composition root) must
never call dispose() on it — the sema studio (/temas/sema) rebuilds its
EngineSemantic on every draft edit and the context survives all of them
because of this rule.