You can not select more than 25 topics Topics must start with a letter or number, can include dashes ('-') and can be up to 35 characters long.
svelte-kit-vice/src/arts/sound/README.md

15 KiB

sound — the sound orchestrator

EngineSound governs every aspect of sound in the framework (redesign 2026-07-31, PLAN-sound-redesign.md): the substrate (one AudioContext per document and its unlock lifecycle), the mix (master → buses), the synthesis machinery with registered voices, the sample path, and media citizenship (the playback transport the service PROVIDES, audio focus, the UI↔content relation, MediaSession). Nobody else creates an AudioContext or touches Web Audio directly, the way every timer goes through uix.timers and every animation through uix.motion. Consumers adapt to the service, not the other way around.

const sound = createEngineSound({ dom, timers }); // dom: a SoundDom port

sound.prime(); // SYNCHRONOUS — call it inside the user gesture
await sound.play(sig, { bus: 'ui', voice: 'sema', gainScale }); // an earcon; never rejects
await sound.preload(urls); // pre-decode samples into the URL-keyed cache
await sound.decode(bytes); // raw bytes → AudioBuffer, on THIS context

sound.master.setGain(0.4); // document-wide policy, the root of the bus tree
sound.bus('ui').setGain(0.5); // per-bus policy: gain / mute / duck holds
sound.bus('ui').setMuted(true);
const hold = sound.bus('ui').duck(0.2); // refcounted; strongest factor wins
hold.release();

sound.registerVoice('sema', voiceSpec); // calibration is the CONSUMER's data
sound.suspend() / sound.resume(); // explicit; wins over the autoSuspend policy
sound.dispose(); // idempotent

sound.state; // 'absent' | 'suspended' | 'running'
sound.context; // AudioContext | null — the raw output

The mix — master → ui + content

"Sound" means two different things in a UI framework, so the graph has two buses under the master: ui carries the earcons (sema's channel) and content carries the work itself (media playback, voice notes). Each bus is a policy surface — gain, muted, refcounted duck() holds (the strongest live factor wins; they don't multiply) — over a GainNode you can also reach raw (bus.node) to hang an analyser or a source.

The structural invariant this exists for: UI-sound policy must never touch the content's volume. Honest about the mechanism (audit AU-4): today prefs.sound acts through sema's channel — BK-REDUCTIONS resolves the level and hands it over as gainScale, always on the ui bus — and nothing wires prefs to a bus directly. The ui bus gain/mute is the graph-side LEVER an app may additionally bind (e.g. prefs.sound → bus('ui').setGain); what is guaranteed by construction is the negative: no UI-side policy ever writes the content bus or a source's own volume. And the three levers never write the same value: the master carries document policy, the bus carries per-bus policy, and gainScale rides one playback's own envelope.

Media citizenship — the service provides the transport

const media = sound.media(el, {
	focus: 'exclusive', // per-source override of `contentFocus`
	metadata: { title, artist, album, artwork } // MediaSession opt-in
});
// media = the player's MediaProvider shape: play/pause/seek/setVolume/
// setMuted/setPlaybackRate/snapshot/subscribe/destroy — plus attach().

Registering a content source is what buys the citizenship (D-SR.4…D-SR.6):

  • Audio focus toward other content sources — contentFocus (service-wide, default 'mixed') or per-source: exclusive pauses the others, duck attenuates them (contentDuckFactor, default 0.25 ≈ −12 dB) until the starter stops. The AVAudioSession / AudioFocus model.
  • duckUiWhileContent (opt-in): while ANY content plays, a duck hold on the ui bus — the "UI must not compete with the work" rule with an owner.
  • MediaSession projection of the active source when metadata is given: the OS lock screen / media keys drive the handle (play/pause/seekto, plus position state); released when the source is destroyed. Players supply metadata; nobody touches navigator.mediaSession directly.
  • Content-aware visibility: autoSuspend never suspends the context while a content source is playing.

Focus semantics, written down: ducking is a state derived from the playing set — while any source with focus 'duck' plays, every other playing source is ducked (a late arrival ducks on arrival); among simultaneous duckers the most recently started one holds the floor. exclusive is an action of the starter (the displaced source does not re-claim later, and is not restarted when the winner ends — the AudioFocus model). Registering the same element again replaces the stale registration; a remount never stacks listeners.

Honest limits, by construction: an unattached element does not pass through the audio graph, so focus-ducking acts on el.volume (the handle's snapshot().volume reports the OWNER's volume so a slider doesn't jerk during a duck; a foreign el.volume write during a duck is clobbered on the next policy write — set volume through the handle). attach() routes the element into the content bus and is irreversible (createMediaElementSource), and a cross-origin source without CORS goes mute — which is why attaching is opt-in and never the default. MediaSession projects play / pause / seekto + position state only; the other actions (stop, seekforward / seekbackward, track skipping) are deliberately not projected — disposition: add them with the player re-plan if its UX asks for them.

Voices — the calibration is the consumer's

The synthesis graph (two oscillators → biquad lowpass → optional tremolo → ADSR-lite) is machinery and lives here. The tremolo sits IN SERIES before the envelope — a gain node at 1 - depth modulated on its own .gain — so the factor peaks at 1 and the trill scales with the note. Never patch a modulator onto envelope.gain: a connection to an AudioParam ADDS to its automation instead of scaling it, which leaves the depth constant while the envelope moves (measured 2026-08-06: an audible tick closing every signal earcon, and a signature declared at gain 0.03 coming out at RMS 0.143). Its perceptual CALIBRATION — the interval ratio and mix, the envelope clamps, the contour sweep, the roughness→AM mapping — is a SoundVoice, registered by whoever owns the vocabulary it was tuned against (registerVoice(name, spec), the same shape as motion's preset registry). Sema registers 'sema'; the engine ships a default voice with the historical values so an unregistered caller still sounds and the redesign changed nothing audible by itself. An unknown voice name falls back to the default and logs once — audio is ornamental, it never throws.

This dissolves the earlier doctrine that baked the calibration into the art as constants (PLAN-sound-engine.md §16.1, revoked): the machinery/doctrine split now holds at the knowledge boundary too.

The split with sema

This art imports nothing from $uix/sema. If it ever needs to, the cut was drawn wrong.

Here (machinery + mix) sema (doctrine)
AudioContext lifecycle · unlock on gesture · bus tree + policy · synthesis + voice execution · sample cache · contour SOUNDS (the named catalogue) · SEMA_SOUND_VOICE (its calibration, as data) · the gesture resolvers · the cascade · which signature for which family × intent · the reduction policy

A SoundSignature here is nine numeric knobs plus a contour — no family, no intent, no evaluative loading. Sema's identically shaped type is structurally assignable to it.

Two doors for consumers, and both lead to the same context

A consumer that needs its own graph (an AnalyserNode for a visualiser, gain above 1, a BufferSource, sample-accurate scheduling) hangs it off sound.context — or off a bus node (sound.bus('content').node) when it should obey that bus's policy. That is the whole reason the art exists: one owner, many consumers.

A consumer that needs the samples — a waveform reading peaks, an analyser measuring an imported file — calls decode(bytes). Without that door it would have to call decodeAudioData, and decodeAudioData needs a context, so it would open its own: the exact second context this art prevents. That is not hypothetical — the sema studio did it until 2026-07-30 (PLAN-sound-engine.md §16, A-1), and the art's own warning never fired, because it only counts contexts it created.

decode returns null instead of throwing, decodes on a suspended context (it never awaits resume(), which outside a user gesture can stay pending forever), does not cache — raw bytes have no key — and lets the platform detach the input buffer.

The ports

The art imports no other art. Both dependencies arrive injected as structural ports — the MotionDom / SceneDom pattern:

  • SoundDom (getDocument + listen) — only to register the global unlock listener. ActiveDom satisfies it. Without it the engine still plays; a context suspended by the autoplay policy is then resumed on the next prime() / play() instead of on the next gesture.
  • SoundTimers (the schedule slice of TimerScheduler) — the earcon's completion timer. ActiveTimers / uix.timers satisfies it; without it the engine falls back to setTimeout (direct unit tests only).

autoSuspend is opt-in — and that is the whole point

autoSuspend: true suspends the context while the tab is hidden and resumes it when it returns. It defaults to false, deliberately:

Suspending on a hidden tab is the right policy for UI earcons and the wrong one for content. People listen to podcasts with the tab hidden. Which of the two you are is a decision of whoever composes — never of the engine.

And it is content-aware: a playing content source (sound.media(...)) vetoes the suspension — people listen to podcasts with the tab hidden. suspend() / resume() called explicitly always win: the policy only auto-resumes what it auto-suspended.

One live context, and it says so

Constructing a second engine is harmless — an engine with no context costs nothing, which is exactly why ActiveUix can create one unconditionally. What is almost always a wiring mistake is a second live context: browsers cap concurrent AudioContexts and the unlock gesture is per-context, so one of the two ends up mute.

So the art counts live contexts per realm and warns through its diagnostics when a second one appears, pointing at the fix (take uix.sound, or pass soundEngine to EngineSemantic). It is a warn, not an error: two contexts are degraded, not broken.

Read the counter for what it is. It only sees the contexts this art creates, so it catches double-engine wiring — not a bare new AudioContext() somewhere else, which is the anti-pattern the message itself describes. Closing that gap would mean monkey-patching the global constructor, which the art will not do. It is written down instead, and the consumers who used to need their own context now have doors that don't require one.

No side effects until it sounds

The constructor creates no context and registers no listener. An app that never plays a sound pays for nothing — and the sound path only enters the bundle when something imports this art.

Diagnostics

The full arts contract (diagnostics.ts, typed catalog, event names in consts.ts) since the redesign — D-SR.9 reverted the earlier "mirror $scene, no catalog" stance. Every failure is still a documented degradation, never a programmer error (levels stay at debug / warn, no errors.ts): audio is ornamental and must never abort its caller.

Behaviour notes

  • Autoplay policy. The context starts suspended in Safari / iOS. prime() creates + resumes it synchronously from inside the gesture; the unlock listener is registered afterwards, on the owner document, capture phase.
  • Synthesis. Two oscillators (sine + the voice's interval) → biquad lowpass (centroid) → ADSR-lite, clamped by the voice's envelope fractions so short notes keep an audible envelope instead of collapsing to a square. Roughness above the voice's threshold adds a fast AM modulator. Contour rides osc.detune over the voice's sweep.
  • Samples. A signature carrying sampleUrl plays the decoded buffer (cached by URL) and falls back to synthesis when the bytes cannot be had, so the perceptual signal is never silently lost. A data: URI is decoded in place rather than fetched: connect-src 'self' blocks fetch('data:…') (the directive matches by scheme) while letting a same-origin path through, so an inline pack routed through fetch degraded to synthesis under exactly the CSP that made someone inline it — and it is ~6x slower besides.
  • Errors are absorbed. Audio is ornamental: a failed earcon must never abort the caller's operation.

Composition

// active-app (attach path) — `defineUixServices()` already declares it
const App = createActiveApp({
	services: { ...defineUixServices({ langs }), cache, session }
});

// active-uix (standalone) creates it directly and exposes uix.sound.

In both boot modes sema's sound channel receives this engine rather than making its own: standalone passes soundEngine when it builds EngineSemantic, and in attach defineEngineSemantic() takes sound as a service dependency and forwards it.

Configuring the policy (masterGain / autoSuspend / contentFocus / contentDuckFactor / duckUiWhileContent): in attach mode through defineEngineSound({ … }) (the EngineSoundPolicyOptions slice). Standalone createActiveUix does not expose it yet — disposition: plumb a sound options slot when the first standalone consumer needs one, rather than growing unused surface now. A hand-rolled service schema that declares events without sound is the one case left where the channel opens its own context — pass soundEngine: App.sound there.

Ownership rule for consumers: whoever creates the engine disposes it. A consumer that receives a shared engine (injected by the composition root) must never call dispose() on it — the sema studio (/temas/sema) rebuilds its EngineSemantic on every draft edit and the context survives all of them because of this rule.

Powered by TurnKey Linux.