You can not select more than 25 topics Topics must start with a letter or number, can include dashes ('-') and can be up to 35 characters long.
svelte-kit-vice/docs/architecture/agent.md

473 lines
29 KiB

This file contains invisible Unicode characters!

This file contains invisible Unicode characters that may be processed differently from what appears below. If your use case is intentional and legitimate, you can safely ignore this warning. Use the Escape button to reveal hidden characters.

This file contains ambiguous Unicode characters that may be confused with others in your current locale. If your use case is intentional and legitimate, you can safely ignore this warning. Use the Escape button to highlight these characters.

---
title: Agent — the delegation axis
type: reference
audience: human + agent
authority: E1 architecture — the agentic axis: the delegation model, the actor primitive, the participation contract, the minimum contracts, the a11y + threat doctrine
status: current — F0 doctrine (track PLAN-agent.md); decisions D-AG.1–11 + ⚖️1/2/3 signed 2026-07-21
sources:
canon: docs/CANON.md (the `delegate` family, ch. 29 of the book)
plan: docs/process/PLAN-agent.md (execution record — process, not doctrine)
review: docs/process/TRIAGE-revision-externa-agente.md (4-reviewer external critique)
---
# Agent — the delegation axis
> An **agent** is another actor. Not a mode, not a chat widget, not a wrapper
> around an LLM SDK. activeUIX already asks, in its semantic canon, _"¿quién
> actúa ahora?"_ — the `delegate` family (book ch. 29). The agent axis is the
> materialization of that family: the machinery by which **an actor other than
> the user** can be invoked from, and participate in, any component of the
> ecosystem — with the same events, the same undo, the same permissions, the
> same accessibility, and the same perceptual semantics that a user's action
> already gets.
This chapter is the standing doctrine — the **WHY**. What a conformant
implementation must actually do is specified normatively, with citable
requirement ids, in
[`spec/delegation-contract.md`](../spec/delegation-contract.md) (`DRAFT`); when
the two disagree about a requirement, the specification governs and this
chapter is corrected. The dated execution record (phases, signed decisions,
per-item work) lives in [`process/PLAN-agent.md`](../process/PLAN-agent.md) —
process, not truth.
---
## 0. Why an axis, not a component (and the reference pass)
The framework's own rule is to **study reference implementations before
designing** (as it did for the framework itself in `comparison.md`). The
agentic-UI field, surveyed 2026-07, converges on a small set of contracts —
and diverges from activeUIX in exactly the places that make this axis worth
building.
| Concern | Field convergence | activeUIX position |
| --------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Readable context | CopilotKit `useCopilotReadable`: NL-description + value, hierarchical (parentId), per-readable gating | `reads` in the per-component manifest; **push→pull** index + `read(instanceId)` on demand, prioritized by focus/viewport (a UIX-native signal none of them have) |
| Frontend tool calling | CopilotKit `useCopilotAction`, Vercel AI SDK tools, MCP tools; roundtrip closes via a `role:'tool'` message | `acts` in the manifest → the component's **existing public API**; result typed by sium, failures mapped to `delegate` verbs |
| Protocol / streaming | AG-UI: 17 events / 5 categories (run · text deltas · tool-calls · state · custom); triple start/content/end; state = snapshot + JSON-Patch delta (RFC 6902), resync by re-snapshot | Own protocol **mirroring AG-UI** (parity baseline, à la `ethereal` ↔ floating-ui); snapshot-per-turn in v1, delta reserved |
| Human-in-the-loop | Two shapes: pre-act approval (CopilotKit renderAndWait) + mid-run interrupt/elicitation (LangGraph interrupts, MCP elicitation). AG-UI models approval as a _confirmation-tool_, no separate primitive | The `delegate` cycle as **observable run states**; on the wire, review/elicitation ride the confirmation-tool shape (states in the motor, interop on the cable) |
| Embedded-UI security | MCP Apps (SEP-1865): sandboxed iframes, JSON-RPC host↔UI, **the host controls authorization** | The app is the ceiling of autonomy (D-AG.7); capabilities never self-authorize |
| Provider coupling | Most are React-first; **none ship a Svelte-native path** | Svelte-5-native by construction; provider adapters isolated behind a port |
**The gaps nobody fills** — and where this axis earns its keep:
- **A perceptual semantics of delegation.** No surveyed framework has a
vocabulary for _who acts_ rendered across sound / motion / a11y. The
`delegate` family is exactly that, and it predates the LLM wave.
- **Transactional undo of a run** (one delegation = one undo step) and
**reversibility declared by effect**, not by capability.
- **Accessibility of autonomous action**: attributed announcements, no
focus-theft, reduced-motion of the presence indicator.
- **Causal actor attribution** across an async, cross-component cascade —
synchronous, unforgeable, without a global async monkey-patch. A
browser-first UI framework that gets this right has no shipping precedent.
The conclusion of the pass: the correct shape is an **orthogonal axis** —
engine (an art) + per-component participation contract + registry + a few own
surfaces — not a monolithic agent component that would import every component
and invert the dependency direction. It instantiates the same anatomy every
other cross-cutting axis (sound, language, theme, motion) already uses.
Rejected alternatives, for the record: _protocol-only, no engine_ pushes the
state machine, budgets and elicitation into every app (re-implemented N times);
_agent-100%-backend_ kills frontend tool calling (the field's most load-bearing
contract) and the perceptual semantics; _generative-UI without manifests_
throws away the typed, auditable acts. Each can be layered on top later; none
can be the base.
---
## 1. The thesis
> **The agent is another actor.** Everything downstream of that sentence is
> the design.
Four principles, none negotiable without re-signing:
1. **Actor-agnostic** (book ch. 29 §1). The axis never asks "is this an LLM?".
A macro, an automation rule, a workflow, a remote human approver and a
model are all just _another actor_. Nothing in the axis depends on a
concrete LLM provider. This is what lets the architecture outlive the LLM
market.
2. **The agent knows no components.** It never names `Palabras`, `Chronos`,
`Chat`. It sees only **capabilities** declared by whatever providers are
mounted. This keeps the dependency direction intact: the canon never
imports the axis.
3. **The golden rule — same public API, actor-dependent concretion.** The
agent acts through the **same public provider API** the app's own code
would call; the call path does not bifurcate. But the **semantic concretion
depends on the actor**: a component's existing events declare `delegate`
capacity (polymorphic events, `allowedFamilies` in the morfo contract) and
the provider concretes `{ family: 'delegate', verb: 'act' }` when the actor
is the agent. So the same `insert()` fires the same `runtime.trigger`, but
sema, eidos and the a11y layer read the correct answer to _"¿quién actúa?"_.
If a capability needs a door the public API lacks, that door is **promoted
to the component** through the normal route — never bypassed.
4. **Total degradation.** With the axis absent, every component works at 100%.
Manifest registration is a no-op; nothing a component owns depends on the
agent existing. See §4.
A corollary the field taught us the hard way: **client-side budgets are UX
courtesy, not a security boundary.** The boundary lives where the credential
lives (a proxy/backend). The motor's budgets exist to drive the state machine
and the UX, and the doctrine says so honestly.
---
## 2. The domain model
The axis is built around a **delegation**, not a conversation. A chat is one
surface that chains delegations; it is not the core (an agent can fill a
document with no conversation at all).
```
ActiveAgent — the session (long-lived; context across runs; identity)
└── Run — a delegation: budget, authorization, undo grouping, control return
├── Turn — one exchange with the model
│ └── CapabilityCall — a delegate.act on a provider (the only thing that mutates)
└── Context — snapshot of reads at turn open (§3)
```
Three readings this ladder settles:
- **The `CapabilityCall` is what mutates the world; the `Run` is the unit of
delegation.** A mutation without the delegation envelope (budget,
authorization, undo grouping, control return) is exactly what the doctrine
forbids.
- **The degenerate Run** — a single auto-authorized call — is the non-AI case
of the book (a macro, a rule, a workflow). It is a valid, cheap Run.
- **Initiative is modeled, not just actor.** Every Run carries
`initiative: 'user' | 'system' | 'scheduled'` — _why_ it acts, not only
_who_ — because it changes trust, announcement, audit and the autonomy
ceiling. The entry verb hints it: a cycle entered by `offer` is
system-initiative.
**Ambient mode** (always-on: autocomplete, predictive suggestions) is **out of
the axis in v1, reserved by name**: a `Subscription`/watcher emitting
`delegate.offer`, with coalesced micro-runs as future design. It is named so a
second, parallel agency system cannot grow in by the back door.
### The run state machine = the `delegate` verbs, literally
The states **are** the verbs — as _custody of control_, not as cognitive
phases. The model's own plan→tool→plan loop lives entirely inside `acting`
(across turns); the machine models who holds control, not how the agent thinks.
```
┌─────────── escalated ──────────┐ (non-terminal; timeout → returned)
▼ │
offered → planned → reviewing ⇄ authorized → acting ──┤
│ │ (re-entrant: (fixes the │ ▼
│ │ partial allowlist + │ returned ← control ALWAYS returns
│ │ approvals) budget) │ (outcome-typed)
└─────────┴──── cancelled ────────────────────────┘
```
- **Authorization fixes scope, not vibes.** `authorized` derives from the plan
a **capability allowlist + budget**; an act outside that scope ⇒ automatic
`escalate` to re-review (never a silent failure, never teatro). This makes
`reviewing` re-entrant from `acting`.
- **`returned` is outcome-typed**: `completed | partial | rolled-back |
aborted | error`, with a `reason` (`superseded` — the premise changed under
the run; `budget-exhausted`; `killed`). Mid-sequence failure **rolls back to
run-start by default** (the user authorized the whole plan, not half); a
capability may declare it accepts partial close.
- **`escalated`** is non-terminal, carries a typed `reason` (`elicitation |
insufficient-scope | budget-exhausted | out-of-plan`), and **auto-returns on
timeout** (an abandoned escalation still needs an exit).
- **Pause needs no new verb**: pause = `return` with resumable context; resume
= a fresh Run chained in the session — which coincides with AG-UI's `resume`
mechanics. Our semantics and its wire converge without trying to.
- **User interruption = immediate `return`**, never an error. A **global kill
switch** on the motor (`disable()`) sends every run to `returned(aborted)`.
---
## 3. The participation contract
A component participates by declaring a **manifest** — the mirror of a sema
pack or a langs catalog, in its own tree (`src/uix/agent/components/{x}.ts`).
The manifest satisfies **structural contracts defined by the art**
(`AgentCapability` / `AgentRead`: string ids, Standard-Schema args via sium,
opaque handler) — the MotionDom / SceneDom port pattern — by typing against the
provider's public API on the uix side. Soma/uix types never cross into the art;
the opaque handler grants location-transparency for free (a reserved path to
distributed delegation).
- **Model-facing projection.** TypeScript never reaches the model. Every read
and act carries a natural-language `description`; args project as **JSON
Schema** (a sium→JSON-Schema emitter), with **size limits per arg** as part
of the contract. Traits annotate each act for the reasoner (cost, latency,
stability, staleness) — the MCP tool-annotations pattern.
- **Instances and context — push→pull.** An always-present **index** (stable id
- type + human label + one-line summary, hierarchical) plus deep reads as a
**pull capability** (`read(instanceId)`), prioritized by **focus/viewport**;
addressing ambiguity ("the editor", with two mounted) ⇒ **elicitation, not a
heuristic**. `available` is a reactive predicate per act; a
`capabilities-changed` event refreshes the toolset between turns.
- **Results and failures** are sium-typed; a result carrying third-party
content is as untrusted as an imported document (same origin doctrine, §5).
`unknown-capability` (a hallucinated tool) returns as a tool-result for
self-correction, bounded by budget.
- **Altitude, not RPC.** Capabilities are **mechanical world-mutations**, few
and at the right altitude (`replaceRange`, not twelve micro-inserts).
Semantic operations (`summarize`, `translate`) are the _model's_ cognition or
a reserved agent-side "skills" tier (the MCP-prompts analogue) — never lodged
in the component, which would invert the dependency direction.
- **Maturity tiers.** A manifest declares its tier — `reads-only` →
`reversible acts` → `full acts` — so the ~110 canon components adopt
incrementally without gating the axis; the guard audits the declared tier.
- **Reversibility is by effect, not by capability.** External effects
(send / publish / notify) are irreversible by definition; the reversion
authority is the component's **native history**, never orca compensation.
- **OCC.** Each read exports a `_version`; each act must carry it; a mismatch
is a recoverable `StateStaleError` (re-read → re-plan) — the answer to TOCTOU
(§5).
---
## 4. Minimum contracts (the §0 row) and degradation
Following the framework's rule that every change derive from the module
contract table:
```text
Module Requires Optional When missing Error
agent AgentTransport port orca, perm, prefs, manifest registration is a NO-OP; only on invoking a
(app supplies it) connection, storage every component stays 100% functional capability with no transport
```
- The agent is an **app-level art** (factory `defineActiveAgent(...)`, the
auth/perm/cache pattern), not something `ActiveUix` builds in standalone.
Promotion to a read-only `uix.agent` accessor in attach mode arrives with the
canonical consumer (the scene-D4 precedent).
- **A component with a manifest MUST behave identically with the agent
absent.** The manifest is inert data until a running `ActiveAgent` reads it.
- Transport is **app-land by default** (the chat-family precedent: "the
transport stays app-land"); a server-authoritative `svrs/agent` is a
conditioned future initiative, never the v1 assumption. Adapters live in
`$agent/adapters/*` outside the main barrel; a **deterministic
`ScriptedAgentTransport`** is first-class, so static demos and component
tests run with **real instances, no mocks**.
---
## 5. The actor primitive (⚖️2) — causal attribution
The load-bearing decision of the axis. When the agent's act on component A
triggers events that cascade to mutate component B, **B's emission must carry
the correct actor** — or sema announces the agent's secondary effect as "user",
a11y and audit-trail catastrophe.
**The answer is the OTel-baggage unification, corrected by the browser caveat.**
The actor is an **opaque, engine-minted token** (`ActorToken`, minted and
resolved only by `EngineAgent`; a non-token `USER` sentinel everywhere else, so
absence can never be confused with a minted token) carried as **one reserved
slot on the causal envelope the framework already propagates** — `OrcaEventMeta`
on the orca path, the bus envelope's `context` bag on the bus path — inherited
depth-to-depth exactly as `traceId` is. It is read **synchronously** via
`actor?: ActorToken` on `TriggerOptions → SemanticSignal`, before the first
`await` in `EngineSemantic.emit`.
This **shares one propagation primitive** with the D-AG.9 trace port
(run = trace, CapabilityCall = span, **actor = baggage**) — distinct payloads,
distinct guarantees: the span export is async and loss-tolerant; the actor read
is synchronous and authoritative. Sharing the primitive avoids forking agency;
keeping the payloads distinct respects that sema cannot tolerate the
fall-to-root a tracer tolerates.
Why **explicit, not ambient**: native `AsyncContext`/ALS is the framework's own
orca mechanism, already adjudicated — in the browser (`als.ts` feature-detects
to `null`) it degrades to a root envelope with no actor, precisely the failure
this decision exists to avoid. A faithful zero-dep reimplementation is
impossible without Zone.js-style global monkey-patching (forbidden, and it
still cannot cover a bare `await`). So **ambient auto-fill is opportunistic
sugar; the explicit envelope is the guarantee** — which is exactly how
OpenTelemetry-JS actually works (baggage on a propagated context, synchronous
`StackContextManager` in browsers, `async_hooks` server-side only). This is the
industry consensus in one line: _explicit provenance bound to the message,
ambient as sugar_ — OTel propagators inject/extract explicitly, CaMeL binds an
unforgeable capability tag to the value, Yjs attributes derived mutations by an
explicit per-transaction `origin`.
Scope and honesty:
- **Anti-forgery** (the reason the token is opaque): no site outside
`arts/agent` may construct an `ActorToken` or write the reserved actor key;
the token is never wire-serialized (the trace-export boundary emits a
non-authoritative `'agent'` label only). The `ActorToken` **type** lives in a
leaf below both orca and agent to avoid an orca→agent cycle; minting
authority stays in `EngineAgent`. On the bus path the actor rides the untyped
`context` bag, so its anti-forgery defense is the `agent-check` lint, not the
type system.
- **The guarantee is scoped.** It is delivered for the **explicitly-authored
A→B bridge** that threads `opts.actor`. An agent driving a stock component
through its public API has no call-site to thread it and, in the browser,
resolves to the `USER` sentinel — best-effort, out of v1. **In-scope
attribution rides the model-derived `data-actor`** paint (eidos reads the
reactive run state during `acting`), which needs no threading; the envelope
baggage is specifically for **out-of-scope secondary effects + audit**.
- **The primitive ships first, the propagation retrofit later.** The minimal
primitive (branded token + `USER` sentinel + private registry + anti-forgery
guard + synchronous `opts.actor` read) is built up front. The
~6-seam envelope-propagation retrofit (`OrcaEventMeta.actor`,
`OrcaALSContext.actor`, `ctx.emit` forwarding, `buildChildEnvelope` copy,
bus-subscription forwarding, `publishCausedBy` parent-merge) is **deferred to
the phase with a real out-of-scope async-cascade consumer**, gated by the
guard's drift check that ambient is never the sole carrier on a
cross-component bridge.
---
## 6. Authorization, consent and autonomy
Three levels — `suggest | review | auto` — resolved prefs-style (the app is the
ceiling; the user adjusts downward; a capability may demand _more_ control,
never less; **deny-overrides**). Irreversible ⇒ `review` always.
- **Dynamic risk by payload**: the manifest declares a risk classifier over
args; `auto` escalates to `review` per invocation (same capability, different
blast radius).
- **Context origin lowers the ceiling** — the unification that dissolves
prompt-injection from an open threat into a policy row: each read carries an
origin tag (`authored-by-user | imported | third-party | other-participant`);
**any untrusted origin in the turn lowers the ceiling** (to `review`, or
`suggest` for sensitive capabilities). This is the industry pattern (forced
human approval when untrusted content is in context) expressed through
machinery already designed.
- **Initiative affects ceilings**: autonomous initiative (`system` /
`scheduled`) defaults stricter than user-requested.
- **Authorization is bound to state, not prose** (TOCTOU): review computes
against a fingerprint `F` of the relevant reads; authorization binds to `F`
**plus a TTL**; drift ⇒ recompute the diff (identical → optional
auto-re-approve; different → re-review). The user↔agent concurrency problem
is the OT/CRDT model (Yjs) — adopt the mental model, not the library.
- **Approval fatigue**: persistent grants ("always allow X") carry a TTL and an
explicit scope (session / component / act type); review shows the **concrete
proposed change in the app's own domain, never prose** (amended 2026-07-22 —
see "The review surface" in §7: the content of a review is app-owned). A
`capabilities-changed` event re-gates grants.
- **perm** governs only server-touching scopes (enforced in the proxy /
server-side tool-loop); the base case is local consent with no perm
dependency. **`reads` are permission subjects symmetric to `acts`** —
sensitive components (credentials/passwords) never register reads.
---
## 7. Semantics, presence and accessibility
- The motor **emits nothing** (an art with no DOM); it exposes the run as
reactive state. **Surfaces** declare the cycle in their morfos (`delegate-*`);
`sustain-processing` belongs to **Aura** (state-bound persistence, cleared on
run close). **Actuated components** concrete `delegate.act` via
`allowedFamilies` with the actor context.
- **A run above `suggest` requires a mounted presence surface** (a minimal
Aura). A run in `acting` with nowhere to emit its cycle is perceptually
invisible — a direct contradiction of _"¿quién actúa?"_. The motor refuses
(or warns) on entering `acting` without one; this is mechanically verifiable.
- **Write-scope is declared**: a run declares its scope (component | part |
range); hold and paint apply only there.
- **The review surface** (amended 2026-07-22 — supersedes the discarded
"review-card" idea, a dev-tool anchoring). The framework/app split is:
the FRAMEWORK owns the cycle states expressed perceptually (`reviewing` /
`escalated` surfaced by the presence surface, with authorize/reject/cancel
affordances — Aura already does this), the "proposed by another actor"
semantics, and the attributed announcements. The APP owns the CONTENT of
what is being reviewed, rendered in its own domain (a document shows its
proposed blocks, a calendar its pending events, a form its suggested
values) — in context, where the user already is. A git-style diff card is
a legitimate construct for a specific application (e.g. a code tool built
ON the framework); it is never canon.
The **a11y contract** — no framework has this; documenting it fixes the pattern
others will cite:
1. **One `polite` live region owned by Aura**; announcements **attributed**
("Agente: …") and **coalesced by plan/batch** — never per delta, never per
act (WCAG 4.1.3).
2. **No focus-theft — a motor invariant** (WCAG 3.2.1 / 3.2.2), not each
surface's courtesy.
3. `aria-busy` on the region being mutated during streaming; a completion
announcement on close **that carries the outcome** — a failure returning
control silently is the same invisibility as an act with nowhere to emit.
The surface expresses it with THREE terminal events, intent intrinsic to
each (a dynamic intent has nowhere to live: the trigger options carry none
and a morfo's `fromProp` binds to a public prop, not to machine state):
`completed` → fulfill · `aborted`/`partial` → no intent (a decline is
absence, not loss) · `error`/`rolled-back` → loss. All three stay
TRANSIENT: once control is back there is no custody to claim, so a
persistent alarm would make the surface lie about who holds it — the
failure's DETAIL belongs to the app, in its own domain. Binding a tone to
this structural family is the sanctioned `BK-FRAME-NO-INTENT` exception
(book-updates A-1) and costs a written `intentRationale`: the engine emits
nothing and no `signal`/`commit` event exists on the presence surface, so
the consolidated return is the outcome's only expressive site.
4. Elicitations are **dismissible / postponable**, not modal by default
(WCAG 2.2.4). An elicitation is a **typed question**, so its surface is a
form in context — not a conversation. The schema the agent asks with is the
same projection its capability args already travel through (sium →
JSON Schema), read in the opposite direction; binding elicitation to a chat
would make a general contract depend on one product shape.
5. **A global keyboard exit** (an Esc-like doctrine): cancel / return reachable
without a mouse. Aura uses `reduce: 'static-frame'`, never `hide`.
---
## 8. Threat model
The axis builds, by construction, the classic **"lethal trifecta"** (Willison):
private data + untrusted content + the ability to act. The doctrine treats it
head-on:
- **Prompt injection.** Reads carry an **origin tag** (per §6); untrusted
origin lowers the autonomy ceiling (the primary mitigation, folded into
policy). Adapters **spotlight** (delimit/mark) non-authored content so the
model treats it as data. Reserved for depth: dual-LLM (a privileged planner
that never sees untrusted content; a quarantined model that returns only
symbolic references) and CaMeL-style capability/data-flow tracking —
philosophically close to the framework's structural ports. Checklist anchors:
OpenAI instruction-hierarchy, OWASP LLM Top 10 (LLM01).
- **TOCTOU.** Authorization bound to a read fingerprint + TTL (§6).
- **Actor forgery.** The opaque engine-minted token + the `agent-check` guard
(§5).
- **Run recovery.** A minimal **write-ahead journal per run** via `$storage`
(sessionStorage suffices in v1) records authorizations + applied acts and
detects orphaned runs on reload. The audit-trail is defined **now** as an
OTel-shaped trace port (run = trace, CapabilityCall = span); the storage
vehicle arrives later. (orca cannot be a store — its own rule.)
- **Kill switch, multi-user, i18n.** A global disable (§2). User identity in
the actor vocabulary is reserved (multi-user arrives with `connection`). The
agent's `identity` / system-context carries the effective locale; the agent
announces in the langs-axis language like any surface.
---
## 9. What this axis is NOT
- **Not a chat.** Chat is one surface (a participant + a composer) that chains
runs; the axis works with no conversation at all.
- **Not an LLM integration.** It is a theory of delegation; a model is one kind
of actor. Nothing depends on a provider SDK (adapters are isolated and
optional).
- **Not a parallel "AI mode".** It preserves every framework invariant —
dependency direction, semantics, events, public API, undo, a11y — instead of
forking them. A component with a manifest is the same component.
- **Not a component that knows components.** It sees capabilities, never types.
---
## 10. To go deeper
- Vocabulary of `delegate` (families, verbs, intent policy): [`CANON.md`](../CANON.md).
- The whole-system layering the axis plugs into:
[`architecture/active-architecture.md`](./active-architecture.md).
- The arts pattern (`Engine*` / `Active*`, ports, degradation):
[`arts/README.md`](../../src/arts/README.md).
- Execution record, signed decisions, phases:
[`process/PLAN-agent.md`](../process/PLAN-agent.md).
- The external critique that hardened this doctrine:
[`process/TRIAGE-revision-externa-agente.md`](../process/TRIAGE-revision-externa-agente.md).

Powered by TurnKey Linux.