You can not select more than 25 topics Topics must start with a letter or number, can include dashes ('-') and can be up to 35 characters long.
svelte-kit-vice/docs/architecture/agent.md

29 KiB

title: Agent — the delegation axis type: reference audience: human + agent authority: E1 architecture — the agentic axis: the delegation model, the actor primitive, the participation contract, the minimum contracts, the a11y + threat doctrine status: current — F0 doctrine (track PLAN-agent.md); decisions D-AG.1–11 + ⚖️1/2/3 signed 2026-07-21 sources: canon: docs/CANON.md (the delegate family, ch. 29 of the book) plan: docs/process/PLAN-agent.md (execution record — process, not doctrine) review: docs/process/TRIAGE-revision-externa-agente.md (4-reviewer external critique)

Agent — the delegation axis

An agent is another actor. Not a mode, not a chat widget, not a wrapper around an LLM SDK. activeUIX already asks, in its semantic canon, "¿quién actúa ahora?" — the delegate family (book ch. 29). The agent axis is the materialization of that family: the machinery by which an actor other than the user can be invoked from, and participate in, any component of the ecosystem — with the same events, the same undo, the same permissions, the same accessibility, and the same perceptual semantics that a user's action already gets.

This chapter is the standing doctrine — the WHY. What a conformant implementation must actually do is specified normatively, with citable requirement ids, in spec/delegation-contract.md (DRAFT); when the two disagree about a requirement, the specification governs and this chapter is corrected. The dated execution record (phases, signed decisions, per-item work) lives in process/PLAN-agent.md — process, not truth.


0. Why an axis, not a component (and the reference pass)

The framework's own rule is to study reference implementations before designing (as it did for the framework itself in comparison.md). The agentic-UI field, surveyed 2026-07, converges on a small set of contracts — and diverges from activeUIX in exactly the places that make this axis worth building.

Concern Field convergence activeUIX position
Readable context CopilotKit useCopilotReadable: NL-description + value, hierarchical (parentId), per-readable gating reads in the per-component manifest; push→pull index + read(instanceId) on demand, prioritized by focus/viewport (a UIX-native signal none of them have)
Frontend tool calling CopilotKit useCopilotAction, Vercel AI SDK tools, MCP tools; roundtrip closes via a role:'tool' message acts in the manifest → the component's existing public API; result typed by sium, failures mapped to delegate verbs
Protocol / streaming AG-UI: 17 events / 5 categories (run · text deltas · tool-calls · state · custom); triple start/content/end; state = snapshot + JSON-Patch delta (RFC 6902), resync by re-snapshot Own protocol mirroring AG-UI (parity baseline, à la ethereal ↔ floating-ui); snapshot-per-turn in v1, delta reserved
Human-in-the-loop Two shapes: pre-act approval (CopilotKit renderAndWait) + mid-run interrupt/elicitation (LangGraph interrupts, MCP elicitation). AG-UI models approval as a confirmation-tool, no separate primitive The delegate cycle as observable run states; on the wire, review/elicitation ride the confirmation-tool shape (states in the motor, interop on the cable)
Embedded-UI security MCP Apps (SEP-1865): sandboxed iframes, JSON-RPC host↔UI, the host controls authorization The app is the ceiling of autonomy (D-AG.7); capabilities never self-authorize
Provider coupling Most are React-first; none ship a Svelte-native path Svelte-5-native by construction; provider adapters isolated behind a port

The gaps nobody fills — and where this axis earns its keep:

  • A perceptual semantics of delegation. No surveyed framework has a vocabulary for who acts rendered across sound / motion / a11y. The delegate family is exactly that, and it predates the LLM wave.
  • Transactional undo of a run (one delegation = one undo step) and reversibility declared by effect, not by capability.
  • Accessibility of autonomous action: attributed announcements, no focus-theft, reduced-motion of the presence indicator.
  • Causal actor attribution across an async, cross-component cascade — synchronous, unforgeable, without a global async monkey-patch. A browser-first UI framework that gets this right has no shipping precedent.

The conclusion of the pass: the correct shape is an orthogonal axis — engine (an art) + per-component participation contract + registry + a few own surfaces — not a monolithic agent component that would import every component and invert the dependency direction. It instantiates the same anatomy every other cross-cutting axis (sound, language, theme, motion) already uses.

Rejected alternatives, for the record: protocol-only, no engine pushes the state machine, budgets and elicitation into every app (re-implemented N times); agent-100%-backend kills frontend tool calling (the field's most load-bearing contract) and the perceptual semantics; generative-UI without manifests throws away the typed, auditable acts. Each can be layered on top later; none can be the base.


1. The thesis

The agent is another actor. Everything downstream of that sentence is the design.

Four principles, none negotiable without re-signing:

  1. Actor-agnostic (book ch. 29 §1). The axis never asks "is this an LLM?". A macro, an automation rule, a workflow, a remote human approver and a model are all just another actor. Nothing in the axis depends on a concrete LLM provider. This is what lets the architecture outlive the LLM market.

  2. The agent knows no components. It never names Palabras, Chronos, Chat. It sees only capabilities declared by whatever providers are mounted. This keeps the dependency direction intact: the canon never imports the axis.

  3. The golden rule — same public API, actor-dependent concretion. The agent acts through the same public provider API the app's own code would call; the call path does not bifurcate. But the semantic concretion depends on the actor: a component's existing events declare delegate capacity (polymorphic events, allowedFamilies in the morfo contract) and the provider concretes { family: 'delegate', verb: 'act' } when the actor is the agent. So the same insert() fires the same runtime.trigger, but sema, eidos and the a11y layer read the correct answer to "¿quién actúa?". If a capability needs a door the public API lacks, that door is promoted to the component through the normal route — never bypassed.

  4. Total degradation. With the axis absent, every component works at 100%. Manifest registration is a no-op; nothing a component owns depends on the agent existing. See §4.

A corollary the field taught us the hard way: client-side budgets are UX courtesy, not a security boundary. The boundary lives where the credential lives (a proxy/backend). The motor's budgets exist to drive the state machine and the UX, and the doctrine says so honestly.


2. The domain model

The axis is built around a delegation, not a conversation. A chat is one surface that chains delegations; it is not the core (an agent can fill a document with no conversation at all).

ActiveAgent  — the session (long-lived; context across runs; identity)
 └── Run      — a delegation: budget, authorization, undo grouping, control return
      ├── Turn          — one exchange with the model
      │    └── CapabilityCall — a delegate.act on a provider (the only thing that mutates)
      └── Context       — snapshot of reads at turn open (§3)

Three readings this ladder settles:

  • The CapabilityCall is what mutates the world; the Run is the unit of delegation. A mutation without the delegation envelope (budget, authorization, undo grouping, control return) is exactly what the doctrine forbids.
  • The degenerate Run — a single auto-authorized call — is the non-AI case of the book (a macro, a rule, a workflow). It is a valid, cheap Run.
  • Initiative is modeled, not just actor. Every Run carries initiative: 'user' | 'system' | 'scheduled' — why it acts, not only who — because it changes trust, announcement, audit and the autonomy ceiling. The entry verb hints it: a cycle entered by offer is system-initiative.

Ambient mode (always-on: autocomplete, predictive suggestions) is out of the axis in v1, reserved by name: a Subscription/watcher emitting delegate.offer, with coalesced micro-runs as future design. It is named so a second, parallel agency system cannot grow in by the back door.

The run state machine = the delegate verbs, literally

The states are the verbs — as custody of control, not as cognitive phases. The model's own plan→tool→plan loop lives entirely inside acting (across turns); the machine models who holds control, not how the agent thinks.

                    ┌─────────── escalated ──────────┐   (non-terminal; timeout → returned)
                    ▼                                 │
offered → planned → reviewing ⇄ authorized → acting ──┤
   │         │      (re-entrant:  (fixes the    │     ▼
   │         │       partial       allowlist +  │  returned  ← control ALWAYS returns
   │         │       approvals)    budget)      │  (outcome-typed)
   └─────────┴──── cancelled ────────────────────────┘
  • Authorization fixes scope, not vibes. authorized derives from the plan a capability allowlist + budget; an act outside that scope ⇒ automatic escalate to re-review (never a silent failure, never teatro). This makes reviewing re-entrant from acting.
  • returned is outcome-typed: completed | partial | rolled-back | aborted | error, with a reason (superseded — the premise changed under the run; budget-exhausted; killed). Mid-sequence failure rolls back to run-start by default (the user authorized the whole plan, not half); a capability may declare it accepts partial close.
  • escalated is non-terminal, carries a typed reason (elicitation | insufficient-scope | budget-exhausted | out-of-plan), and auto-returns on timeout (an abandoned escalation still needs an exit).
  • Pause needs no new verb: pause = return with resumable context; resume = a fresh Run chained in the session — which coincides with AG-UI's resume mechanics. Our semantics and its wire converge without trying to.
  • User interruption = immediate return, never an error. A global kill switch on the motor (disable()) sends every run to returned(aborted).

3. The participation contract

A component participates by declaring a manifest — the mirror of a sema pack or a langs catalog, in its own tree (src/uix/agent/components/{x}.ts). The manifest satisfies structural contracts defined by the art (AgentCapability / AgentRead: string ids, Standard-Schema args via sium, opaque handler) — the MotionDom / SceneDom port pattern — by typing against the provider's public API on the uix side. Soma/uix types never cross into the art; the opaque handler grants location-transparency for free (a reserved path to distributed delegation).

  • Model-facing projection. TypeScript never reaches the model. Every read and act carries a natural-language description; args project as JSON Schema (a sium→JSON-Schema emitter), with size limits per arg as part of the contract. Traits annotate each act for the reasoner (cost, latency, stability, staleness) — the MCP tool-annotations pattern.
  • Instances and context — push→pull. An always-present index (stable id
    • type + human label + one-line summary, hierarchical) plus deep reads as a pull capability (read(instanceId)), prioritized by focus/viewport; addressing ambiguity ("the editor", with two mounted) ⇒ elicitation, not a heuristic. available is a reactive predicate per act; a capabilities-changed event refreshes the toolset between turns.
  • Results and failures are sium-typed; a result carrying third-party content is as untrusted as an imported document (same origin doctrine, §5). unknown-capability (a hallucinated tool) returns as a tool-result for self-correction, bounded by budget.
  • Altitude, not RPC. Capabilities are mechanical world-mutations, few and at the right altitude (replaceRange, not twelve micro-inserts). Semantic operations (summarize, translate) are the model's cognition or a reserved agent-side "skills" tier (the MCP-prompts analogue) — never lodged in the component, which would invert the dependency direction.
  • Maturity tiers. A manifest declares its tier — reads-only → reversible acts → full acts — so the ~110 canon components adopt incrementally without gating the axis; the guard audits the declared tier.
  • Reversibility is by effect, not by capability. External effects (send / publish / notify) are irreversible by definition; the reversion authority is the component's native history, never orca compensation.
  • OCC. Each read exports a _version; each act must carry it; a mismatch is a recoverable StateStaleError (re-read → re-plan) — the answer to TOCTOU (§5).

4. Minimum contracts (the §0 row) and degradation

Following the framework's rule that every change derive from the module contract table:

Module   Requires            Optional                 When missing                         Error
agent    AgentTransport port orca, perm, prefs,        manifest registration is a NO-OP;    only on invoking a
         (app supplies it)   connection, storage       every component stays 100% functional capability with no transport
  • The agent is an app-level art (factory defineActiveAgent(...), the auth/perm/cache pattern), not something ActiveUix builds in standalone. Promotion to a read-only uix.agent accessor in attach mode arrives with the canonical consumer (the scene-D4 precedent).
  • A component with a manifest MUST behave identically with the agent absent. The manifest is inert data until a running ActiveAgent reads it.
  • Transport is app-land by default (the chat-family precedent: "the transport stays app-land"); a server-authoritative svrs/agent is a conditioned future initiative, never the v1 assumption. Adapters live in $agent/adapters/* outside the main barrel; a deterministic ScriptedAgentTransport is first-class, so static demos and component tests run with real instances, no mocks.

5. The actor primitive (⚖️2) — causal attribution

The load-bearing decision of the axis. When the agent's act on component A triggers events that cascade to mutate component B, B's emission must carry the correct actor — or sema announces the agent's secondary effect as "user", a11y and audit-trail catastrophe.

The answer is the OTel-baggage unification, corrected by the browser caveat. The actor is an opaque, engine-minted token (ActorToken, minted and resolved only by EngineAgent; a non-token USER sentinel everywhere else, so absence can never be confused with a minted token) carried as one reserved slot on the causal envelope the framework already propagates — OrcaEventMeta on the orca path, the bus envelope's context bag on the bus path — inherited depth-to-depth exactly as traceId is. It is read synchronously via actor?: ActorToken on TriggerOptions → SemanticSignal, before the first await in EngineSemantic.emit.

This shares one propagation primitive with the D-AG.9 trace port (run = trace, CapabilityCall = span, actor = baggage) — distinct payloads, distinct guarantees: the span export is async and loss-tolerant; the actor read is synchronous and authoritative. Sharing the primitive avoids forking agency; keeping the payloads distinct respects that sema cannot tolerate the fall-to-root a tracer tolerates.

Why explicit, not ambient: native AsyncContext/ALS is the framework's own orca mechanism, already adjudicated — in the browser (als.ts feature-detects to null) it degrades to a root envelope with no actor, precisely the failure this decision exists to avoid. A faithful zero-dep reimplementation is impossible without Zone.js-style global monkey-patching (forbidden, and it still cannot cover a bare await). So ambient auto-fill is opportunistic sugar; the explicit envelope is the guarantee — which is exactly how OpenTelemetry-JS actually works (baggage on a propagated context, synchronous StackContextManager in browsers, async_hooks server-side only). This is the industry consensus in one line: explicit provenance bound to the message, ambient as sugar — OTel propagators inject/extract explicitly, CaMeL binds an unforgeable capability tag to the value, Yjs attributes derived mutations by an explicit per-transaction origin.

Scope and honesty:

  • Anti-forgery (the reason the token is opaque): no site outside arts/agent may construct an ActorToken or write the reserved actor key; the token is never wire-serialized (the trace-export boundary emits a non-authoritative 'agent' label only). The ActorToken type lives in a leaf below both orca and agent to avoid an orca→agent cycle; minting authority stays in EngineAgent. On the bus path the actor rides the untyped context bag, so its anti-forgery defense is the agent-check lint, not the type system.
  • The guarantee is scoped. It is delivered for the explicitly-authored A→B bridge that threads opts.actor. An agent driving a stock component through its public API has no call-site to thread it and, in the browser, resolves to the USER sentinel — best-effort, out of v1. In-scope attribution rides the model-derived data-actor paint (eidos reads the reactive run state during acting), which needs no threading; the envelope baggage is specifically for out-of-scope secondary effects + audit.
  • The primitive ships first, the propagation retrofit later. The minimal primitive (branded token + USER sentinel + private registry + anti-forgery guard + synchronous opts.actor read) is built up front. The ~6-seam envelope-propagation retrofit (OrcaEventMeta.actor, OrcaALSContext.actor, ctx.emit forwarding, buildChildEnvelope copy, bus-subscription forwarding, publishCausedBy parent-merge) is deferred to the phase with a real out-of-scope async-cascade consumer, gated by the guard's drift check that ambient is never the sole carrier on a cross-component bridge.

Three levels — suggest | review | auto — resolved prefs-style (the app is the ceiling; the user adjusts downward; a capability may demand more control, never less; deny-overrides). Irreversible ⇒ review always.

  • Dynamic risk by payload: the manifest declares a risk classifier over args; auto escalates to review per invocation (same capability, different blast radius).
  • Context origin lowers the ceiling — the unification that dissolves prompt-injection from an open threat into a policy row: each read carries an origin tag (authored-by-user | imported | third-party | other-participant); any untrusted origin in the turn lowers the ceiling (to review, or suggest for sensitive capabilities). This is the industry pattern (forced human approval when untrusted content is in context) expressed through machinery already designed.
  • Initiative affects ceilings: autonomous initiative (system / scheduled) defaults stricter than user-requested.
  • Authorization is bound to state, not prose (TOCTOU): review computes against a fingerprint F of the relevant reads; authorization binds to F plus a TTL; drift ⇒ recompute the diff (identical → optional auto-re-approve; different → re-review). The user↔agent concurrency problem is the OT/CRDT model (Yjs) — adopt the mental model, not the library.
  • Approval fatigue: persistent grants ("always allow X") carry a TTL and an explicit scope (session / component / act type); review shows the concrete proposed change in the app's own domain, never prose (amended 2026-07-22 — see "The review surface" in §7: the content of a review is app-owned). A capabilities-changed event re-gates grants.
  • perm governs only server-touching scopes (enforced in the proxy / server-side tool-loop); the base case is local consent with no perm dependency. reads are permission subjects symmetric to acts — sensitive components (credentials/passwords) never register reads.

7. Semantics, presence and accessibility

  • The motor emits nothing (an art with no DOM); it exposes the run as reactive state. Surfaces declare the cycle in their morfos (delegate-*); sustain-processing belongs to Aura (state-bound persistence, cleared on run close). Actuated components concrete delegate.act via allowedFamilies with the actor context.
  • A run above suggest requires a mounted presence surface (a minimal Aura). A run in acting with nowhere to emit its cycle is perceptually invisible — a direct contradiction of "¿quién actúa?". The motor refuses (or warns) on entering acting without one; this is mechanically verifiable.
  • Write-scope is declared: a run declares its scope (component | part | range); hold and paint apply only there.
  • The review surface (amended 2026-07-22 — supersedes the discarded "review-card" idea, a dev-tool anchoring). The framework/app split is: the FRAMEWORK owns the cycle states expressed perceptually (reviewing / escalated surfaced by the presence surface, with authorize/reject/cancel affordances — Aura already does this), the "proposed by another actor" semantics, and the attributed announcements. The APP owns the CONTENT of what is being reviewed, rendered in its own domain (a document shows its proposed blocks, a calendar its pending events, a form its suggested values) — in context, where the user already is. A git-style diff card is a legitimate construct for a specific application (e.g. a code tool built ON the framework); it is never canon.

The a11y contract — no framework has this; documenting it fixes the pattern others will cite:

  1. One polite live region owned by Aura; announcements attributed ("Agente: …") and coalesced by plan/batch — never per delta, never per act (WCAG 4.1.3).
  2. No focus-theft — a motor invariant (WCAG 3.2.1 / 3.2.2), not each surface's courtesy.
  3. aria-busy on the region being mutated during streaming; a completion announcement on close that carries the outcome — a failure returning control silently is the same invisibility as an act with nowhere to emit. The surface expresses it with THREE terminal events, intent intrinsic to each (a dynamic intent has nowhere to live: the trigger options carry none and a morfo's fromProp binds to a public prop, not to machine state): completed → fulfill · aborted/partial → no intent (a decline is absence, not loss) · error/rolled-back → loss. All three stay TRANSIENT: once control is back there is no custody to claim, so a persistent alarm would make the surface lie about who holds it — the failure's DETAIL belongs to the app, in its own domain. Binding a tone to this structural family is the sanctioned BK-FRAME-NO-INTENT exception (book-updates A-1) and costs a written intentRationale: the engine emits nothing and no signal/commit event exists on the presence surface, so the consolidated return is the outcome's only expressive site.
  4. Elicitations are dismissible / postponable, not modal by default (WCAG 2.2.4). An elicitation is a typed question, so its surface is a form in context — not a conversation. The schema the agent asks with is the same projection its capability args already travel through (sium → JSON Schema), read in the opposite direction; binding elicitation to a chat would make a general contract depend on one product shape.
  5. A global keyboard exit (an Esc-like doctrine): cancel / return reachable without a mouse. Aura uses reduce: 'static-frame', never hide.

8. Threat model

The axis builds, by construction, the classic "lethal trifecta" (Willison): private data + untrusted content + the ability to act. The doctrine treats it head-on:

  • Prompt injection. Reads carry an origin tag (per §6); untrusted origin lowers the autonomy ceiling (the primary mitigation, folded into policy). Adapters spotlight (delimit/mark) non-authored content so the model treats it as data. Reserved for depth: dual-LLM (a privileged planner that never sees untrusted content; a quarantined model that returns only symbolic references) and CaMeL-style capability/data-flow tracking — philosophically close to the framework's structural ports. Checklist anchors: OpenAI instruction-hierarchy, OWASP LLM Top 10 (LLM01).
  • TOCTOU. Authorization bound to a read fingerprint + TTL (§6).
  • Actor forgery. The opaque engine-minted token + the agent-check guard (§5).
  • Run recovery. A minimal write-ahead journal per run via $storage (sessionStorage suffices in v1) records authorizations + applied acts and detects orphaned runs on reload. The audit-trail is defined now as an OTel-shaped trace port (run = trace, CapabilityCall = span); the storage vehicle arrives later. (orca cannot be a store — its own rule.)
  • Kill switch, multi-user, i18n. A global disable (§2). User identity in the actor vocabulary is reserved (multi-user arrives with connection). The agent's identity / system-context carries the effective locale; the agent announces in the langs-axis language like any surface.

9. What this axis is NOT

  • Not a chat. Chat is one surface (a participant + a composer) that chains runs; the axis works with no conversation at all.
  • Not an LLM integration. It is a theory of delegation; a model is one kind of actor. Nothing depends on a provider SDK (adapters are isolated and optional).
  • Not a parallel "AI mode". It preserves every framework invariant — dependency direction, semantics, events, public API, undo, a11y — instead of forking them. A component with a manifest is the same component.
  • Not a component that knows components. It sees capabilities, never types.

10. To go deeper

Powered by TurnKey Linux.