22 KiB
| title | type | audience | authority | version | status | doctrine | plan |
|---|---|---|---|---|---|---|---|
| The Delegation Contract — activeUIX agent axis | specification | human + agent + implementors | NORMATIVE — this document defines conformance for the delegation axis | 2026-07-28 | DRAFT | docs/architecture/agent.md (the WHY; this document is the WHAT) | docs/process/PLAN-agent.md (§5 · D-AG.12 — the decision that created this document) |
The Delegation Contract
An agent is another actor. This document specifies what an implementation must do so that a delegation — a bounded transfer of control from a person to another actor — is observable, attributable, bounded and reversible.
It is deliberately not an LLM specification. Nothing here depends on a model, a provider or a prompt. A macro, an automation rule, a workflow and a language model are all another actor, and they conform the same way.
0. About this document
0.1 Conformance language
The key words MUST, MUST NOT, REQUIRED, SHALL, SHALL NOT, SHOULD, SHOULD NOT, RECOMMENDED, MAY and OPTIONAL are to be interpreted as described in RFC 2119 / RFC 8174, when and only when they appear in all capitals.
0.2 Requirement identifiers
Every normative statement carries a stable identifier of the form AG-n.
Identifiers are never reused and never renumbered. A requirement that is
withdrawn keeps its number and is marked Withdrawn; the number is not
recycled. Conformance failures — from either vehicle in §0.5 — MUST cite
the identifier they violate.
0.3 Stability ladder
Each section declares a stability level. The level tells an implementor what they may build against:
| Level | Meaning |
|---|---|
stable |
Will not change incompatibly within a major version. |
provisional |
Shipped and testable, but may change while this document is DRAFT. |
reserved |
The shape is declared and the hole is typed; behaviour is NOT yet specified. Implementations MUST NOT repurpose reserved identifiers or wire categories for other meanings. |
0.4 Versioning
This document is versioned by date (YYYY-MM-DD). The date changes when a
normative statement changes incompatibly; editorial changes do not move it.
While status: DRAFT, any section may change.
0.5 Conformance vehicles
Conformance is verifiable by two distinct vehicles, because two distinct things are being verified:
| Vehicle | Verifies | Nature |
|---|---|---|
| Protocol fixture kit | The wire and the custody machine | Portable — a third party runs it against their implementation. Scripted turns in, expected custody transitions out. |
agent-check |
The participation contract of a component | Repo lint over the manifest tree. |
Both MUST cite requirement identifiers in their failures. A vehicle that reports "failed" without naming what it failed is not a conformance vehicle.
0.6 Honest position
Stated here rather than discovered later:
- Capability calls are sequential in v1; concurrency is not specified.
- The axis is client-first. A server-authoritative implementation is permitted and its portability is required by §7, but it is not the assumed topology.
- Generality is under test. At the version above, the contract has been exercised by one presence surface and is being proven against two domain shapes (schema and set). Until §A.3's gate is met, the generality of the manifest format is a claim, not a result.
1. Model and terminology
Stability: stable
Session — long-lived; identity; context across delegations
└── Run — ONE delegation: budget · authorization · undo group · control return
├── Turn — one exchange with the acting party
│ └── CapabilityCall — the ONLY thing that mutates the world
└── Context — snapshot of reads taken when the turn opens
- Run — a bounded transfer of control. The unit of delegation, of undo and of audit.
- CapabilityCall — an invocation of a declared capability. The only construct permitted to mutate anything.
- Read — a declared, versioned, origin-tagged view a participant exposes.
- Manifest — the declaration a component makes to participate: its capabilities and its reads.
- Initiative — why a run exists:
user | system | scheduled. - Autonomy — how far the run may go without a human:
suggest | review | auto.
AG-1 — Every mutation performed by a non-user actor MUST occur inside a CapabilityCall that belongs to a Run. A mutation without the delegation envelope (budget, authorization, undo grouping, control return) is non-conformant.
AG-2 — A Run MUST carry an
initiative.
AG-3 — A single auto-authorized CapabilityCall MUST be expressible as a valid Run. (The non-model case — a macro, a rule — is not a second system.)
2. Custody of control
Stability: stable
The run states model who holds control, not how the acting party thinks.
Any internal plan/act/replan loop lives entirely inside acting.
┌─────────── escalated ──────────┐ (non-terminal)
▼ │
offered → planned → reviewing ⇄ authorized → acting ──┤
│ │ (re-entrant) │ ▼
└─────────┴──── cancelled ───────────────► returned (control ALWAYS returns)
AG-4 — Control MUST always return. Every Run MUST reach a terminal state (
returnedorcancelled) — including on failure, on budget exhaustion, on timeout and on kill. A Run that can end without returning control is non-conformant.
AG-5 —
returnedMUST carry a typed outcome from at least{completed, partial, rolled-back, aborted, error}, and SHOULD carry a typed reason.
AG-6 —
escalatedMUST be non-terminal, MUST carry a typed reason, and MUST auto-return on timeout. An abandoned escalation still needs an exit.
AG-7 — Authorization MUST derive a capability allowlist and a budget. A CapabilityCall outside that scope MUST escalate for re-review. It MUST NOT fail silently and MUST NOT proceed.
AG-8 — A mid-sequence failure MUST roll back to run-start, unless the capability declares that it accepts partial close. The person authorized the plan, not half of it.
AG-9 — User interruption MUST be modeled as a return, never as an error.
AG-10 — The implementation MUST provide a global kill that sends every active Run to
returned(aborted).
AG-11 — Pause and resume MUST NOT introduce new states: a pause is a return with resumable context; a resume is a fresh Run chained in the session.
3. The participation contract
Stability: provisional — the manifest format is being proven against two
domain shapes; see §A.3.
AG-12 — A participant declares capabilities and reads in a manifest. The manifest MUST be inert data: it MUST NOT have any effect until a running delegation reads it.
AG-13 — A component carrying a manifest MUST behave identically when no agent is present. Conformance to this requirement is verifiable by running the component's own test suite with no agent configured.
AG-14 — Every capability and every read MUST carry a natural-language description intended for the acting party. Implementation types MUST NOT be assumed to reach it.
AG-15 — Capability arguments MUST be typed by a schema and MUST be projectable to JSON Schema for the acting party.
AG-16 — The contract MUST declare a size limit per argument. Exceeding it MUST produce a typed invalid-arguments failure, not a truncation.
AG-17 — A manifest MUST declare a maturity tier from at least
{reads-only, reversible-acts, full-acts}. The declared tier MUST be auditable against what the manifest actually exposes.
AG-18 — Reversibility MUST be declared per effect, not per capability. External effects (send, publish, notify) MUST be treated as irreversible.
AG-19 — The reversion authority MUST be the participant's own native history. Compensating transactions MUST NOT be presented as reversal.
AG-20 — Each read MUST export a version. Each act MUST carry the version it was planned against. A mismatch MUST be a recoverable stale error that permits re-read and re-plan — not a run failure.
AG-21 — An invocation of an undeclared capability MUST be returned to the acting party as a result it can correct from, bounded by budget. It MUST NOT terminate the Run by itself.
AG-22 — Capabilities MUST be mechanical mutations at the participant's own altitude. Cognitive operations (summarize, translate, classify) MUST NOT be lodged in the participant.
AG-23 — When addressing is ambiguous (two instances of the same participant are mounted), the implementation MUST elicit. It MUST NOT resolve by heuristic.
4. Authorization, consent and autonomy
Stability: stable
AG-24 — Autonomy MUST resolve as: the application sets the ceiling; the person may only lower it; a capability may demand more control, never less. Conflicts MUST resolve deny-overrides.
AG-25 — A capability with an irreversible effect MUST require at least
review, regardless of the resolved ceiling.
AG-26 — Non-user initiative (
system,scheduled) MUST NOT run atauto; it MUST be clamped to at mostreview.
AG-27 — Every read MUST carry an origin tag distinguishing at least author-produced from imported or third-party content.
AG-28 — An untrusted origin present in the turn's context MUST lower the autonomy ceiling. This is the specified answer to prompt injection: it is a policy row, not an open threat.
AG-29 — Authorization MUST bind to a fingerprint of the reads it was computed against, plus a time to live. Drift MUST trigger recomputation: an identical diff MAY auto-re-approve; a different one MUST re-review.
AG-30 — A persistent grant ("always allow X") MUST carry a time to live and an explicit scope, and MUST be re-gated when the capability set changes.
AG-31 — A review surface MUST present the concrete proposed change in the application's own domain. Prose descriptions of a change MUST NOT substitute for it.
5. Actor attribution
Stability: stable
AG-32 — Every CapabilityCall MUST carry an actor reference identifying who caused it. The absence of a reference MUST mean the person, never an unattributed actor.
AG-33 — Only the delegation engine MUST be able to produce an actor reference that resolves. A structurally forged reference MUST resolve to nothing.
AG-34 — Code outside the engine MUST carry a received actor reference verbatim. It MUST NOT construct actor contexts. (Verified by
agent-check; the type system alone cannot enforce this, because a cast type-checks.)
6. Presence and accessibility
Stability: provisional — AG-35 carries an open decision; see §D.
This section is the axis's distinguishing contract. No surveyed framework specifies it, and it is the reason the delegation is observable rather than merely logged.
AG-35 — A Run above
suggestautonomy SHOULD have an attached presence — some surface that is expressing its cycle. An implementation MUST make this attachment observable, so that a Run enteringactingwith no presence attached is detectable rather than silent.A presence is not a specific component. Any surface that renders the cycle satisfies this requirement by attaching; the requirement is that something is observing, not that a particular widget is mounted.
⚠ OPEN — see §D.1. Whether an implementation MUST refuse such a Run, or SHOULD warn, is not yet decided.
AG-36 — The axis MUST own exactly one polite live region. Multiple concurrent announcement channels are non-conformant.
AG-37 — Announcements MUST be attributed (they say who acted) and MUST be coalesced by plan. They MUST NOT be emitted per act or per delta.
AG-38 — The engine MUST NOT move focus. This is an engine invariant, not a courtesy of each surface.
AG-39 — The announcement that closes a Run MUST carry its outcome. A failure that returns control silently is indistinguishable from success and is non-conformant.
AG-40 — A region being mutated by the acting party MUST be marked busy while the mutation is in flight.
AG-41 — Elicitations MUST be dismissible or postponable, and MUST NOT be modal by default.
⚠ NOT YET MET — see §D.2. The channel that would carry an elicitation's question and its answer does not exist yet.
AG-42 — A keyboard-reachable exit from the delegation MUST exist at all times while a Run is active.
AG-43 — Under reduced-motion preferences, presence MUST be preserved in a static form. It MUST NOT be hidden.
AG-44 — No state of the cycle MUST be distinguishable by colour alone.
7. Wire protocol
Stability: stable for the active categories · reserved where marked
AG-45 — A transport MUST declare the protocol version it speaks.
AG-46 — The protocol MUST carry at least: turn lifecycle, text deltas, tool-call lifecycle, and an escape hatch for implementation-specific events.
AG-47 — State synchronisation is reserved. Its events carry a typed direction, because context flows from the interface to the acting party while agent state flows the other way; an implementation MUST NOT collapse the two directions into one channel.
AG-48 — Elicitation is reserved at the wire level, with its shape declared (a request identifier, a prompt, and an optional schema). Implementations MUST NOT repurpose it.
AG-49 — A transport is replaceable. An implementation MUST NOT make the delegation machine depend on a specific provider dialect; dialect translation belongs in an adapter.
8. Degradation
Stability: stable
| Requires | Optional | When missing | Error |
|---|---|---|---|
| A transport port (the application supplies it) | orchestration, permissions, preferences, connection, storage | Manifest registration is a no-op; every participant stays fully functional | Only on invoking a capability with no transport |
AG-50 — With no transport configured, manifest registration MUST be a no-op and every participant MUST remain fully functional. An error MUST be raised only when a capability is actually invoked.
Annex A — Frontier claims
These are the two claims this contract makes beyond the state of the field. Each carries an explicit status, because a specification that claims what it has not shipped is not a reference.
A.1 — Context by attention · claimed
Reads are not a flat bag handed wholesale to the acting party. They are an always-present index (identity, type, label, summary) plus deep reads pulled on demand, prioritized by what the person is attending to — focus and viewport. This is a signal a user-interface framework has and a server-side agent framework does not.
Status: claimed. The push→pull index is specified (§3) and not yet
implemented.
A.2 — Delegation as a transaction · claimed
One delegation is one undo step. The Run is the transaction envelope: budget, authorization, optimistic concurrency per read, write scope, and supersession when the premise changes underneath.
Status: claimed. Requirements AG-8, AG-19 and AG-20 specify it; it has
not yet been exercised against a set-shaped participant, which is where it
either holds or breaks.
A.3 — Promotion gate
This document leaves DRAFT and claims reference status when, and only when:
- ≥2 distinct domain shapes have conformant manifests — at minimum one schema-shaped and one set-shaped participant. Two text-shaped participants do not satisfy this gate: a manifest format proven only over sequential text silently assumes sequential text.
- Both frontier claims above are
shipped. - The protocol fixture kit passes.
Annex B — Threat model
Stability: provisional · informative
The axis constructs, by definition, the classic lethal trifecta: private data, untrusted content, and the ability to act. The contract's answer is distributed rather than centralized:
| Threat | Where the contract answers it |
|---|---|
| Prompt injection via context | AG-27 · AG-28 (origin lowers the ceiling) |
| Time-of-check/time-of-use | AG-20 (versioned reads) · AG-29 (fingerprint + TTL) |
| Actor impersonation | AG-33 · AG-34 |
| Silent action | AG-35 · AG-39 · AG-4 |
| Scope creep mid-run | AG-7 |
| Approval fatigue | AG-30 |
| Irreversible surprise | AG-18 · AG-25 |
Annex C — Conformance checklists
C.1 — For an implementation of the axis
Shipped as $agent/conformance (outside the main barrel):
import { runConformance, formatConformanceReport } from '$agent/conformance';
const report = await runConformance(mySubjectFactory);
if (report.failed) throw new Error(formatConformanceReport(report));
An implementor adapts ConformanceSubject — start · run · authorize · reject ·
resolveEscalation · stop · disable — over their own engine. The kit never
imports ours; subject.ts is the worked example of that adapter, not a
dependency. Every failure line leads with the AG-n it violates.
Covered today: AG-2, AG-4, AG-5, AG-6, AG-9, AG-10, AG-21,
AG-25, AG-26, AG-50. AG-35 and AG-41 are deliberately absent while
§D holds their decisions open — a kit that asserted an undecided requirement
would be inventing the decision.
C.2 — For a participant (a component)
agent-check; every failure cites its AG-n.
§D — Open decisions
Listed here rather than silently resolved. While this document is DRAFT, an
open decision is a legitimate state; on promotion, §D MUST be empty.
D.1 — Presence: SHOULD or MUST? (AG-35)
The doctrine states the engine "refuses (or warns)". The parenthetical is unresolved.
- Refusing gives the axis a hard guarantee: a delegation can never act invisibly. It also breaks every headless and test usage that does not attach a presence, and removes the application's right to choose its own surface.
- Warning preserves both, and makes the violation discoverable in development — but a warning enforces nothing in production.
Recommendation: SHOULD by default, with an application-level policy that
promotes it to MUST for applications that want the hard guarantee. This is
expressible in RFC-2119 without weakening the claim, and it is the only shape
that does not punish an application for rendering the cycle its own way.
Blocks: the implementation of AG-35.
D.2 — The elicitation channel (AG-41)
The requirement is specified and not met, for a reason that is structural rather than cosmetic: an escalation carries a typed reason but not a question, and the resolution channel is binary — it has nowhere to carry an answer.
Materializing AG-41 requires both directions:
- the question travelling out (prompt + schema, the shape AG-48 already reserves), and
- a typed answer travelling back in.
Blocks: AG-41, and with it the completeness of §6.
To go deeper
- Why any of this:
docs/architecture/agent.md. - The execution record and signed decisions:
docs/process/PLAN-agent.md. - The semantic vocabulary the cycle is expressed in:
docs/CANON.md.