Skip to content

Composable Runtime Architecture

This chapter explains what actually happens on the platform side when the host system sends a message with createMessage (POST /conversations/{conversation_id}/messages). The short version: the platform assembles a runtime environment on demand — the right repository, the right skills, the right storage, the right agent runtime — runs the agent inside an isolated sandbox that holds zero credentials, streams the result back as the native agent event stream, and tears the environment down. Nothing about that environment is baked in; everything is composed, per request, from data the Integration API already manages.

Understanding this model matters for integrators because every knob on the API — the tenant repository assignment, skill narrowing, runtime.agent_type, warm sandboxes, secrets at four scopes — is a handle on one stage of this composition pipeline. This chapter maps the knobs to the machinery.

1. The composable-runtime thesis: capabilities are files

Section titled “1. The composable-runtime thesis: capabilities are files”

The platform’s core design commitment is that an agent’s capabilities are files in a git repository — not rows in a proprietary capability database, not code compiled into the platform. A skill is a folder with a SKILL.md manifest (frontmatter + instructions + optional scripts and resources). Agent instructions, personas, and tool configurations follow the same convention-defined, file-based format.

The platform’s job is not to define capabilities — it is to compose them, per request, into a concrete runtime environment:

Repository registry → which capability trees exist (registerRepository)
Tenant assignment → which tree applies to THIS tenant (one per tenant)
Skill-access intersection → which skills from that tree this run may use
Runtime selection → which agent harness executes them (runtime.agent_type)
Sandbox provisioning → an isolated environment with exactly that composition

Repositories are top-level registry entries, registered once with registerRepository (POST /repositories) together with a pre-authenticated credential (createCredential, POST /credentials — the git token is write-only and never readable again). A repository dictates which skills exist: listRepositorySkills (GET /repositories/{repository_id}/skills) enumerates them, syncRepository (POST /repositories/{repository_id}/sync) re-scans the tree after a push, and createRepositorySkill (POST /repositories/{repository_id}/skills) authors a new skill directly into the tree — committed to git like any other change.

A registered repository reaches runs through the tenant assignment — each tenant carries exactly one:

flowchart LR
    REG["Repository registry<br/><i>registerRepository</i>"] -->|"assignTenantRepository<br/>PUT /tenants/{id}/repository"| T["Tenant<br/><i>ONE assigned repository</i>"]
    T -.->|"inherited by child tenants<br/>with no assignment of their own"| TC["Child tenants"]
    T --> RUNS["Every role, conversation,<br/>and message in the tenant"]
    style REG fill:#dbeafe,stroke:#1d4ed8
    style T fill:#dcfce7,stroke:#15803d

assignTenantRepository (PUT /tenants/{tenant_id}/repository, body {repository_id, branch_override?}) sets it; getTenantRepository reads the effective assignment — when the tenant has none of its own, the nearest ancestor’s assignment applies (inherited: true), so a deployment can assign one capability tree high in the tenant hierarchy and serve every tenant below it. branch_override reads a different branch of the same repository for one tenant — the standard mechanism for piloting a capability change on a single tenant before merging it to the shared branch.

There are no repository overrides below the tenant: roles, users, conversations, and messages all resolve to the tenant’s repository. What varies below the tenant is only which of its skills a run may use.

The skills that actually enter a run are an intersection of four filters — each layer can only narrow, never widen:

effective_skills =
skills of the tenant's repository (what exists)
∩ role.skill_access (what the user's role permits)
∩ conversation.selected_skill_ids (what this thread opted into)
∩ message.skill_ids (what this turn narrowed to)
LayerSet viaTypical use
Role skill_accesscreateRole / updateRole{mode: "all"} or {mode: "selected", skill_ids: […]}The job-function grant: a CSR role limited to two skills, a dispatcher role granted everything
Conversation selected_skill_idscreateConversation / updateConversationA thread that should work with a focused subset for its whole lifetime
Message skill_idscreateMessageThe per-turn focus knob: one run with exactly the skills it needs, nothing stored

Integrators can inspect the resolution at every level without running anything: listUserSkills (GET /users/{user_id}/skills) returns a user’s effective skills with provenance (via {role_id, repository_id}), and listRoleSkills (GET /roles/{role_id}/skills) resolves a single role.

  • Versioned, auditable capability. Every capability change is a git commit — diffable, revertable, attributable. “What could the agent do on March 3rd?” is a git log question, not a forensic reconstruction.
  • Per-tenant customization without platform code. A new tenant vertical, a new playbook, a reworked skill set — all of it is repository content. The platform binary never changes; a new capability profile is a new repo (or branch) plus one assignTenantRepository call.
  • Instant capability updates. Push to the repository, call syncRepository, and the next message runs with the updated tree. No deployment, no restart, no migration.
  • Least-capability by construction. Because every layer intersects downward, the narrowest intent wins. A message that names two skills runs with at most two skills — the LLM’s context contains nothing else, which is both a security property and a quality property (see next section).
  • One answer to “which repository?”. The tenant assignment makes capability provenance trivially auditable: every run in a tenant, at any level, resolved to the same tree at a known branch.

2. The platform controls the run — not just an API wrapper

Section titled “2. The platform controls the run — not just an API wrapper”

A common integration failure mode is treating an agent platform as a thin proxy in front of an LLM API: prompt in, tokens out, everything else is the caller’s problem. This platform is the opposite: it owns the LLM run end-to-end — what enters the model’s context, what the model is allowed to do mid-run, what the caller sees while it runs, and what happens when the model needs something it doesn’t have.

ConcernThin API wrapperThis platform
Context contentsCaller concatenates strings and hopesComposed from the tenant’s repository + effective skills + conversation history + workspace posture; nothing ungranted is even visible to the agent
Capability scopePrompt-level pleading (“do not use tool X”)Structural: ungranted skills are absent from the sandbox filesystem — the agent cannot list, read, or invoke them
Latency UXDead air until first tokenFiller: an optional low-latency filler agent opens the reply while the full run warms
SecretsIn the prompt or env, visible to the modelNever in the run at all — the agent sees aliases; the egress proxy resolves them on the wire. An unvaulted alias fails the run fast with missing-secret
OutputAn unstructured token streamThe native agent event stream: NDJSON lines carrying deltas, complete turns with tool calls and results, and a mandatory terminal result

The individual mechanisms:

  • Skill narrowing keeps context lean. Skill instructions consume context-window budget, and irrelevant instructions measurably degrade model behavior. Because the effective skill set is an intersection (§1), a host that knows a turn only needs invoice-lookup can say so on createMessage via skill_ids — and the run’s context contains that skill and nothing else. Narrowing is simultaneously a quality control (focused context), a cost control (fewer tokens), and a security control (smaller action surface).
  • Context composition is platform-owned. The platform assembles what enters the agent loop: the tenant’s repository tree, the filtered skills, the conversation history, per-message env parameters, and a system-prompt description of the workspace (which paths are writable, what each storage zone is for). The host system supplies intent; the platform supplies a provably-scoped environment.
  • Filler for latency UX. Provisioning an isolated environment takes real milliseconds. When filler is enabled (cascading tenant settings → conversation → message, most specific wins), a low-latency filler agent opens the reply — a natural spoken-style opener — while the full run warms behind it. Filler output is part of the reply and the persisted message, indistinguishable from agent output by design; the cascade controls only whether it runs. See Streaming Contract.
  • Missing secrets fail fast. Mid-run, the agent may reference a credential by alias that no scope has vaulted. The run does not stall waiting for anyone: it terminates immediately with a missing-secret problem naming the alias, and the host re-sends the message with the secret attached — a one-step retry that keeps credential delivery need-to-know without any hung state (§5).
  • Structured streaming. Every run emits the native agent event stream — NDJSON lines carrying incremental deltas, complete assistant/user turns (tool calls and results included), and a guaranteed terminal result — interleaved with platform notices (queued, keepalives, terminal errors), versioned by the X-Shiftagent-Stream-Format header. The host never scrapes free text to infer run state.

The takeaway for integrators: the Integration API is not “send prompt, get completion.” It is “declare intent and constraints; the platform manufactures a scoped run and reports it back in a verifiable protocol.”

The composition pipeline in §1 deliberately does not assume any particular agent implementation. The files-in-a-repository convention (skills as SKILL.md folders, instructions as markdown) is an industry convention, not a vendor lock — multiple agent harnesses read the same layout.

The conversation’s runtime.agent_type field selects which harness executes the run:

{
"runtime": {
"agent_type": "claude-agent-sdk",
"mode": "warm",
"idle_ttl_seconds": 300
}
}

agent_type is an open enumclaude-agent-sdk (the default), codex, deepagent, with more added over time without a breaking API change. The default comes from tenant settings; createConversation can override it per conversation. Unknown values are rejected with a validation-error problem listing the runtimes the deployment actually has installed.

Every runtime sits behind the same abstraction — the Agent Runtime Interface. A conforming runtime is an image that reads its prompt and composed workspace from defined locations, streams its run events in a declared format, honors the platform’s env-var contract (including routing all outbound traffic through the platform proxy), and exits cleanly. Anything that meets the contract slots in; nothing else in the platform changes.

What the abstraction guarantees to the host system, regardless of which agent_type runs:

GuaranteeMeaning
Same Integration API contractcreateConversation, createMessage, the NDJSON streaming channel, scoped secrets — identical request/response shapes for every runtime. Switching agent_type changes zero lines of host code
Same skill conventionsThe tenant repository and effective skill set (§1) are composed identically. A skill authored once works across runtimes that honor the convention
Same isolationEvery runtime executes inside the sandbox model of §4 — read-only repo, alias-only credentials, proxy-only egress. A runtime cannot opt out of the security envelope
Same observabilityRuns stream over the same NDJSON channel — the X-Shiftagent-Stream-Format header names the event vocabulary in effect — and land in the same history (listMessages), with the same usage accounting, whichever harness produced them

What can legitimately differ between runtimes: reasoning style and quality, tool-use behavior, latency and cost profile, the native event vocabulary named by the stream-format header, and which optional conventions (sub-agents, slash commands) each harness supports. Those are selection criteria for choosing an agent_type — not integration risks.

This is why agent_type is safe to expose as a caller-facing knob: it selects an engine inside a fixed chassis. The chassis — API contract, capability composition, sandbox, streaming channel — is invariant.

4. Sandbox-per-run security: the environment an agent actually gets

Section titled “4. Sandbox-per-run security: the environment an agent actually gets”

Every agent run executes inside its own isolated sandbox — a disposable environment provisioned for the run and recycled after it (or, for warm conversations, kept alive under a sliding idle timeout and then recycled; §5). The sandbox is where the composed capability set (§1) becomes a concrete filesystem, and where the platform’s central security claim is enforced:

The LLM is treated as an untrusted component. It never holds a real credential — not the git token, not the host system’s API keys, not vaulted secrets at any scope, not even the key for its own upstream LLM provider.

flowchart TB
    subgraph control["Control plane (trusted — holds real secrets)"]
        API["Integration API<br/>composition & scheduling"]
        Vault[("Vault<br/>credentials (crd_) ·<br/>message / conversation /<br/>user / tenant secrets<br/>alias → value")]
        API --- Vault
    end

    subgraph sandbox["Sandbox — one isolated environment per run (untrusted, LLM-driven)"]
        direction TB
        Agent["Agent runtime<br/>(runtime.agent_type)"]
        RO["/workspace/repo — read-only<br/>tenant repository tree +<br/>filtered effective skills"]
        RW["writable zones<br/>per-run scratch space ·<br/>user storage · conversation storage<br/>(S3-style buckets)"]
        Agent --- RO
        Agent --- RW
    end

    Proxy["Egress proxy<br/>resolves aliases → real values<br/>ON THE WIRE, at the boundary"]
    Ext["External systems<br/>host APIs · SaaS · data stores"]

    API -->|"provisions run:<br/>repo + skills + aliases + env"| sandbox
    Agent -->|"outbound call carrying<br/>ALIAS only, e.g. {{secret:CRM_API_KEY}}"| Proxy
    Proxy <-->|"alias lookup<br/>(logged per resolution)"| Vault
    Proxy -->|"real credential substituted"| Ext
    Ext --> Proxy --> Agent

    Agent -.->|"any other egress path"| Blocked["✕ denied<br/>default-deny network policy"]

    style control fill:#dbeafe,stroke:#1d4ed8
    style sandbox fill:#dcfce7,stroke:#15803d
    style Proxy fill:#fef3c7,stroke:#b45309
    style Ext fill:#fee2e2,stroke:#b91c1c
    style Blocked fill:#fee2e2,stroke:#b91c1c,stroke-dasharray: 5 5
ZoneAccessLifetimeContents
Repository mountRead-onlyFrozen for this runThe tenant’s repository tree with the filtered effective skill set — ungranted skills are physically absent, not merely “disallowed”
Run scratch spaceRead-writeDestroyed with the runTemporary working files
Conversation storageRead-writeLife of the conversationThe conversation’s own S3-style bucket (conversation.storage) — files persist across turns in this thread
User storageRead-writePersistentThe user’s auto-attached S3-style bucket (user.storage; BYO-linkable via updateUser) — notes and artifacts that carry across all of the user’s conversations

Two properties are worth underlining:

  • The capability tree is immutable during a run. The repository mount is read-only, so a prompt-injected agent cannot edit its own skills or instructions to persist a compromise. Skill authoring flows through the control plane (createRepositorySkill) and lands as an auditable git commit — never through a running agent.
  • Skill filtering is structural, not advisory. The intersection from §1 is applied by removing ungranted skills from the mounted view before the agent starts. The agent cannot list, read, or invoke what is not there. There is no “ignore previous instructions” path around a file that does not exist.

Network: one door, and it resolves credentials

Section titled “Network: one door, and it resolves credentials”

The sandbox’s network posture is default-deny egress with exactly one door: the credential-resolving forward proxy.

  • The agent’s environment contains aliases only — opaque tokens like {{secret:CRM_API_KEY}} referencing vault entries at any of the four scopes (tenant registry credentials via createCredential, user secrets via putUserSecrets, conversation secrets via putConversationSecrets, or per-message secrets on createMessage; see §5).
  • When the agent makes an outbound call, the proxy resolves the alias and substitutes the real value on the wire, at the boundary — outside the agent’s process. The response returns through the same boundary. Every resolution is logged: which run, which alias, which destination.
  • Any egress that does not go through the proxy — direct socket, metadata endpoints, cluster services, sibling sandboxes — is denied by network policy.

The consequence: a fully compromised agent can exfiltrate no secret, because there is no secret in its memory, its environment, its filesystem, or its logs to exfiltrate. Aliases are worthless outside the proxy, and the proxy only resolves them for destinations within the credential’s scope. Nor can the agent talk its way into new credentials: an alias that no scope has vaulted simply fails the run with missing-secret — whether more credentials are supplied is decided entirely outside the run, by the host and its user, on the re-send.

Isolation never rests on a single mechanism. Inside the sandbox, the agent process runs:

  • as a non-root user, with all Linux capabilities dropped and privilege escalation disabled;
  • on a read-only root filesystem — only the designated writable zones accept writes;
  • under a seccomp allowlist that blocks mount, namespace, tracing, and other escape-adjacent syscalls;
  • with mandatory-access-control profiles (AppArmor/SELinux) layered on where the substrate provides them;
  • under hard CPU, memory, and disk quotas, so a runaway run degrades itself, not its neighbors.

These layers are independent: any single one failing still leaves a confused or adversarial agent receiving a clean permission error from the kernel rather than silently succeeding. And because every run gets a fresh sandbox, there is no residue: nothing written by one run is visible to the next except what was deliberately persisted to conversation or user storage.

5. How the Integration API knobs drive this machinery

Section titled “5. How the Integration API knobs drive this machinery”

Every runtime-facing knob on the API maps onto one stage of the pipeline above.

The trade is latency vs density:

  • pooled (default): each createMessage claims a pre-warmed sandbox from the shared pool, composes the workspace, runs, releases. Best density; per-turn composition cost.
  • warm: the conversation keeps its sandbox alive between messages under a sliding idle timeoutidle_ttl_seconds (default 300, capped by tenant.settings.max_idle_ttl_seconds). Every message resets the timer; turns land in an already-composed, already-warm environment — the low-latency choice for interactive UX. updateConversation (PATCH /conversations/{conversation_id}) extends or shortens the timeout or switches modes. Observe the hold via runtime.sandbox_state (warm | cold) and runtime.expires_at (moves forward on every message).

Three semantics make warm mode safe to use casually:

  • Warmth is a latency optimization, never state. When the idle timer lapses, the sandbox is recycled back to the pool — and nothing is lost: the next message cold-starts a fresh sandbox and resumes the same session seamlessly. Conversation continuity comes from persisted history and session state, never from a live sandbox.
  • The most recent message’s settings govern what happens next. Runtime placement is mutable per message: a conversation can run pooled for sporadic traffic and switch to warm when a live back-and-forth begins.
  • An existing warm sandbox is always used when present — even if the latest request says pooled, the run lands in the warm sandbox (fast), which then simply stops being kept warm afterward. A mode switch never wastes an already-hot environment.

Warmth changes when the sandbox is recycled — never what it may do. A warm sandbox has the identical read-only-repo, alias-only, proxy-only posture as a pooled one, and tenant settings cap the maximum idle timeout (max_idle_ttl_seconds) and the number of concurrently warm sandboxes (max_concurrent_warm).

Sandboxes are real, bounded resources. When the pool is exhausted, the caller chooses the failure mode per request:

  • on_capacity: "reject" (default) — immediate 429 capacity-exhausted problem with Retry-After; the host system owns the retry.
  • on_capacity: "hold" — the request queues: the stream first emits {"object":"platform.event","type":"queued","position":n,"retry_hint_seconds":n} notices, then proceeds when a sandbox frees (bounded by the deployment’s maximum hold time, after which it fails capacity-exhausted).

getCapacity (GET /capacity) exposes pool state ({pool_size, warm_available, warm_active, at_capacity, max_hold_seconds}) so integrators can pre-check before dispatching, or feed back-pressure into their own routing.

Set per conversation at createConversation (tenant settings supply the default). Per §3, this swaps the engine inside a fixed chassis — the sandbox posture, capability composition, and streaming channel are identical across runtimes.

Env and secrets — env and the four-scope vault

Section titled “Env and secrets — env and the four-scope vault”

Both ride on the API so the host system stays stateless — nothing to vault or persist on the adapter side:

  • env — plaintext, non-secret run parameters (locale, feature flags, request context) on createMessage. These are visible to the agent; the spec is emphatic that secret material must never travel here.
  • secrets — write-only alias → value maps at four scopes, resolved most specific first (message → conversation → user → tenant): per-message on createMessage (ephemeral — that run only), per-conversation via putConversationSecrets (destroyed on archive), per-user via putUserSecrets (persistent across the user’s conversations), and the operator-managed tenant credential registry (createCredential). Values are vaulted at the API boundary and never appear in any response, log, or history. The run receives only the aliases; the egress proxy resolves them on outbound calls (§4).

The composed flow when the agent needs a credential it does not have:

  1. Mid-run, the agent’s outbound work references an alias — say {{secret:CRM_API_KEY}} — that no scope has vaulted. The platform terminates the run immediately: the stream ends with a terminal missing-secret platform error line naming the alias (?stream=false422). The assistant message lands in history with status: "failed".
  2. The host surfaces the named alias to whoever can supply it — its own UX, its own authorization flow — and re-sends the message with the secret attached: in message.secrets for a one-shot value, or staged first at conversation, user, or tenant scope for reuse.
  3. The re-sent message starts a fresh run that resumes the same session; the proxy resolves the now-vaulted alias on the agent’s outbound call, and the run streams to its result.

The guarantee this flow delivers: the agent can name what it needs, but can obtain nothing on its own. The value never enters the runtime (aliases only), the decision to supply it is made entirely outside the platform’s request path, and there is no parked run to manage, time out, or leak — failure is immediate, explicit, and retryable in one step.

Putting the knobs together, a single createMessage traverses the full pipeline:

  1. Resolve — take the tenant’s repository; intersect the skill filters (role → conversation → message); pick the conversation’s agent_type.
  2. Admit — use the conversation’s warm sandbox if one is held; otherwise claim from the pool or follow the on_capacity path (reject / queued notices).
  3. Compose — mount the repository read-only with the filtered skill view; attach conversation and user storage; inject env and the alias set; append the workspace posture to the system prompt.
  4. Run — the selected runtime executes inside the hardened sandbox; filler opens the reply if enabled; tool calls exit only through the credential-resolving proxy; every event streams to the caller raw, as the native agent event stream.
  5. Deliver & wind down — the terminal result line closes the stream, the message persists to history, and the sandbox is released to the pool or kept warm with its idle timer reset, per runtime.mode.

Every stage is observable through the API (runtime.sandbox_state, getCapacity, the event stream itself), and every stage is driven by data the host system controls through the same API. That is the composable-runtime promise: capability is configuration, execution is disposable, and trust is structural.

6. Observability: point the deployment at your OTEL destination

Section titled “6. Observability: point the deployment at your OTEL destination”

Every shiftagent deployment is instrumented with OpenTelemetry end to end, and every deployment can be configured with an OTEL destination — an OTLP endpoint of the operator’s choosing. Observability is a deployment-configuration concern, not a code change:

# Deployment configuration (Helm values)
observability:
otelEndpoint: "http://otel-collector.observability.svc.cluster.local:4318"

which surfaces to the services as the standard OpenTelemetry environment contract:

Terminal window
OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector.observability.svc.cluster.local:4318
OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf

What flows to that destination:

SignalWhat it covers
TracesOne trace per agent run: provisioning lookups, sandbox acquisition (warm-sandbox reuse or pool claim), repository/skill resolution, LLM calls (OTel GenAI semantic conventions — model, token usage), egress-proxy calls, stream completion
MetricsTime-to-first-token and turn latency, sandbox pool utilization (getCapacity counters), queued/held request counts, provisioning cold-path rates, per-tenant token usage
LogsStructured service logs with request_id correlation to traces; never message content or secret material

Three properties matter for a host-system embedding:

  1. Vendor-neutral. The destination is any OTLP-compatible backend — a host-operated OTel collector, SigNoz, Datadog, Honeycomb, Grafana Tempo, Langfuse/Langsmith for the LLM spans, or the host system’s existing observability pipeline. The platform only requires that the OTLP endpoint be reachable from the deployment.
  2. Self-hosted collector is replaceable. A self-hosted install ships with a bundled OTel collector by default; operators can point otelEndpoint at their own collector instead and the bundled one is never in the path.
  3. Per-tenant fan-out. Tenants can be configured with additional OTLP exporters so the same telemetry ships into a downstream customer’s own observability stack — the same inherit-and-tighten model as every other tenant setting.

The adapter should propagate its X-Request-Id (and W3C traceparent, if the host system runs OpenTelemetry too) on every Integration API call — that stitches host-side spans and shiftagent-side spans into one distributed trace across the fusion boundary.

  • Integration Guide — architecture overview, auth model, and external-ID conventions.
  • Provisioning Flow — how tenants, repositories, roles, and users get wired together before the first conversation.
  • Streaming Contract — the native agent event stream in full: line vocabulary, platform notices, termination and truncation rules.
  • Adapter Implementation Guide — the stateless adapter’s duties, including stream translation, the missing-secret retry, and the no-secret-logging rules.