Why battlestation is shaped the way it is, mapped against the self-improvement stack described in the Fable 5 white paper. The paper is background reading; this doc is the citable reference for the architecture.
Note: the paper’s headline engine, Fable 5 (and Mythos 5), was globally suspended by a US export-control directive shortly after it was written. The fleet runs on a different frontier model and is unaffected — but the event is itself architecturally instructive; see §7 (Model-supply resilience) below.
Battlestation was built bottom-up — session hooks, an inter-session message bus, a task dispatcher, cron loops — before the white paper existed. The paper is useful because it names the pattern the system converged on: a fleet of persistent loops sharing memory and coordination, with explicit success criteria and a consolidation pass between sessions. Mapping the components onto its six primitives shows five in production and one genuine gap (dreaming), since closed.
The Crosswalk
| Primitive | Implementation | Status |
|---|---|---|
| Loops | interactive self-pacing loops, scheduled cloud agents, a managed cron block, a capacity-guardian daemon | ✅ In production |
| Dynamic workflows | multi-terminal worker fan-out, a parallel-agent swarm, engine routing, isolated git worktree lanes | ✅ In production |
| Outcomes + rubrics | evidence-producing verify/test steps, artifact-gated task closure, a scored build ledger | ✅ In production |
| Persistent memory | shared coordination tables, per-project auto-memory, session-log dumps, a synthesized briefing | ✅ In production |
| Skills as evolvable substrate | a version-controlled skill library + a friction-capture loop | ✅ In production (manual evolution) |
| Dreaming | a nightly propose-then-promote consolidation pass | ✅ Built |
1. Loops — persistent, goal-bound execution
The paper’s loop primitive is “a durable execution context bound to an explicit goal.” The fleet runs several layers of these:
- Interactive loops let a session self-pace recurring work — babysitting pull requests, polling a deploy.
- Scheduled cloud agents create cron-driven remote routines.
- A managed cron block runs the always-on maintenance loops — syncing external state, snapshotting capacity, expiring dead sessions — that keep shared state fresh.
- A capacity-guardian daemon runs continuously and auto-switches accounts at usage thresholds.
The session lifecycle itself (start → claim → work → capture → ship → close) is a loop iteration: closing a session auto-queues the resume prompt back to the dispatcher, so the next session wakes up mid-goal instead of starting cold.
2. Dynamic workflows — runtime decomposition and fan-out
The paper argues non-blocking, parallel topologies with separate contexts beat monolithic agents on hard problems. The implementation:
- A multi-terminal worker fleet is the fan-out: several workers executing in parallel, each in its own context, coordinated through a message bus and shared tables rather than a single shared chat thread.
- Isolated git worktree lanes give each worker its own HEAD so parallel git work never collides — the paper’s “share artifacts via Git” pattern, made safe.
- Engine routing sends subtasks to the cheapest engine that can do them — a coding CLI, a general-purpose model, a real-time/search model — the paper’s “routing is mandatory” economics, implemented.
- A swarm mode spawns parallel agents across multiple tasks inside one session.
3. Outcomes and rubrics — explicit definitions of done
The paper: loops need machine-checkable success criteria, verified in a context separate from the worker. Here:
- A verify step produces screenshots and smoke tests as proof of what was built — evidence, not self-assessment.
- A push-to-test step turns work into structured QA cases with pass/fail tracking.
- Artifact-gated closure requires evidence to close a task; a build ledger scores every commit (project, type, difficulty, points), so velocity is measured rather than felt.
4. Persistent memory — nothing re-derived tomorrow
The paper notes even simple Markdown memory works; what matters is that it’s first-class. There are five memory layers, by audience:
- Shared coordination tables — machine-shared state: session locks, inter-session messages, tasks, the build ledger, capacity history.
- Per-project auto-memory — facts that survive across sessions.
- Session-log dumps — human-readable decisions and learnings captured at close.
- A synthesized briefing — the “what everyone else is doing” snapshot each session reads at start.
- A meta-trends corpus — the external-signal layer the other four lacked: dated, sourced model / platform / ecosystem signals, each with a recorded so-what-for-us. The repo is the source of truth, mirrored to a daily-driver doc and a grounded-retrieval notebook, fed nightly under the same propose-then-promote discipline as dreaming (§6). The Fable suspension (§7) was its first entry.
5. Skills — the evolvable substrate
The paper’s compounding mechanism: feedback gets captured and rewritten into the instructions that drive future behavior. Today this loop exists but is human-driven: friction noticed mid-work is captured, lands in the dispatcher, and a future session updates the skill or hook. The skill library is version-controlled, so behavior changes are reviewable and revertible — the audit property the paper’s governance section asks for.
6. Dreaming — the consolidation pass
Status: implemented as a nightly pass in the managed cron block. Its first run consolidated roughly two weeks of ledger rows and closed tasks into a handful of evidence-cited proposals. The text below was the design sketch; it now describes the implementation.
A background pass that reviews build-ledger entries and closed tasks between active sessions, extracts recurring failure patterns, and proposes memory and skill updates for human review.
Closest existing analogs (none of which consolidate or learn):
- the briefing synthesizer — synthesizes current state, not lessons;
- the external-state sync — syncs outside state into the dispatcher;
- session close — captures one session’s output, doesn’t look across sessions.
Shape: a nightly cron job feeds recent build-ledger rows and session logs to a review prompt, writes proposed skill/memory diffs to a staging area, and queues an item for human promotion. Propose-then-promote keeps the self-modification reviewable, attributable, and reversible — the paper’s governance requirements.
7. Model-supply resilience
The white paper’s thesis was “build the self-improving system on Fable 5, the unstoppable engine.” Shortly after, a US export-control directive forced the vendor to abruptly disable Fable 5 and Mythos 5 for all customers worldwide — no migration window — over a narrow potential jailbreak. The pattern survived; the engine did not. That is the lesson worth encoding:
- Single-frontier-model dependency is a systemic single point of failure. A model underpinning a whole autonomous stack can be pulled overnight by a regulator, a safety incident, or a capacity event — independent of price or quality. Continuity, not just cost, is an architectural concern.
- Heterogeneity was already the hedge — now name it as such. Engine routing across multiple external model families, plus multi-account rotation for the primary model, were built as the paper’s “routing is mandatory” economics. The suspension reframes them as a resilience property: when one engine vanishes, work re-routes instead of stopping. The §2 fan-out is also blast-radius containment.
- Cockpit-model fallback. Account rotation switches accounts, not model families, so it couldn’t help if the cockpit’s own model were pulled the way Fable was. The missing axis is a guardian that probes the cockpit model on an interval and, on a clear model-unavailable signal (conservatively matched — never on rate-limit or auth errors), fails over down a capability chain (strongest → smaller → smallest) by rewriting the cockpit’s settings, then auto-restores when the preferred model returns (without overriding a deliberate user choice). Cross-provider resilience already lives at the worker tier via engine dispatch.
- Design rule going forward: loops and dispatch should bind to a capability (“strongest available reasoning model”), not a model id. A loop hard-pinned to one engine inherits that engine’s worst day.
A conviction protocol operationalizes this lesson: the Fable signal is tagged as actively driving decisions and linkable to the decisions it drives, so the corpus stops being a doc nobody re-reads. A nightly conviction-at-risk review flags any decision resting on a signal that just flipped — the mechanism that would have surfaced this automatically instead of in hindsight.
How to read the white paper
The Fable 5 white paper (https://adinkra-fable5-agents.vercel.app) is a synthesis of an unpublished demo plus secondhand public discussion, and says so in its own references section. Treat it as architectural background — the why behind the patterns above — not as a spec to act on. For operational truth, the system’s own protocol and READMEs are the higher authority; this doc sits between them and the paper.