The supervision engine
How firstmate watches a fleet without burning tokens, what wakes it, why no turn ends blind, and how the durable wake queue survives restarts.
The first mate cannot poll the fleet with LLM turns. That would burn quota and bury your conversation. Instead:
a bash watcher sleeps on the fleet and wakes the first mate only when something needs you — firstmate README
The watcher: bin/fm-watch.sh#
It is a single-shot bash loop. Each cycle it sleeps FM_POLL seconds (default 15), then classifies what changed. What happens next depends on the wake:
- Benign (no-op heartbeat, routine polling): it logs to
state/.watch-triage.logand keeps blocking. No LLM turn is spent. - Actionable: it writes a durable record to
state/.wake-queue, prints one reason line, and exits. Whatever owns re-arming for your harness starts the next cycle.
| Reason line | Fires when |
|---|---|
signal: <file>... |
A worker appended a captain-relevant status: done:, needs-decision:, blocked: or failed: (or legacy PR ready, checks green, ready in branch, merged) |
stale: <window> |
A pane has been idle for FM_STALE_ESCALATE_SECS (240s) with no proof it is working. After FM_WEDGE_DEMAND_INSPECT_COUNT (3) consecutive stale wakes it adds demand-deep-inspection. |
check: <script>: <out> |
A PR merge poll, Relay mention, mail poll, process-event source, or a task's custom check fired |
heartbeat |
Backstop review: every FM_HEARTBEAT (600s), backing off to FM_HEARTBEAT_MAX (7200s) |
A busy pane is exempt from stale detection up to FM_BUSY_TURN_MAX_SECS (3600s) per turn.
Things that suppress false alarms#
paused: <why> until <ISO time>: the worker declared a known wait: an external one (a rate limit, an upstream release) or its own background shell, pipeline run or long command. Workers must declare the second kind before ending a turn on it. A stopped pane still surfaces one first-sight stale wake; after that the watcher rechecks everyFM_PAUSE_RESURFACE_SECS(4h) instead of nagging.captain-held: the item has been handed to a durable captain hold. It is treated like paused.- Dead or missing agents: the watcher reads the
dead/missingverdict once and stops re-escalating. Before this fix, two finished lanes racked up 226 and 203 stale escalations. - Opt-in:
config/wedge-defer-parked-gatedefers wedge alarms while a no-mistakes gate is waiting on the first mate's own decision.config/turnend-churn-absorbcounts pane-content churn as evidence of activity.
Busy state is semantic#
Whether a worker is "busy" comes from harness events, not screen scraping, wherever a hook exists. state/<id>.busy-state holds one line: v1 gen=... seq=... state=busy|idle|unknown source=... event=... ts=.... It is written by per-harness adapters via bin/fm-busy-event.sh:
| Harness | Source |
|---|---|
| Claude | UserPromptSubmit / Stop / StopFailure / SessionEnd hooks |
| Pi, pi-signed | extension agent_start / agent_settled |
| omp | agent_start / agent_end |
| OpenCode | plugin session.status |
| Cursor, Muse | conversation transcript / session log |
| Codex, Kimi | explicit probes |
| Grok, Rovo, AGY | isolated rendered-tail fallback |
"Unknown is never promoted to busy", and missing or stale state is unknown, never idle.
Who re-arms the watcher#
The watcher is one-shot by design, so each harness has an arm owner that keeps it alive between turns:
| Harness | Arm owner | How a wake reaches the model |
|---|---|---|
| Claude Code | .claude/settings.json Stop hook → fm-claude-stop-autoarm.sh (asyncRewake) |
The hook exits 2 with a rewake banner, delivered as Stop-hook feedback |
| Pi / pi-signed | .pi/extensions/fm-primary-pi-watch.ts (fm_watch_arm_pi tool) |
The extension delivers the wake; the supervision branch absorbs routine ones |
| omp | .omp/extensions/fm-primary-omp-watch.ts |
Extension |
| OpenCode | .opencode/plugins/fm-primary-watch-arm.js on session.idle |
client.session.promptAsync |
| Cursor | .cursor/hooks.json stop hook parks on one watcher cycle |
A follow-up turn |
| Grok | The model runs the arm with run_terminal_command background:true |
Grok's synthetic task_completed message |
| Codex | The model loops fm-watch-checkpoint.sh --seconds 180 |
Foreground output |
Adapters start the successor watcher before delivering the current wake, verify it is ready, and retry with backoff. If restoration fails, they still deliver a typed "continuity failure" wake: "the model is never left blind." A recovery episode (state/.watcher-down) keeps resurfacing at session start until it is acknowledged.
The turn-end guard: "no turn ends blind"#
bin/fm-turnend-guard.sh is the backstop. When the first mate tries to end a turn while work is under way and no live watcher has a fresh beacon, the harness hook blocks the stop or forces a follow-up.
| Harness | Mechanism |
|---|---|
| Claude | Stop hook; re-blocks up to FM_CLAUDE_TURNEND_BLOCK_BUDGET (3) times, then one loud fail-open |
| omp | Blocking session_stop hook that compels a continuation (the strongest) |
| Pi | agent_settled listener |
| Cursor | Cannot block (exit 2 is a no-op), so it parks the hook open on the watcher; bounded by FM_CURSOR_TURNEND_LOOP_CEILING (180) |
| OpenCode | session.idle → promptAsync once |
| Grok | Stop hook in .grok/hooks/, blocking when the capability fields are present |
In away mode, the away daemon satisfies the guard instead of a watcher.
The durable wake queue#
Actionable wakes land in state/.wake-queue before anyone acts on them. bin/fm-wake-drain.sh presents:
- the raw records,
OPEN DECISIONS— needs-decision / blocked keys still unresolved,UNREAD STATUS,RECORD DIVERGENCE— the status log and the backlog disagree about a captain call.
The first mate must then acknowledge with the exact --ack-through command the drain printed. If it is interrupted before acknowledging, the work stays queued and is handled again idempotently. This is how "restart is a non-event" holds for supervision itself.
Two supervisors in one home: main and branch#
On Pi, the supervision branch is a second AgentSession inside the same pi process. It takes eligible task-local wakes and heartbeat scans off the main conversation, so your chat isn't interrupted for routine events. It hands captain-facing outcomes back as one processing turn. You can run it on a cheaper model with /supervision-model.
Kun described the pattern when he built it:
when your agent is juggling a lot of such loops, eventually you will see it becoming too busy to even talk to you... the background session will decide whether an event needs to interrupt the main session or not... most events don't need to, but they don't get silently dropped. they get "merged" into the main session like git commits merge between branches — @kunchenguid, 2026-08-24
For other harnesses, the supervision host runs the same branch contract on a headless engine beside the primary (bin/fm-supervision-host.sh). A Claude primary runs it by default, with no config file. Other non-Pi primaries opt in with config/supervision-host, and a file saying off opts any home out. It works while you are away on Claude, Cursor, OpenCode, omp, Grok and Codex. Pi does not need it, because Pi has the native branch. While you are attended, it is active only on Claude and Cursor, which have a proven dialog mirror (bin/fm-host-mirror.sh).
Leases (bin/fm-lease.sh) keep the two actors from colliding. Merging, local landing, spawning and answering your decisions are main-only while you're attended. During confirmed away mode, merge, spawn and decision authority move to the branch.
Knobs worth knowing#
| Env var | Default | Effect |
|---|---|---|
FM_POLL |
15 | watcher cycle, seconds |
FM_HEARTBEAT / FM_HEARTBEAT_MAX |
600 / 7200 | backstop cadence and backoff ceiling |
FM_STALE_ESCALATE_SECS |
240 | idle-without-proof threshold |
FM_BUSY_TURN_MAX_SECS |
3600 | longest tolerated single busy turn |
FM_PAUSE_RESURFACE_SECS |
14400 | recheck cadence for paused / captain-held |
FM_GUARD_GRACE |
300 | beacon freshness window for the guard |
FM_TASK_INBOX_GRACE_SECS / FM_TASK_INBOX_RING_MAX |
90 / 3 | steering-message re-ring |
FM_CHECK_INTERVAL / FM_CHECK_TIMEOUT |
300 / 30 | PR, Relay and custom check cadence |