⚓ firstmate field guide
Using firstmate

Troubleshooting

Stuck or looping crewmates, a dead watcher, "the pipeline is dead" claims, session-start diagnostic codes, merge refusals, teardown refusals, and restart recovery.

The first mate handles most of this itself: skills load at the right trigger. This page is for when you want to know what it is doing, or you are debugging by hand.

First, ask it

"Why is task X stuck?" or "why did you refuse to merge?" The first mate can read its own state, scripts and skills. VISION.md: "When something is not working well, the captain can ask the first mate and it figures it out."

Inspection commands#

export FM_HOME=$PWD                   # in the firstmate checkout
bin/fm-fleet-view.sh                  # whole fleet
bin/fm-crew-state.sh <id>             # current truth for one task
bin/fm-peek.sh <id> 80                # last 80 lines of its pane, non-disruptive
cat state/<id>.status                 # event history (not current state!)
ls state/<id>.inbox/                  # unacknowledged steering messages
bin/fm-watch-arm.sh                   # is supervision alive?
bin/fm-lock.sh status                 # who owns this home's session?
tail state/.watch-triage.log          # wakes the watcher absorbed
tail state/.watch-cycle-exits.log     # watcher lifecycle

A crewmate is stuck, looping or confused#

The stuck-crewmate-recovery skill loads on:

  • a stale wake, a looping pane, or repeated confusion,
  • a worker asking something its brief already answers,
  • an unresponsive worker or a failed steer,
  • a dead endpoint reported at session start.

It escalates in this order:

  1. Look. fm-peek.sh <id>, and check state/<id>.inbox/ for a message it never acknowledged.
  2. Answer. If it's waiting on something the brief answers, fm-send.sh the answer.
  3. Interrupt and redirect. fm-control.sh <id> interrupt, then one corrective fm-send.
  4. Relaunch. fm-control.sh <id> relaunch --note '<progress so far>'. Add --harness/--model/--effort to switch runtimes. It keeps the same worktree and any uncommitted work.
  5. Give up honestly. If a second relaunch fails, it marks the backlog item failed and tells you plainly.

"A low context reading is not wedging — modern harnesses auto-compact." A worker near its context limit is not stuck.

Dead endpoint? A dead pane only says the worker isn't present. It is not proof the work is gone. The first mate checks fm-crew-state.sh first. A still-running no-mistakes run for that branch means it keeps supervising rather than starting a duplicate. It relaunches only after proving no live agent owns the task and the worktree survives. A dead endpoint for a task whose PR already landed is just normal teardown, never --force.

"The no-mistakes pipeline is dead"#

A crewmate's timed-out drive call is usually a slow response, not a dead daemon. The shared daemon serves every lane, so it is never restarted on a crewmate's say-so. Verify first:

no-mistakes daemon status               # authoritative
no-mistakes axi status --run <run-id>
bin/fm-crew-state.sh <id>

Only a refused connection or missing socket from daemon status counts as evidence. That goes to you.

Supervision stopped#

Symptoms: the first mate ends turns while work is under way, or wakes never arrive.

bin/fm-watch-arm.sh --restart    # restarts THIS home's watcher only

Don't

  • pkill -f bin/fm-watch.sh kills every home's watcher on the machine.
  • bin/fm-watch-arm.sh & inside another command dies silently. It must be a tracked background task. The repo documents a ~30-minute outage from exactly this.

The adapters are built never to leave the model blind. A failed restoration still delivers a typed wake. A recovery episode (state/.watcher-down) keeps resurfacing at each session start until the drain's printed --ack-through command acknowledges it.

Claude-specific history. Two incidents (2026-08-14 and 2026-08-26) involved the Stop-hook auto-arm deferring to a stale lock. The first left every turn blind for 40+ minutes. The current epoch/generation design exists because of them. If you're on an old checkout, /updatefirstmate.

Session-start diagnostic codes#

These are handled by the bootstrap-diagnostics skill. Its rule: "detect, then consent, then install".

Code Meaning / fix
MISSING: <tool> (install: ...) Approve once and it runs fm-bootstrap.sh install ...
MISSING_MANUAL: You install it by hand, then restart the session
NEEDS_GH_AUTH Run gh auth login yourself. It can also mean the network is down.
PRESENTATION_UNAVAILABLE: lavish-axi Visual boards are off. Everything else continues in plain text.
BACKEND_INVALID: Bad config/backend or FM_BACKEND
TANGLE: The primary checkout was left on a feature branch. The fix is git checkout <default>, the only git write the first mate may start on its own checkout.
CREW_DISPATCH: / STARTUP_MEMORY_BUDGET: Malformed config. Correct it; there is no fallback.
FLEET_SYNC: ... STUCK A project clone needs hands-on attention (diverged or dirty) before it falls further behind
NETWORK_CHECKS: The deferred network stage didn't finish. Rerun the printed idempotent command.
BACKLOG_RECONCILE: A pending backlog close couldn't replay. Never hand-delete state/<id>.backlog-close.
SECONDMATE_SYNC/LIVENESS/HANDOFF:, NUDGE_SECONDMATES: Secondmate convergence issues
HOME_SUMMARY:, FMX: Remote summary cache; Relay

"Read-only session" usually means another live session holds state/.lock. bin/fm-lock.sh status shows the holder.

Merge refused#

fm-pr-merge.sh re-reads the PR live and refuses unless:

  • it is open and not a draft,
  • mergeable == MERGEABLE and mergeStateStatus != DIRTY (DIRTY means conflicts),
  • every unwaived required check is green at the current head.

It never auto-resolves conflicts. The fix is a crewmate rebase: "have the crewmate rebase the branch". To merge past a red check, your current instruction must name it (--allow-red <check>). Auto-merge, --admin and branch-deletion flags need --attended-override, which requires your explicit instruction.

fm-pr-check.sh refuses a draft PR. Mark it ready and re-register.

Teardown refused#

This is by design. Hard rule 3: never tear down unlanded work. fm-teardown.sh refuses if commits aren't reachable from a remote-tracking branch or a merged PR (it understands squash merges), or if there are uncommitted changes.

  • Find out why. It is "a stop-and-investigate result, never an obstacle to bypass."
  • If you truly want the work gone, say so explicitly. The first mate then uses --force.
  • A task that was an open captain call keeps its backlog row after teardown, with the deliverable attached. That is expected.

Trust dialogs blocking a worker#

The first mate's steering can only send Enter, Escape and Ctrl-C, not arrow keys, so it can't navigate most dialogs.

  • Claude: bin/fm-claude-trust.sh pre-registers workspace trust for each worktree. It pre-approves the "Allow external CLAUDE.md file imports?" dialog only if you already approved it on the primary checkout. Otherwise the dialog still renders in the worker, which is expected. Leave that pane alone and answer the dialog yourself; do not interrupt it. Escape on this dialog records a permanent decline, and after a recorded decline the script refuses to register the project at all, so spawns for it fail. To recover, remove hasClaudeMdExternalIncludesApproved and hasClaudeMdExternalIncludesWarningShown from the project's entry in ~/.claude.json, then approve the imports dialog once by hand.
  • Cursor must run with --trust. Kimi is refused on cmux and Orca (its folder-trust dialog needs a capture those backends lack).

Worker refused to launch: account pin#

If config/claude-account or config/pi-account pins a login that isn't signed in, spawns are refused. The first mate reports which login is needed. It never edits the pin. Sign in, or change the pin yourself.

After a crash or reboot#

Just start the harness again in the checkout. bin/fm-session-start.sh takes the lock, reconciles the backlog, drains queued wakes, lists every task's meta and endpoint liveness, and relaunches dead secondmates. Crewmates whose panes died are recovered through the relaunch path, in their original worktrees. There is no general "resume" verb. Resume behaviour differs per harness, so recovery always goes through relaunch.

Herdr-specific#

  • Below Herdr 0.8.0, closing a background workspace steals focus for about 1/7s. Presentation spaces are gated to ≥0.8.0, and a one-time marker tells you to upgrade.
  • A pane frozen on an old agent's status is caused by Herdr's sticky per-session authority. A relaunch must reuse the old session binding, which firstmate does for Pi via --session.
  • A stale agent=pi registration after exit comes from a nested shell under the pane (upstream Herdr issue #4115). firstmate reads the process table instead of trusting the registration.
  • On remote hosts, a ~/.local/bin/herdr can shadow a newer package-managed client. The adapter detects protocol_mismatch, but you should remove the old copy.
Unofficial guide built 2026-09-26 from kunchenguid/firstmate, the AXI repos, and Kun Chen's public posts. The repo moves daily; when this guide and the repo disagree, the repo wins.