⚓ firstmate field guide
The AXI ecosystem

Herdr & the wider toolkit

Herdr, gnhf, /vision, backpass, kun, compact-adviser and Jev — the neighbouring tools in Kun's stack, what each is for, and which ones firstmate actually wires in.

Only some of Kun's tools are dependencies of firstmate. The rest are part of the same working style. Here is which is which:

Tool Wired into firstmate? One line
no-mistakes, treehouse, gh/chrome-devtools/tasks/quota-axi yes (required) see the previous pages
lavish-axi yes (optional) visual boards and reviews
Herdr yes (a backend option) agent-native terminal, Kun's daily driver
Jev (typesafe.ai) yes (optional, TYPESAFE_API_KEY) cheap classifier for dispatch routing
compact-adviser explicitly disabled in workers (COMPACT_ADVISER_DISABLE=1) when to /compact
gnhf, backpass, kun, /vision no separate tools, same philosophy

Herdr#

Herdr is a terminal and session manager built for agents. Per-pane agent state and push events are what make it a better host than tmux for a fleet. Kun on why it fits:

herdr provides the lifecycle management and organization of all the agents. firstmate serves as a single point of contact to orchestrate all of them without going insane. — @kunchenguid, 2026-09-09

He describes it as "more like an SDK that happens to come with a default UI", with programmatic and remote access. He sees that as the difference from UI-first alternatives. Setup and caveats: Runtime backends. Herdr is not Kun's project; firstmate integrates it.

gnhf — "good night, have fun"#

Before I go to bed, I tell my agents: good night, have fun

gnhf is a long-running loop orchestrator for unattended overnight work, inspired by Karpathy's auto-research pattern. It is on npm and was Kun's first project to pass 1,000 stars (April 2026). It covers one-off loops:

'one off loops' = a loop that babysits an ephemeral task and will end once the task is done. 'durable loops' = a cron job that keeps running forever. — @kunchenguid, 2026-06-08

firstmate doesn't call gnhf. It supports the same harnesses, Cursor CLI included.

/vision#

This is a skill (kunchenguid/vision) that mines a repo's history to draft a VISION.md. It stress-tests the draft with 8–12 hard hypothetical feature ideas and iterates with you on a review board. firstmate's own VISION.md shows the output: principles an agent can use to adjudicate "should we build this?".

what i've been experiencing for the past couple of months is telling me that eventually we won't even be able to review all the plans... our influence has to go one level up - we need to define the vision. — @kunchenguid, 2026-08-17

This matters for firstmate because initiative beyond your literal request is legitimate "only where the captain has committed a vision precise enough to adjudicate it".

backpass — train your AGENTS.md#

From Kun's post Your AGENTS.md is a Neural Net. The analogy:

  • the memory file is the weights, and its token budget is the model size,
  • each session is a forward pass,
  • editing the file from transcript evidence is the backward pass.

npx -y backpass harvests local transcripts from Claude Code, Codex, Pi, OpenCode, Grok and Cursor. It aggregates "gradients" and drops anything seen in fewer than two sessions. It proposes at most 5 edits per run, each backed by a verbatim quote, and a human approves every write in lavish. Suggested cadence: weekly per active repo.

Case study: firstmate's own AGENTS.md (2026-09-27). Kun ran backpass over 100 recent firstmate session transcripts. He rejected two of its proposed edits and accepted the rest, and the always-loaded AGENTS.md (previously ~30k tokens) shrank by 48%. "biggest win came from a lot of instructions getting moved into skills, based on observation that they weren't actually needed in most transcripts". The resulting change is firstmate 49a218b (#5872), which moved the home layout, session-start recovery, validation and landing supervision, away/quiet supervision and more into on-demand skills. He checked it against his private eval set for regressions.

Inside firstmate, /stow plays the same role for data/captain.md and data/learnings.md: evidence-only reinforcement, decay, and a token budget.

kun — Kun, distilled#

npx skills add kunchenguid/kun -g, then /kun <question>. It is a daily-updated distillation of his public writing and activity. Per his post I distilled myself, the motive was that his fleet kept escalating judgment calls to him. It is handy for questions like "how would Kun set up dispatch for X?".

Jev and compact-adviser#

Jev is a small, cheap model from typesafe.ai that returns a decision with a confidence rather than generated text. With TYPESAFE_API_KEY set, firstmate uses it for dispatch resolution. Kun measured −71% cost and −90% wall time on that step, with matching choices (2026-09-17). His framing:

LLMs make us all build agent loops. this enables us to explore a different architecture - a combination of deterministic logic (code), intelligent decision making (Jev), and occasional generation of content (LLM).

compact-adviser is a Jev-powered plugin that says when it is safe to /compact. Every firstmate worker launch sets COMPACT_ADVISER_DISABLE=1, because an unattended session should never act on it.

Unofficial guide built 2026-09-26 from kunchenguid/firstmate, the AXI repos, and Kun Chen's public posts. The repo moves daily; when this guide and the repo disagree, the repo wins.