Kun's playbook
How the author actually runs firstmate — machines, backends, model org chart, routing rules, reasoning effort, and the habits behind 50–100 PRs a day. Sourced from his posts, dated, because his stack changes monthly.
Dated advice
Kun re-tunes his model lineup almost monthly. Everything below is dated. Treat later entries as replacing earlier ones, and treat the reasoning as more durable than the model names.
The shape of his day#
i just talk to firstmate for whatever i want to do, and it'll manage the rest of the crew to get work done. decisions and PRs get surfaced to me through firstmate where i apply judgment. that's it. — @kunchenguid, 2026-08-20
- One agent he talks to. "firstmate: the only agent i directly talk to." (2026-06-27)
- Intent in, PRs out. "what I do for each PR is really just specifying my intent and requirements clearly. then the agents are pretty good at getting to a 90% correct implementation. no-mistakes then takes it to 100%." (2026-04-20)
- He doesn't read every line. "I started asking the agent what changed, how it's tested, what are some risky areas i should indeed take a look at etc. all that workflow is now packaged into no-mistakes." (2026-04-20)
- Parallelism. With firstmate handling context switches, he went "from 3-5 to now 10-20" parallel sessions. He noted that "token quota is the clear bottleneck", not speed. (2026-06-26)
- Plans and reviews happen in HTML. "review the PR with me in lavish". Key decisions render as buttons.
His machines#
From the 2026-08-05 setup thread:
- MacBook — the interactive first mate.
- Headless Mac mini — a remote secondmate for iOS work, "so it never competes for interactive resources".
- Hetzner VPS boxes — more secondmates.
- All on one Tailscale network with SSH keys. The VPSes can't reach his personal devices.
- Herdr replaced tmux in July: "i have completed replaced tmux with herdr".
- Scale by late August: "1 firstmate and 7 second mates".
On hardware: "agents are very light weight by themselves... hardware spec depends more on what kind of projects you are trying to work on - build and test can be resource intensive."
On renting versus owning, he came down hard on owning in late September. His Mac mini cost $2,500; a similar VPS spec is about $300 a month, so nine months of rent exceeds the purchase price. "if you are considering a long term setup, i can’t think of any good reason to be renting VPS. own your hardware - my desk is my cloud" (2026-09-27).
Harness by capability#
for video generation and real time info, i use grok build. image generation uses codex - this is because of harness capability. for claude models i use claude code - because their TOS does not allow any other harness. everything else is pi.
His model org chart (latest: 2026-09-25)#
After the Opus 5.5 release he regrouped models into five buckets (2026-09-25):
| Bucket | Models |
|---|---|
| Interactive orchestrators (first mate, secondmates) | opus 5.5, grok 4.7 |
| Premium intelligence (escalations) | gpt 6 astra, fable 5.1 |
| Planners | opus 5.5 |
| Implementers | opus 5.5, gpt 6 sol, grok 4.7 (budget: muse spark 1.3, deepseek v4 flash) |
| Trivial fixers | gpt 6 luna |
Update, 2026-09-29: Sonnet 5.5 as first mate. After a day of real use he moved his own first mate to Sonnet 5.5: "i also tested heavily using sonnet 5.5 as my firstmate and it’s working super well! i even prefer it over opus, because it’s faster, cheaper, and smart enough to understand my intent and handle the orchestration" (2026-09-29). His split in the same post: opus plans, sonnet implements, haiku will take "fixes" once Haiku 5.5 ships, and fable advises when something is not going well. He does not treat Sonnet as a cheaper Opus for planning: on ambiguous product problems, Opus asks what the real goal is while Sonnet looks at how to get it done.
The lineup changes. The principles have held steady:
- The orchestrator needs speed, judgment, and to be pleasant to talk to, not max IQ. "i don't think fable class models are the best orchestrator because a lot of mundane orchestration work really doesn't need that much intelligence. It's better to use a fast and efficient model for orchestration, and let it delegate tough questions to fable/astra when needed" (2026-09-11).
- Crewmates are routed by hard capability first (images, video, realtime data), then by ambiguity (well-defined work goes to cheap models, ambiguous investigation to strong ones), then by live quota via quota-axi.
- Expensive planners need explicit approval. In August he gated fable- and kimi-class planning behind explicit approval, because those models are expensive.
- Slow-but-cheap models belong in the background. GPT-6 Luna was unusable as an interactive first mate: "by the time it finishes the current turn there are already two more turns worth of events piled up". It is excellent as a trivial-fix crewmate (2026-09-23).
- Review gets its own pin. no-mistakes review stays on whichever model is most thorough per dollar.
Earlier snapshots, for context:
- 2026-06-30 — first mate on opus 4.8 max; crewmates on sonnet 5 / grok / codex; all PRs through no-mistakes on gpt 5.5.
- 2026-07-11 — first mate on GPT-5.6-Sol xhigh (Pi) or Grok 4.5 high; secondmates on Fable 5.
- 2026-08-05 — first mate on grok 4.5, falling back to opus 4.8 when quota runs out; secondmates on opus 4.8.
How he evaluates a new model: the "blind date"#
my way of really evaluating a model has become increasingly like a blind date - we hang out, we do a few things together. if it clicked, we get a second date... don't tell me your IQ, your SAT or your GPA - they mean nothing on a date. — @kunchenguid, 2026-09-04
His test is to run the model as firstmate for a full day. Firstmate's contract is long and strict, which makes it a good probe: "firstmate stretches frontier models' reasoning capability and is a really good test that can quickly reveal how good a model is at following instructions" (2026-07-17).
Reasoning effort is a ceiling, not a floor#
reasoning effort level in mainstream LLMs today means a "ceiling", not a "floor"! setting a high effort level does NOT mean every prompt you send will use a lot of thinking tokens. — @kunchenguid, 2026-09-25
His defaults:
- medium for the first mate. Ambiguous work has already been routed to crewmates.
- xhigh for planning and investigation.
- He mostly avoids max.
A cache gotcha from the same thread: "switching reasoning effort level will change either the shape of the request or a tiny part of the system prompt and breaks prompt caching." That is one reason to pin effort per role in dispatch profiles rather than flip it mid-session.
Why routing lives in the first mate#
the complexity of a task only reveals itself when you start working on it... routing should work at the boundary between agents and subagents, not per each LLM request. — @kunchenguid, 2026-07-23
i ideate and plan with my firstmate, and firstmate decide which model/harness/reasoning to use when dispatching tasks to crewmates, based on some preferences i wrote. it also takes into account which subscription quota i have the most of. works extremely well. — same thread
Prompt caching is the second reason routing can't be per request:
the second problem is prompt caching. it means we can’t frequently change the model during a session in an economic way. one or two missteps here and there, and you find yourself spending more cost than simply using the best model all along — @kunchenguid, 2026-09-26
Habits worth copying#
- Wait until it hurts. "don't use firstmate until you feel the pain of juggling between many sessions."
- Write intent, not instructions. The brief's
## Captain's intentis your words verbatim. Clear intent is the highest-leverage thing you type. - Put durable things on disk. "anything that needs to serve as a durable record should not be trusted to stay only in the context window. they must be pushed and persisted onto the disk, and made discoverable by the agent." Use
/stowanddata/captain.md. - Use one session for intent. Don't split into siloed agents: "it's easier to express all your intent in a single session so this session can make the right decisions on behalf on you. if you created a few silos, then you will have to carry the context around."
- Let the first mate fix itself. When something is off, ask it why, then ask it to fix its own distro. That is how the 500+ community PRs happened.
- Write a VISION.md for projects where you want the crew to take initiative.
- Use guardrails, not micro-management. "any human organization that produces a large amount of code WILL have bad code slipping in. this is not an AI problem - it's a complexity and scale problem. fight it with good culture and guardrails, not micro-management." (2026-07-03)
Community Q&A worth knowing#
- Do crewmates share context with the first mate? "they are separate context windows, so the context is different. but firstmate doesn't "gatekeep" any context - crewmate can get whatever context from firstmate." (link)
- Is it a harness? "firstmate is not a harness :)" (link)
- Remote machines? They're built in, as "remote second mates" — ask your first mate to set one up. (link)
- Isn't Claude's dispatch the same thing? "isn't claude dispatch just firstmate, but way more constrained to a specific scenario and locked into a specific vendor?" (link)
- Where does the Relay
.envgo? The repo root, notconfig/.env.