Tickets, frozen tests and loop budgets: our engineering discipline
One ticket from start to merge, and the engineering rules behind it. Tests first, frozen tests, time boxes, guards that must fail, review and one merge owner.
The overview says discipline, not headcount, is how a small team ships fast, and that we use AI tools to move faster on the repetitive parts. This post shows the machinery: the rules our engineers set, and the checks that hold every change, including AI-assisted ones, to them. The examples come from Gbege, our game, which has the strictest version of these rules. Our other repos use the same ideas with lighter tooling.
1. The ticket
A ticket is a short file. It has:
- Scope: two or three sentences on what changes.
- Depends on: tickets that must be done first. A ticket can only be picked up when all of them are done.
- Files: the paths the change may touch. Anything else needs a reason in the hand-off.
- Tests: a table with one row per behaviour, each with its exact test name.
- Shares: which other tickets touch the same files, so two builds never collide.
Every build, AI-assisted or not, starts by restating the scope in two sentences. If the restatement and the file disagree, the file wins. Nobody reads other tickets “for context”. If the ticket is missing something, that is a reason to stop, not to explore.
A ticket also has a size limit. In the game, a docs check refuses a ticket that names more than five paths. We once wrote a ticket that named seven. It was split in two, one for the rules engine and one for the server, and the split went into the decision log before any code was written.
2. Red, then green
The tests come first, one per row, named exactly as the row says. They are committed on their own and run. They must fail.
Then comes the smallest change that makes them pass, touching only the named files. Then the ticket’s check runs: formatting, lints, a scan for slop, and the ticket’s own tests.
That check fails if no test carries the ticket’s id. A green run with zero tests isn’t green. We also read passing output for the counts we expect. “Nothing failed” and “the checks ran” are different claims.
3. Frozen tests
When a test file reaches main, its rows are frozen. Frozen tests sit on a list of protected files. The merge script refuses a branch that modifies, deletes or renames one, unless a commit on the branch carries an approval line that only the owner writes.
This sounds rigid. It is meant to be. Any build that is stuck, by a person or with a tool, has an easy way out: change the test so the code passes. Freezing closes that door.
Behaviour does change, of course. When it does, the ticket retires the old rows by name, on a “Retires” line, and adds new ones. It never edits an old row to mean something new. The record of what the game promised stays readable.
4. Guards that can fail
Every guard in the toolchain, from the slop scan to the merge checks, has its own test. That test feeds the guard a violation and asserts that it refuses. Change a guard, and you run those tests.
The rule behind it: a guard that cannot fail is worse than no guard, because it looks exactly like success.
5. Loop budgets
Each ticket has a time box across all attempts: five attempts, five dollars and two hours. An attempt ends when the ticket’s check runs.
When an attempt fails, we step the tooling up. A ticket starts on a cheaper model, moves to a stronger one, then to the strongest at higher effort.
When the time box runs out, the work stops. What blocked it and what was tried go on the ticket, and an engineer reads that. Often the fix is a clearer ticket, not more attempts.
6. Stop and ask
The work stops and we ask when:
- a row needs a fact nobody wrote down,
- two rules conflict,
- a finished dependency doesn’t provide what the ticket assumes,
- the change needs a protected file,
- the budget is gone.
Guessing costs more than asking. When a build does make a small call instead of asking, such as a default or a limit, it writes it on the ticket under “Calls made”, with how to undo it.
7. Style that keeps slop out
Code generators like to add things, and so do tired engineers. Our rules push the other way:
- No defensive code for cases the ticket doesn’t name. Raise them instead.
- No feature flags, settings or abstractions the ticket doesn’t ask for.
- No folders called
utils,helpersormisc. Name the folder after what it does. - No TODOs. A deliberate shortcut gets one comment: what was left out, its limit, and when to add it.
- Facts about game items live in data, not in code that names one item.
8. One merge owner
At most two builds run at once, on tickets that don’t share files. Each works in its own worktree, on a branch named after its ticket.
Only one session merges. Each build pushes its branch and queues it. A merge daemon on a separate machine takes branches one at a time. It refuses a branch that is behind main, carries AI attribution in its commits, or touches a protected file without approval. Then it runs the pre-merge suite on that exact commit, in a fresh worktree, and fast-forwards main.
Nothing reaches main another way. That gives us a linear history, and one place to look when something breaks.
9. Review
Engineers set the rules, and an engineer reviews every change against its ticket, not against their own idea of good code. Branches that touch a protected file, or that needed the strongest tooling, also get an AI-assisted second read before the merge. Visual changes are checked by a person looking at them. Approval only ever comes from the owner, never from a tool.
What it costs, and what it buys
This is more process than most small teams run. Precise tickets take time to write. Frozen tests mean some changes need an approval and a new ticket.
What it buys is trust in speed. When a branch merges, we know which rows it proved, that those rows failed first, that an engineer reviewed it, and that nothing older was quietly rewritten. That is what lets us use AI tools to go fast without handing them the decisions.