Harness engineering is the work of shaping the environment around an AI agent: instructions, tools, permissions and checks. Harness-Driven Development (HDD) is a working method built on it: every agent mistake becomes a rule of that environment.
In this exampleA prompt fixes one answer, context fixes what the model knows, and the harness fixes every future agent run.
In generalHarness-Driven Development is work at the third stage: the developer writes less code and spends more time setting up the environment where the agent writes, checks and fixes it.
Next stepWrite down the agent's last three mistakes and mark the stage at which each could have been fixed for good.
The HDD loop
Mitchell Hashimoto's principle: when the agent makes a mistake, change the harness so that mistake can't happen again Faros
Takeaway
In this exampleA mistake caught by the sensors goes into the harness as a new rule, and the agent gets that rule on every following run.
In generalA result fixed by hand helps once. A rule in the harness (an instruction, a test or a hook) applies to all future tasks, so the harness gets more precise over time without changing the model.
Next stepAdopt a rule: any agent mistake you had to fix by hand becomes a line in the instructions or a test on the same day.
How a harness is laid out: 3 × 3
Three domains of the loop and three contexts in each — nine canons. Every canon has exactly one domain and one context.
→skill → spec of its domain·a commit hook script checks the direction
snapshota table, updated after each new measurement
livea protocol, fresh every time, nothing stored
Takeaway
In this exampleThe loop splits into Expectation, Observation and Loop, each with its own spec, skill and memory. Loop reads the other two domains; they never read Loop.
In generalIt is a negative feedback loop: the target value is compared with the measured one, and something changes only on a mismatch. A strict reference direction keeps the domains from tangling, and the same concept repeating in two domains is not a defect.
Next stepSort the files of your harness into the nine cells and check that no Expectation or Observation file refers to Loop.
The double loop
single loopfix the action under the same rules
double loopchange the rules themselves
Patches without a new predictionfix it with a rule
Several complaintsfirst test for a shared cause
A new ruleaccepted only if old cases follow from it
No external checkdiscard it, with a record
The loop applies to itself: every run ends with the residual risk, never with a “done” status.
Takeaway
In this exampleA series of patches, complaints, a new rule and an external check are four checkable principles the Loop uses to decide whether to fix the action or change the rules.
In generalA single loop fixes the result within the old rules. A double loop changes the rules themselves, and does it by checkable principles, so the harness learns without drifting from random edits.
Next stepWhen the same fix comes up a third time, stop and ask which rule causes it. Change the rule only if it also explains the old cases.
In this exampleFour cells: guides set the course before the agent works, sensors catch deviations afterwards, and each can be computational or AI-based.
In generalComputational checks run in seconds and don't make mistakes, so use them wherever a rule can be written down formally. Keep AI-based checks for meaning: style, clarity, fit with the intent.
Next stepFor each rule in the agent's instructions, ask whether a test or linter could check it. If so, move it there.
In this exampleCode quality is the easiest to check today, and product behavior the hardest.
In generalLinters and formatters cover the first layer almost immediately. Architecture is checked with dependency rules and performance requirements. Behavior still needs good scenario tests and a human eye.
Next stepStart with the first layer: run a linter and formatter on every agent run, then add one rule for module boundaries.
Earlier is cheaper
Before workinstructions · specs$
On writehooks · linter$$
In CItests · build$$$
At reviewa person$$$$
Takeaway
In this exampleThe same mistake costs least before the agent starts and most at human review.
In generalPut fast checks as early as possible: the agent fixes the problem itself while the task context is still in front of it, and only the debatable points reach a person.
Next stepWire tests and the linter into a hook the agent runs after every edit, in addition to CI.
Where people matter
Harness
style · types · tests · boundaries
Person
architecture · product · new rules
"A good harness keeps the human in the loop and directs their attention to where it matters most" — paraphrasing Martin Fowler
Takeaway
In this exampleThe harness takes most of the checks; the person keeps the decisions and the work of improving the harness.
In generalAn experienced developer carries an implicit harness in their head: knowledge of the architecture, team habits, a feel for risk. HDD moves that knowledge into explicit rules, so every agent run gets it.
Next stepAfter each review of the agent's code, write down one comment that keeps coming up and turn it into a rule.
In this exampleThe spec answers "what are we building", the harness answers "how do we know we built it".
In generalWithout a spec the harness has nothing to check, and without a harness the spec stays a document. Together they let you accept the agent's code based on checks instead of reading every line.
Next stepFor each requirement in the spec, write down which test or check will confirm it is met.
Where to start
1
An instructions file for the agent
AGENTS.md or CLAUDE.md in the project root: how to build, how to test, what not to do
2
Linter and tests after every edit
a hook the agent runs itself before handing over the result
3
An agent mistake log
every mistake fixed by hand gets a line with its cause
4
Mistake → rule
once a week, move log entries into instructions, tests or hooks
5
A progress file for long tasks
a feature list with statuses and the commit history, so a new session picks up where the last one stopped Anthropic
Takeaway
In this exampleFive steps: two set up the harness right away, three start the "mistake → rule" loop.
In generalNobody designs a harness in full up front. It grows out of the agent's real mistakes in your project, and within a few weeks the agent stops repeating the most common ones.
Next stepToday, create the instructions file and the mistake log. In a week, move the first log entries into rules.