I closed out a recent post on agent terminology with a fourth term that keeps showing up next to “agent”: harness. It deserves its own post, because “harness” names a specific idea that’s easy to blur into “agent setup.”

Agent equals model plus harness

The shorthand comes from LangChain’s anatomy of an agent harness: Agent = Model + Harness. The harness is everything except the model’s actual reasoning — the execution loop, the tools it can call, the checks on what it produces, the decision about whether to try again or stop.

The split exists because every call to an LLM is stateless. Anthropic’s Messages API is stateless by design. You own the conversation history and resend it with every request, because the model has no memory between calls. OpenAI’s Responses API can look stateful (store: true, resume with previous_response_id). But that’s the vendor storing your context server-side and replaying it back; the model itself still sees only what is sent with each call. Left alone, a stateless model can’t do anything that spans more than one turn. So reduce a harness to one sentence and it’s this: a harness decides what’s in the model’s context window for any given call. Everything below is a variation on that.

Birgitta Böckeler’s harness engineering writeup for Thoughtworks gives the idea real vocabulary, and both halves of it turn out to be context management wearing different names. Guides (feedforward) are context added before the model acts: conventions, specs, reference docs. Sensors (feedback) are context added after: linters, test failures, review comments, fed back in for the next turn. Each can be computational (deterministic, cheap, run by a CPU — a type checker) or inferential (semantic, expensive, run by a model — an AI code reviewer). Böckeler’s own words for the relationship: “engineering a user harness for a coding agent is a specific form of context engineering.” Harness richness and autonomy are separate axes: autonomy is how much control you hand the model, and the harness is what largely determines how much you can trust what comes out.

Böckeler also makes a point worth stealing: harnesses nest. There’s the inner harness a vendor builds into their coding agent (system prompt, retrieval, orchestration), and there’s the outer harness you build around it for your own repo. GitHub Copilot ships with a harness already inside it. Your AGENTS.md, your pre-commit hooks and your custom lints form a second harness wrapped around the first.

Memory is just delayed context

A model has two sources it draws on directly: its weights, and whatever’s in the current context window. That makes memory the same context-management job as everything above, just stretched across time instead of held within one call. “The agent remembers your preferences” describes the harness, not the model: it writes state to a file, then re-injects that file into context the next time it starts. AGENTS.md (the cross-vendor convention) and CLAUDE.md (Claude Code’s own) are exactly this: a memory file the harness loads automatically at session start.

The job also runs in reverse: trimming, not just adding. Context rot is what happens when a model gets worse as its context window fills up. Harnesses fight this with compaction (summarizing and offloading older context before it overflows) and tool-output offloading (keeping only the head and tail of a noisy tool result, writing the rest to disk where the model can fetch it back if needed). Deciding what stays, what gets summarized, and what gets evicted is as much the harness’s job as deciding what goes in. Claude Code’s skills, below, are one direct answer to it.

The range, in practice

OpenAI’s Codex team built the heaviest harness I’ve seen documented. They shipped an internal product with zero manually written lines of code, roughly a million lines, five months, a team that started at three engineers and grew to seven. A docs/ directory serves as the system of record instead of one giant AGENTS.md; in their words, “give Codex a map, not a 1,000-page instruction manual.” A layered architecture keeps dependency directions enforced by custom linters and structural tests. A full observability stack is wired directly into the agent, so it can query its own logs and metrics with LogQL and PromQL. And a recurring “garbage collection” pass has background agents scan for architectural drift and open their own cleanup PRs. The team’s own framing: humans steer, agents execute, and the engineering work moved almost entirely into the scaffolding.

Claude Code’s harness is the most explicitly modular one I’ve seen a vendor document. Anthropic breaks it into seven distinct pieces. CLAUDE.md handles always-on project memory; rules hold path-scoped constraints; skills are reusable procedures that stay out of context until actually invoked; subagents delegate work into their own isolated context window; hooks fire on lifecycle events (file edits, tool calls, session start) to run deterministic checks a model shouldn’t be trusted to self-police; output styles and system-prompt appends handle finer control. Skills in particular answer the context-rot problem described earlier: only the short description loads at session start, and the full instructions load on demand.

GitHub Copilot ships several harnesses under one name, and they differ mainly in where the sensors live (the surfaces are laid out in job two of the agents post). When Copilot runs locally, in VS Code’s Agent Mode or the CLI, you’re the sensor: you watch the output and interrupt when it drifts. The cloud agent runs unattended in an ephemeral GitHub-hosted environment, so the harness has to supply its own sensors: it executes tests and linters, and as of mid-2026 does a self-review pass before opening the PR, leaving CI and a human reviewer as the final gate. Custom agents (.agent.md files, usable from both the IDE and the CLI) are the guide side: a system prompt and a tool allowlist.

On the memory side, Copilot’s code review and VS Code’s agent mode read a repository’s AGENTS.md, the same convention OpenAI’s Codex uses (Claude Code reads its own CLAUDE.md), so the outer harness you build around the repo carries across surfaces.

Stripe’s Minions run a different harness at higher volume. Stripe’s homegrown coding agents merge more than a thousand PRs a week. The final gate is a human: every change is reviewed before it merges, on top of CI runs. Speed comes from the front half of the loop (agents write start to finish, one-shot), and the check at the end stays. That’s a different bet from OpenAI’s: trust the harness for generation, keep a human as the final gate.

In my experience, most teams get the inner harness for free and build very little of the outer one. Claude Code, Copilot and Codex all ship a capable harness out of the box, and everyone using them gets the same one. What differs is what you wrap around it: for most repos, that’s an AGENTS.md file, whatever linters CI already runs, and maybe a custom review prompt. Plenty of guides and almost no sensors, or the reverse, and little of it shared across the team. Böckeler’s term for what limits you here is harnessability: not every codebase affords the same controls. A strongly typed language gives you type-checking as a sensor for free; a codebase with fuzzy module boundaries can’t get architectural fitness checks nearly as cheaply. The harness you can build is partly a function of decisions made years before anyone was harnessing anything.

Building the outer harness pays off in three ways. Speed: the standing context is already assembled, so what you type can be a short instruction instead of a re-explanation of your conventions every time. Consistency: your team’s own patterns are written down where the agent reads them, so output tracks how you build things rather than the model’s defaults. Token use: a short map plus skills that load on demand costs far less than pasting the same background into every prompt, and a linter or hook that catches a mistake costs almost nothing to run, where a model re-checking its own work spends tokens to do it. The last one cuts both ways. A bloated always-on instruction file spends tokens on every call and brings the context rot back, so the outer harness has to stay lean.

Part of your outer harness doesn’t have to be written by you: community assets. Collections like Addy Osmani’s agent-skills (skills, slash commands, reviewer personas, and checklists that install across Claude Code, Copilot, Cursor, and others) give a team a reviewed starting point instead of a blank AGENTS.md. They deserve the same scrutiny as any dependency, though. You’re putting someone else’s instructions into your model’s context, and any hooks they ship run code on your machine.

The terms, and who owns them

Here’s how the pieces above sort into responsibilities. Each one is split between the inner harness (what the vendor ships, the same for everyone) and the outer harness (what you build around it for your repo and team).

Responsibility What it means Inner harness (vendor) Outer harness (yours)
Context assembly Deciding what’s in the window for each call System prompt, retrieval, orchestration Which docs and files you point it at
Guides (feedforward) Context added before the model acts Built-in system prompt and tool descriptions AGENTS.md / CLAUDE.md, rules, skills, custom agents (.agent.md), community assets
Computational sensors Deterministic, cheap checks run by a CPU Tests and linters run by the cloud agent (Copilot) Your linters, type checker, structural tests, CI, hooks
Inferential sensors Semantic checks run by a model Self-review pass, Copilot code review Custom review prompts, reviewer personas, review subagents, scheduled cleanup agents (Codex)
Memory State written to disk and re-injected later Loading the memory file at session start; chat-app memory (Dreaming, Claude memory, Microsoft 365 Copilot) The memory files that live in your repo
Context trimming Fighting context rot Compaction, tool-output offloading Short maps instead of manuals, skills that load on demand
Final gate Who decides the work is acceptable None; the vendor hands off to you CI plus human review (Stripe, Copilot cloud agent)

Where does a chat app fit into this?

Everything above is a coding-agent harness, built for long-horizon work. The same vocabulary applies to the desktop and web chat apps most people actually use every day: ChatGPT, Claude, Microsoft 365 Copilot. They’re harnesses too, just built for a different job, and almost entirely inner harness: you can’t wrap your own around them the way you can around a coding agent, which is much of why the coding tools are where harness engineering happens.

A chat app’s harness stays turn-by-turn by design: a human reads every response and decides what happens next, so autonomy stays low no matter how good the harness gets underneath. Harness richness and autonomy are separate axes, as noted above. But real harness engineering still shows up in three places: tools (web search, code execution, file analysis), safety layers, and, increasingly this year, memory.

ChatGPT, Claude, and Microsoft 365 Copilot all now have dedicated memory systems. OpenAI’s Dreaming (June 2026) curates memories from chat history into an evolving, editable summary. Claude’s memory went free for all users on March 2, 2026, and it’s a plain text file of inferred preferences and context that you can open and edit yourself. Microsoft 365 Copilot has had memory since 2025, and separately grounds its answers in a user’s files, meetings, and email through Microsoft Graph. That’s retrieval rather than memory, but it fills the same slot in the context window.

Memories don’t carry over automatically: what ChatGPT learns about you doesn’t reach Claude (Claude does offer a memory import), and what Copilot learns about your working style stays inside Microsoft 365. It’s the same mechanism as everywhere else in this post, state on disk re-injected into context, just kept per vendor instead of in a repo you control.

Why the range matters

Put a scrappy harness and OpenAI’s harness around the same model on the same class of task and I’d expect very different results, even though “agent” and “harness” apply equally well to both. It’s the same trap as the vocabulary problem in the agents post: two people can say “we have an agent for that” and mean setups that aren’t in the same league of trustworthiness.

It also shows where OpenAI’s effort went. Their effort went into guides (the docs map, AGENTS.md), architectural constraints, sensors (linters, structural tests, an observability stack the agent could query), and a cleanup loop that runs without being asked, rather than into better prompts.

The next question is how to build an outer harness the whole team shares, instead of one each developer improvises. That’s its own post.

If “prompt engineering” was the 2023 skill, harness engineering (deciding what belongs in a model’s context for any given call, and building the machinery that puts it there) looks like the one that compounds.


Sources checked October 2026. Agent = Model + Harness framing, and the memory/context-rot mechanics, from LangChain’s anatomy of an agent harness. Guides/sensors, computational/inferential, and the “harness engineering is a form of context engineering” framing from Birgitta Böckeler’s Harness engineering for coding agent users (Thoughtworks, April 2026). API statelessness from the Anthropic Messages API and OpenAI Responses API docs. Codex details from OpenAI’s Harness engineering: leveraging Codex in an agent-first world (February 2026). Claude Code’s component breakdown from Anthropic’s Steering Claude Code blog post. GitHub Copilot’s cloud agent from GitHub Docs, and AGENTS.md support in code review from the GitHub Changelog. Stripe Minions details from Minions: Stripe’s one-shot, end-to-end coding agents (February 2026). ChatGPT’s Dreaming memory from OpenAI’s announcement (June 2026); Claude’s free-tier memory rollout (March 2, 2026) and Microsoft 365 Copilot’s memory, grounded via Microsoft Graph, per contemporary reporting.