Industry Insights

Harness Engineering: What It Changes About Legal AI for In-House Teams

Last updated:
October 8, 2026
Written by:
Vanessa Davis, J.D.
,
Chief Product Officer

For the last few months, I've watched legal AI tools the way an engineer watches a coding agent. Not just what the output looks like, but what happens around it before anyone trusts it. A coding agent can write code that compiles, runs, and looks clean, and still be wrong in ways that only show up later. Engineers learned not to judge the AI's output alone, but the system built around it: the tests, reviews, and guardrails that catch problems before they ship. There's a name for that system now. Software teams call it the harness, and building it well is harness engineering.

What harness engineering is

At a high level, an AI agent system can be understood as a model (or set of models), while everything around it is the harness. Writing on Martin Fowler's site, Thoughtworks' Birgitta Böckeler defines the harness as the instructions, tools, context, and checks that decide how the model behaves and how it catches itself when it's wrong. 

Harness engineering is the process of building and improving everything that defines how the model behaves and how it arrives at a reliable and accurate output.

OpenAI's engineering team learned this by building an internal product with no hand-written code. Their takeaway: humans steer, agents execute. Their hardest problems weren’t about the model, they were about the harness. 

Böckeler organizes the harness into two kinds of controls: guides and sensors. Guides steer the AI before it acts, so it starts in the right place. Sensors check it after, so it can detect its own mistakes before you see them. The human's job isn't just to write clever prompts. It's to notice when the system makes the same mistake twice, then fix the guide or the sensor so it stops. A guide can be as simple as a project instruction. Building sensors that detect errors is where a purpose-built legal AI tool earns its place in an attorney’s tech stack.  

Why harness engineering matters for legal AI

Harness engineering is the difference between a general chatbot that produces an indemnification clause that merely appears right, and a purpose-built legal AI tool that uses attorney-vetted guides and sensors to steer and check its work. Ask a general-purpose AI tool what standard it applied and what positions it was steering toward, and you get either a shrug in paragraph form or an overly confident answer. 

A general-purpose chatbot has a capable model, but no legal harness you can trust across a team. You can bolt a harness on (a well-engineered ChatGPT prompt, a project instruction, a custom Claude skill, even Claude for Legal's practice-area plugins) but that harness is only as good as the person who built or configured it. It doesn't travel to the rest of your team, and the output still varies from run to run.

Purpose-built legal AI is where harness engineering comes into play. It wraps that same frontier model in attorney-built guides and sensors, so the output reflects a real legal standard before a lawyer sees it.

Harness engineering is even more important for legal work than coding. Unlike legal teams, software developers have deterministic feedback mechanisms. Bad code often fails a test or throws an error. Legal work isn’t like that. You can't instantly test whether a liability cap is reasonable for your counterparty on a specific deal. That is where harness engineering can help put the model to work for you before a lawyer has to step into the loop (and a lawyer always should). 

Harness engineering in practice: contract review

Out of all in-house use cases, harness engineering is most mission-critical in contract review, where a confident wrong answer on the wrong clause becomes real legal and financial risk.

LegalOn's 2026 Contract Review Benchmark  put 11 leading AI models head-to-head against LegalOn across 3,282 contract reviews on 21 precision-critical provisions. An independent judge, blind to which system produced each review, scored each output for accuracy, completeness, and usefulness. LegalOn ranked first on all 21 provision types. The models it beat were the strongest general-purpose systems from Anthropic, Google, and OpenAI — the same frontier models millions of people use every day.

Handed a contract, a general-purpose model reads the whole thing in one pass and returns a verdict. LegalOn structures the job differently, running 21 provision-level checks in parallel, each isolated to a single guideline, catching the detail a single-pass read misses.

This is harness engineering in action. Used directly, general-purpose models can identify general topics but miss the detail. LegalOn uses these same models, but engineers the harness to produce outputs that the independent LLM judge preferred up to 79% of the time.

Accuracy results from the 2026 Contract Review Benchmark

Three ways harness engineering can improve your legal AI work

1. Treat your playbook as your feedforward guide

In harness engineering, guides steer the AI before it generates a single response.

The closest legal analogue to a guide already exists: the contract playbook. A playbook steers the AI before it reviews, taking it toward your preferred positions and your fallbacks. The clearer the guide, the better the first draft.

The practical move is to get those positions out of senior lawyers' heads and into a form the tool can act on. In LegalOn, playbooks are written in plain English, so the standard you'd apply by hand is the standard the AI reviews against, instead of one it improvises. A good guide means the redlines come back closer to how you would have marked them up yourself on the first pass.

2. Build a feedback loop with sensors

A sensor does the opposite job of a guide: it detects issues after the model is finished producing an output. Engineers have a rule. Catch problems as far to the left as possible, because that's where they're cheapest to fix.

A legal team can do the same with a purpose-built tool that provides sensors. These run checks after the AI drafts and flag what looks wrong, so the problem surfaces while you can still fix it. LegalOn Review does this against an attorney-built library of 10,000+ legal issues across 50+ contract types, surfacing the missing cap, the one-sided indemnity, or the off-market term while there is still time to negotiate it.

That said, a sensor makes a judgment call, not a final ruling. A lawyer must still review it. It might catch that the indemnity runs one way, but it can't tell you whether that's the right call for this deal. Only the lawyer can do that.

3. Fix the system, not the prompt

After an AI-generated error, the first instinct is to fix the individual prompt. Harness engineering does better: it shapes how the model behaves before a prompt enters a chat window.

Fixing the harness starts with what OpenAI's team calls legibility. What the agent can't see doesn't exist. A negotiated position that lives only in a senior lawyer's memory, a Slack thread, or a PDF nobody opens cannot steer the AI. Encoded into a playbook, it can.

Every time you refine a playbook after a miss, you are moving from marking up individual contracts to shaping the system that marks them up for you.

What harness engineering means for in-house legal teams

In-house counsel should stop judging  legal AI by the model alone and start judging the harness around it. Ask what steers the AI before it acts, what checks it after, and how easily you can improve both when something goes wrong.

LegalOn is built on that premise. Unlike the frontier models, LegalOn runs with attorney-built guides and sensors, so your team spends less time cleaning up first drafts and more time on the calls only a lawyer can make.

To see what an attorney-built harness looks like in practice, book a demo.

Related Posts

View all
Industry Insights
September 25, 2026
Construction Contracts: What to Include and How to Review One
Industry Insights
September 23, 2026
Assogestioni's Roberta D'Apice on AI, Regulation, and the Harder Question of Legal Value
Industry Insights
September 11, 2026
Best Automated Contract Review Software Tools of 2026
View all

Experience LegalOn Today

See how LegalOn can save you time, reduce legal risk, and free you from tedious work.
Book a Demo