Agentic AI Compliance for Law Firms: What a 36% Score Means

Agentic AI compliance infographic showing a benchmark comparison chart with a 36.2% best-performing AI agent pass rate versus under 25% for most frontier AI models, alongside compliance checklist, security shield, and AI governance icons.

There’s a number worth sitting with before your firm signs off on another agentic AI pilot: 36.2 percent.

That’s the best score any AI agent configuration managed on a new benchmark called HANDBOOK.md , published on arXiv, July 28 by a team that includes researchers from Surge AI. Most frontier models didn’t even get that far; they landed under 25 percent. The test itself is almost too boring to describe- drop an AI agent into a simulated company, give it a real employee handbook (some ran 124 pages), and ask it to do routine work while following the rules written down in front of it. Finance tasks. Medical billing. Insurance claims. Logistics. HR. Sixty-five tasks total, across ten fictional companies, graded against 824 separate pass or fail criteria checking not just what the agent did, but what it was supposed to avoid doing.

 

I think that second part is the one worth pausing on. This isn’t a benchmark about whether AI can draft a memo or summarize a contract. It’s about whether an agent, once you hand it a policy document and some autonomy, actually stays inside the lines you drew.

What HANDBOOK.md Actually Tested

The setup is closer to a real workplace than most AI benchmarks bother getting. Each agent had access to files, a simulated inbox, chat, a calendar, an issue tracker, the kind of tool sprawl your own staff deals with daily. Then it got a dense, expert written standard operating procedure and one instruction: follow the company’s rules while doing the work.

Under strict grading, the best of thirty tested model configurations passed 36.2 percent of trials. Most frontier configurations, meaning the models’ firms are actually deploying right now, stayed below a quarter.

Some of the specific failures are the kind you’d flag in a heartbeat if a junior associate did them. Agents fired employees without authorization. They cleared expense reports that the same agent had submitted itself. They forwarded expired medical records to insurers rather than flagging the expiration and stopping. None of these were edge cases buried in the fine print, they were straightforward violations of policies the agent had direct access to read.

Why Frontier Models Keep Overriding Their Own Rules

Two failure patterns stood out in the research, and both should sound familiar if you’ve watched an AI tool operate for more than a few minutes.

 

The first: agents would override a standing policy the moment a user pushed back in the conversation. Someone says “just this once” or reframes a request slightly, and the agent drops the rule it was following seconds earlier. Which, honestly, sounds a lot like how a poorly trained employee might behave under social pressure, except this employee never gets tired of being asked and never reports the pattern to anyone.

The second: agents lost track of rules over long tasks. The handbook might be 100 plus pages, and somewhere around page 40 the agent stops applying what it read on page 12. Long context windows get marketed as a strength. This benchmark suggests they’re also where a lot of the actual governance risk lives.

What This Means for Agentic AI Compliance in Law Firms

Here’s where it gets pointed for a law firm specifically. Adoption of agentic tools inside legal has moved fast this year. Intapp’s Celeste went generally available. Thomson Reuters rolled out its next generation CoCounsel with an agentic brief drafting tool built in. Harvey keeps expanding what its agents can touch. None of that is inherently bad, I’m not arguing firms should freeze and wait. But agentic AI compliance in a law firm setting isn’t a marketing checkbox, it’s the thing that determines whether a conflict check actually ran, whether a billing entry got the right narrative, whether client intake screening followed the ethical wall it was supposed to respect.

 

A 36% best case pass rate on a general enterprise benchmark doesn’t map one to one onto legal specific tools, and it would be unfair to claim otherwise. But it’s the first hard, quantified signal that “the agent followed our policy document” is not something you should assume just because you handed the agent a policy document. Vendors will tell you their tool is different, better tuned, purpose built for legal workflows. Maybe some are. The point of a benchmark like this is that you shouldn’t take that claim on faith, you should ask for evidence.

Building Real AI Risk Mitigation Controls, Not Just Policies

We’ve written before about why AI risk mitigation for law firms has to go beyond a written AI use policy sitting in a binder somewhere (our piece gets into this). A policy document is only as good as the system’s actual behavior when nobody’s watching closely, and this benchmark is a pretty direct illustration of that gap.

 

What seems to actually work, based on what’s held up so far, is layering in checkpoints the agent can’t talk its way around. Human approval gates before anything with financial or client impact goes out. Logging that captures not just what the agent did, but what it was told to do and whether the two match. Scoped access, so an agent handling document review doesn’t also have standing permission to send client communications or approve payments. None of this is exotic. It’s the same least privilege thinking that’s shaped IT security for two decades, just applied to a new kind of user that happens to be software.

Where to Start This Week

If your firm has agentic tools in pilot right now, or is about to, a few places worth looking at first:

  • Map what each agent can truly touch versus what it should be able to touch. We covered the access side of this in our piece on AI agent security for law firms, and it’s a shorter audit than most firms expect once you sit down and list it out.

 

  • Check whether “shadow” AI tools, meaning one’s staff picked up on their own outside of anything IT approved, have the same gaps this benchmark found in sanctioned tools. Our shadow AI protection piece walks through how to find those before a client or regulator does.

 

  • And ask your vendors directly whether they’ve run anything resembling this kind of adversarial, long context compliance test, rather than just a demo that shows the happy path.

 

None of this means agentic AI doesn’t belong in a law firm. It clearly has real upside, and the firms sitting this out entirely are going to fall behind the ones using it well. But “well” must include actually verifying the thing follows your rules, not assuming it does because the sales deck said so.

Ready to See Where Your Firm Actually Stands

If you’re not sure whether your current AI deployments would hold up under a test like this one, that’s worth finding out before a client or a regulator forces the question.  Our Zero Trust Assessment looks at exactly this kind of gap, what your AI tools and agents can access versus what they should, and gives you a clear picture of where the real exposure sits.

Recent Posts

Have Any Question?

Call or email Cocha.  We can help with your cybersecurity needs!

About the Author:

Picture of Steve Combs

Steve Combs

Co-Founder & Managing Director, Cocha Technology

Steven is a fractional CIO/CISO with 30+ years of enterprise IT and security leadership. He has built AI governance frameworks for organizations with 1,700+ users, led enterprise Microsoft Copilot deployments, and conducted security assessments across law firms, energy companies, financial institutions, and PE-backed manufacturers.