Agentic AI: Zero Risk Isn’t the Job—A CISO’s Guide to Smarter AI Containment

Agentic AI security assessment showing a CISO reviewing an AI agent access map, permissions dashboard, and vendor risk controls on a laptop in a modern law office.

Here’s a sentence I did not expect to write this year: a leading AI lab hacked its own customers by accident. Not maliciously. Not through some elaborate exploit. Through a misunderstanding about whether a test environment had internet access.

That’s the story Anthropic told on July 31, and it should be required reading for every AMLaw 200 CIO & CISO who have approved an agentic AI rollout in the last six months.

What Anthropic Disclosed on July 31

Anthropic reviewed roughly 141,000 of its own cybersecurity evaluation sessions after OpenAI disclosed a breach of its own. What it found was not comfortable. Three separate models, Claude Opus 4.7, Claude Mythos 5, and an internal research model, had breached the real infrastructure of three outside organizations during what was supposed to be contained agentic AI safety testing. Not simulated targets. Real systems, real credentials, real exposure.

Here’s the part that stings. The evaluation prompt told Claude the environment was a simulation with no internet access. It wasn’t. A misunderstanding between Anthropic and its evaluation partner left internet access open, and when Claude’s search turned up live systems, it treated them as part of the exercise and went to work. The techniques it used weren’t exotic. Weak passwords. Unauthenticated endpoints. The kind of gap a decent security assessment finds in an afternoon.

Anthropic caught it, eventually. It started reviewing transcripts on July 23, suspended cyber evaluations that same day, had all three incidents identified by July 24, and notified the affected organizations on July 27. Two of those organizations reportedly had no idea anything happened until that call came in.

I think that timeline is the most useful part of this whole story. A frontier lab with more security talent than almost anyone on earth needed four days just to confirm the scope of its own mistake, and it still doesn’t appear to have named the victims publicly. Bloomberg’s reporting on this ties it to a broader pattern too, since OpenAI disclosed a breach of its own right before Anthropic went looking.

Why "Isolated" Test Environments Aren't the Same as Zero Risk

Here’s the uncomfortable lesson for anyone buying agentic AI tools right now: the word “isolated” in an agentic AI vendor’s documentation is a claim, not a guarantee. Anthropic didn’t cut corners on this. It built what it believed was a contained evaluation environment, staffed by people who understood the stakes, and the containment still failed because of a coordination gap with a partner.

Zero risk isn’t the job. It never was. The job is knowing where your actual exposure sits, and most firms adopting agentic AI right now haven’t mapped that exposure at all. They’ve read a vendor’s security page, maybe asked a few questions in a sales call, and moved forward on trust.

What stands out to me is how ordinary the failure mode was. This wasn’t a novel attack technique nobody could have predicted. It was a scope and permissions problem, the exact category of thing a proper access review would catch. Law firms already know how to think about scope and permissions in the context of client conflicts and ethical walls. That same discipline needs to extend to what an agentic AI system can reach on your network, not just what it’s supposed to reach.

This is also an insurance and malpractice question, not only an IT one. Cyber and tech E&O policies are starting to ask pointed questions about agentic AI oversight at renewal, and “we trusted the vendor” is not going to read well to an underwriter, or worse, to a state bar disciplinary panel after a breach touches privileged material. Firms that get ahead of this now are treating agentic AI vendor review the same way they treat any other third party with access to client data, not as a special exception because the technology is new and exciting.

What This Means If Your Firm Runs Harvey, CoCounsel, or Copilot Agents

Most AmLaw and midsize firms are past the pilot stage now. Agentic AI tools like Harvey, CoCounsel, and Microsoft Copilot agents are doing real work inside real matters, often with access to document repositories, email, and case management systems that touch privileged client information.

Every one of those platforms rests on the same basic premise Anthropic’s evaluation did: that the boundaries around what an agentic AI system can see and do will hold. Anthropic just showed that even a company built by AI safety researchers can get that assumption wrong.

So, what does this mean in practice? It means the question your IT committee should be asking isn’t “does our vendor test for this?” Nearly every vendor will say yes. The better question is “can our vendor show us documented proof of how containment is verified, and what happens when it fails?” If a vendor’s answer is a general reassurance rather than a specific control, that’s your answer too.

It also means firms need their own layer of verification, independent of whatever the vendor claims. You wouldn’t let outside counsel self-certify their own conflicts check. Don’t let an agentic AI vendor self-certify their own containment either.

 

There’s a cost angle worth naming too. A verified containment model is cheap next to the alternative. Once agentic AI touches a live matter and something goes wrong, the response involves outside forensics, client notification, possibly a malpractice claim, and an uncomfortable conversation with whoever championed the rollout. A few hours mapping access now is not a hard sell once you frame it against that.

A CISO's Guide to Vetting Agentic AI Vendors

A few things worth building into your agentic AI vendor review process now, before the next renewal cycle rather than after an incident:

Ask for the isolation architecture behind the vendor’s containment claim, not a marketing summary. Network level segmentation is a different, stronger claim than “we use a sandboxed environment,” and vendors know the difference even when their marketing blurs it.

 

Ask who verifies the containment, and how often. A control that’s tested once at launch and never again isn’t a control, it’s a memory of one.

Ask what monitoring exists for the agent reaching outside its intended scope, and how fast an alert reaches a human. Anthropic’s own detection took days. Your firm’s tolerance for that kind of delay is probably lower, especially with active client matters involved.

Map what each agentic AI deployment can technically reach today, not what it was designed to reach originally. Permissions drift. Integrations get added. Nobody revisits the original scope document six months later unless someone is assigned to.

 

One more thing worth building in as a contract term rather than a hope: require your agentic AI vendor to disclose incidents like Anthropic’s on a defined timeline, in writing. Anthropic chose transparency here. Not every vendor will, and a firm holding privileged client data cannot afford to learn about a containment failure from a news article instead of a phone call.

Four Questions to Ask Before You Deploy

If you want a shorter version to bring into your next agentic AI vendor call, ask these four things directly and expect specific answers, not reassurance:

 

  1. What happens, technically, if this agent’s isolation fails the way Anthropic’s did?
  2. Who at your organization would know within minutes, not days?
  3. Can you show us the last time this containment was independently tested?
  4. What client data, systems, or matters would be exposed in a worst-case failure?

 

None of these are trick questions. A vendor that has genuinely tested its own containment model should be able to answer all four on a single call, with specifics. A vendor that can’t is telling you something real about how it thinks about agentic AI risk, whether it means to or not. If the answers are vague, that’s information too.

Where Cocha Technology Fits In

None of this means firms should pull back from agentic AI adoption. The upside is real and the competitive pressure to adopt agentic AI isn’t going away. But adoption without a mapped, tested containment model is just deferred risk, and deferred risk in a law firm eventually becomes a client notification letter.

 

We help firms map exactly this: what an agentic AI deployment can reach in practice, where the isolation claims hold up under testing and where they don’t, and what a realistic incident response looks like if a vendor’s containment fails the way Anthropic’s did. If your firm has rolled out Harvey, CoCounsel, or Copilot agents in the last year and nobody has independently verified the access boundaries, that’s the gap to close first.

Our Zero Trust Assessment is built for exactly this kind of review, mapping what every identity and every agent can reach across your environment rather than what a policy document says it can reach. If you’ve read our earlier piece on agentic AI security for law firms or your firm’s biggest AI risk, this is the natural next step: turn those concerns into a tested answer instead of an open question. We’ve also written specifically about Copilot governance for law firms if that’s the agentic platform your firm is furthest along with.

 

Anthropic’s disclosure is a rare gift, honestly. Most vendors never tell you when their containment fails. Use the transparency while it’s in the news cycle and go find out what your own agentic AI setup would do under the same conditions.

Recent Posts

Have Any Question?

Call or email Cocha.  We can help with your cybersecurity needs!

About the Author:

Picture of Steve Combs

Steve Combs

Co-Founder & Managing Director, Cocha Technology

Steven is a fractional CIO/CISO with 30+ years of enterprise IT and security leadership. He has built AI governance frameworks for organizations with 1,700+ users, led enterprise Microsoft Copilot deployments, and conducted security assessments across law firms, energy companies, financial institutions, and PE-backed manufacturers.