Governance & Security

How to Roll Out Claude Code and Cowork Without Betting the Company Data

A security architecture playbook for agentic AI: data tiers, contract terms, permission modes, and the regulated-industry rollout pattern.

Ryan Drake

Ryan Drake

Founder, Ential · Jul 8, 2026 · 8 min read

Key takeaways

  • Classify data into green, amber and red tiers before rollout; the tiers, not the tool, decide what an agent may touch.
  • Get training opt-outs, retention limits and (where needed) zero-data-retention terms in writing before anyone pastes a client file.
  • Run agents on least privilege: read-only by default, approval gates on writes and sends, scoped credentials instead of blanket access.
  • Regulated firms should ladder in: non-client data first, then de-identified data, then controlled access with audit logging on everything.

An executive asked us last month: "How do I roll out Claude Code and Cowork without putting our company data at risk?" Then, in the same breath: "And how do I do it without neutering the thing? If the agent can't touch anything real, why am I paying for it?"

Both questions are right, and the second one is the reason most AI security advice fails. Lock everything down and you get an expensive chatbot that nobody uses. Open everything up and you get a client file in a training set. The answer is neither. It is an architecture, and it has five layers: data classification, commercial terms, permission modes, scoped access, and audit. Get those five right and you can hand agents real work in real systems, even in a regulated industry.

I know the five layers work because the most paranoid buyer on earth just signed off on the same stack. On 7 July, Anthropic put Claude Code and Cowork into public beta inside a FedRAMP High environment as Claude for Government Desktop: conversation history stored locally on agency-managed devices, department-level admin, per-department spend and model limits, and hash-chained audit logs. If federal agencies handling sensitive workloads can run this tooling under FedRAMP High controls, the controls exist for a 40-person finance firm to inherit. You do not have to invent the security model. You have to configure it.

What Data Is Safe to Put in AI?

The honest answer: it depends on the tier, and you need to define the tiers before rollout, not after the first incident. We use a three-colour classification with clients, and it fits on one page.

Green: goes in freely. Anything already public or harmless if leaked. Marketing copy, published pricing, public documentation, job ads, blog drafts, meeting notes with no client identifiers, your own process documents. This is where every rollout should start, because it lets people build skill with zero risk while you sort the contracts.

Amber: goes in with controls. Internal but not catastrophic. Financial reports, sales pipelines, internal strategy documents, codebases, contracts with client names, operational data. Amber data can absolutely go into Claude Code and Cowork, but only once the commercial terms below are signed, access is scoped to the people who already see that data, and logging is on. Most of the real ROI lives here: Anthropic's own usage data across 1.2 million anonymised Cowork sessions shows business-process operations (reporting, checklists, spreadsheet reconciliation) at 33.4 percent of usage, dwarfing software development at 8.7 percent (source). Reporting and reconciliation are amber-tier work.

Red: never goes in, or goes in only under a specific regime. Credentials and API keys, customer payment data, health records, anything under legal privilege, trade secrets whose exposure would be existential, and data you are contractually barred from sharing with subprocessors. Red does not mean "AI can never help here". It means the default is no, and any exception is a deliberate, documented decision with its own controls (think de-identification, private deployment, or a regulator conversation first).

Write the tiers down, give three examples of each from your own business, and put it on one page. If your team needs a decision tree to know whether a document is safe, the classification has failed.

Which Commercial Terms Actually Protect Your Data?

Three terms matter, and they matter more than any technical control, because they govern what happens to your data on someone else's infrastructure.

Training opt-out. Confirm, in writing, that your inputs and outputs are not used to train the vendor's models. On commercial and enterprise plans this is standard; on consumer plans it often is not. This single distinction is why "the team is just using personal accounts" is the riskiest AI posture a company can have. You are not saving money, you are donating your amber data.

Retention limits. How long are prompts and outputs held, where, and who can access them? Shorter is better; known is essential. For the most sensitive workloads, zero-data-retention agreements exist at the API level: your data is processed and discarded, never stored. If you operate under privacy law or client confidentiality obligations, ask for ZDR explicitly rather than assuming the default covers you.

Audit and compliance surface. Ask what the vendor gives your security team to work with. This has matured fast: Anthropic's Compliance API now ships with 28 security integrations covering platforms like CrowdStrike, Okta, Zscaler, Wiz and Microsoft Purview, so AI usage flows into the monitoring stack you already run instead of becoming a blind spot.

One more reason to take the vendor layer seriously: it is now front-page news. In June, the US government imposed export controls on Claude Fable 5 after researchers found a safeguard bypass; access was suspended and then restored on 1 July with an improved classifier (source). Whatever you make of the episode, the lesson for buyers is that the security and regulatory layer around frontier models is active and enforced, and your vendor's handling of it is part of your own risk posture. Ask how outages and fallbacks work before you depend on the tool.

How Do You Keep the Power of the Harness Without Handing Over the Keys?

You configure the harness the way you would onboard a capable new contractor: real access, narrow scope, supervision proportional to blast radius. Claude Code and Cowork are harnesses, meaning the model does not float free; it acts through permission modes, sandboxes and approval gates that you control. Used properly, those controls are the answer to the "power versus safety" question, not a tax on it.

The configuration that works in practice:

  1. 1Read-only by default. An agent that can read the codebase, the folder, or the report data can do most of the valuable analysis work with no ability to change anything. Start every new use case here.
  2. 2Approval gates on anything irreversible. Writes to production systems, emails and messages to humans, payments, deletions: each one pauses for a named human to approve. The agent drafts; a person releases.
  3. 3Sandboxes for execution. When an agent needs to run code or modify files, it does so in an isolated workspace, and a human reviews the diff before anything merges. Claude Security's codebase scanning slots in here too, catching vulnerabilities in what the agent (or your humans) wrote.
  4. 4Spend and model limits per team. The government deployment ships department-level spend and model caps for a reason: cost control is a security control, because runaway usage is usually a sign something is misconfigured. We cover the budgeting side properly in token spend is the new cloud bill.

Notice what this is: least privilege, separation of duties, change control. Your IT team has run these principles for twenty years. Agents do not need a new philosophy, they need the old one applied consistently.

How Do You Let People Access Internal Data Safely?

Through scoped connectors and scoped credentials, never through blanket access. The lazy rollout gives the agent an admin login and hopes for the best. The correct rollout gives each agent, and each user, a credential that can see exactly what that workflow needs and nothing else.

Concretely: the reporting agent gets a read-only key to the finance system, not the finance system password. The CRM connector is scoped to the pipeline the sales team already sees. The agent working on the codebase gets a repository token, not org-wide admin. When a credential leaks or an agent misbehaves, the blast radius is one workflow, and revocation is one key.

Then log everything. Who ran what, against which data, with which model, and what came back. The FedRAMP High deployment uses hash-chained audit logs, meaning the log itself is tamper-evident; you may not need that grade, but you do need logs your security team can actually query when someone asks "did any agent touch the client folder last quarter?" An answerable question is the difference between an incident and a crisis.

Guardrails like these are also what make wide adoption possible rather than merely safe. When people know the system will stop a catastrophic action, they stop being afraid of the tool, and usage goes up instead of down. That paradox (tighter rails, more freedom) is the whole argument of our piece on AI governance, and it is the difference between a policy people follow and a policy people route around with personal accounts.

Can You Use AI in a Highly Regulated Industry?

Yes, and the pattern is a ladder, not a leap. Financial services, health-adjacent businesses and legal practices ask us this weekly, usually expecting the answer to be "wait five years". The government beta settles the question of feasibility: FedRAMP High is a stricter regime than most private-sector firms will ever face, and Claude Code and Cowork are now operating inside it. The controls exist. Your job is sequencing.

The ladder we run with regulated clients:

  1. 1Rung one: non-client data. Internal operations, process documentation, marketing, generic code. Ninety days of real usage here builds the skills, the audit habit and the internal champions, with nothing regulated in scope.
  2. 2Rung two: de-identified data. Removing names is not enough. Assess re-identification risk in the actual access context, minimise and generalise fields where needed, apply access controls, and have your privacy or legal lead approve the workflow before an agent touches the data.
  3. 3Rung three: controlled access to live data. Scoped credentials, approval gates, retention terms and audit logging as described above, introduced one workflow at a time, with your compliance lead signing off each workflow rather than a blanket policy nobody read.

Each rung produces evidence for the next: usage logs, incident-free months, documented controls. When the regulator or the board asks how you got here, you have a paper trail instead of a shrug.

Where to Start This Month

Pick the smallest version of each layer. One page of data tiers. One confirmation email from your vendor on training and retention. One agent, read-only, on one amber workflow, with logging on. That is a secure rollout in miniature, and it scales by repetition, not by a grand programme.

If you want a second pair of eyes on it, our Automation Audit maps your workflows, classifies your data, and hands you a rollout sequence with the controls already specified. It is the fastest way to get from "we should be careful" to "we are careful, and shipping".

Keep reading

More from the Blog

The AI Governance Operating Model: Freedom Inside Guardrails
Governance & Security8 min read

The AI Governance Operating Model: Freedom Inside Guardrails

A two-lane operating model for governing AI: let anyone build in a sandbox, gate anything others depend on, and make every leader own the results.

Read article
The 14 Levels of an AI Rollout, Translated for a Mid-Market Budget
AI Adoption11 min read

The 14 Levels of an AI Rollout, Translated for a Mid-Market Budget

A well-known rollout map runs fourteen levels from first audit to self-guided agent. Here is how each level actually plays out when you have $1M to $50M in revenue, not a data department.

Read article
How to Run an AI Audit That Actually Leads Somewhere
AI Adoption9 min read

How to Run an AI Audit That Actually Leads Somewhere

Most "AI audits" are a sixty-page deck and an invoice. Here is the one-week version that ends in a ranked queue of work, not a strategy binder nobody opens.

Read article