Pause before you paste.
Nobody opens the AI policy at the moment they paste a spreadsheet into a chatbot, so I wrote the policy into the chatbot itself: a guardrail that pauses, names the risk and offers a safer route.
DisclosureWritten by Claude Opus 5.5. Curated, fact-checked and edited by Kevin Clubb.
Most AI policies live in a PDF. Staff meet it at induction, tick a box and never see it again. The moment it matters comes later, with a spreadsheet on the clipboard and a chat window open, and nobody goes looking for the PDF then.
So I put the policy in the chat window. This note covers how it was built, what it catches, and where it stops being any use.
The setting is a UK organisation running a small Claude Team plan. Its staff handle the kind of data that makes headlines when it leaks: personal details of the people it serves, staff records, payment information, and contracts that say where data is allowed to live. Before the plan existed, plenty of them were already using AI on personal accounts. Case Study 001 was the survey that found them.
Three readers in particular. Hannah, the HR business partner who is the only person in her building really using AI and has become the unofficial helpdesk as a result. Sanjay, who owns a 14-person business and wants to know whether any of this applies below enterprise scale. And Marcus, who runs information security and will spot the weakness before this paragraph ends. He is right, and it gets its own section.
The PDF nobody opens.
Microsoft’s UK research in October 2025 found that 71% of UK employees had used unapproved consumer AI tools at work, and 51% still did every week. The more useful numbers sit further down the same release. Only 32% said they were concerned about the privacy of company or customer data they typed in. And 28% said their employer did not offer an approved tool at all.
The use of Shadow AI poses security risks because sensitive company or customer data may not be protected effectively, leaving organisations vulnerable to data leaks, regulatory non-compliance, and increased risk of cyber-attack.
Microsoft UK · Rise in Shadow AI · 13 Oct 2025 · Censuswide, 2,003 UK employeesA company that sells the approved tool paid for that research, so weigh it accordingly. The shape still holds. Buying licences deals with the 28%. It does nothing about the fact that only about a third were worried about what they pasted. Hand those people a sanctioned tool and they paste the same things, only into something with better terms.
Training helps, once. A policy helps, if someone reads it. Neither is in the room when the pasting happens, and that is where the risk sits.
What a Team plan gives you.
My first question to Claude was blunt: could I put guardrails in at team level, so that a prompt containing personal data is blocked or at least warned? The answer was no, not natively. The Team plan has no built-in prompt filter. True blocking needs data loss prevention (DLP) tooling on the network, the browser or the device, and the admin feed of what people typed, the Compliance API, is an Enterprise feature.
What Team does give an owner is two levers.
- Organisation instructions. Up to 3,000 characters that sit on top of every conversation for every member. They win where they conflict with personal preferences, members cannot see or edit them, and a change can take up to an hour to reach everyone.
- Organisation skills. A skill is a written rulebook Claude loads when a task matches its description (Case Study 004 covers how they work). An owner can provision one for everyone. It arrives switched on, and each member can switch it off.
Neither lever enforces anything. Organisation instructions do take precedence over a member’s own preferences, and only owners can view or edit them. But precedence is not enforcement, and Anthropic’s own help page says so plainly.
In rare edge cases involving directly contradictory instructions, behavior may vary. Test your instructions to confirm they produce the results you expect.
Claude Help Center · Set organization instructions · Anthropic’s own caveatYou are asking the model to behave, and it usually does. That is the whole mechanism, and it shaped everything that followed.
Two layers and a floor.
The first version was about 450 words of organisation instructions trying to spell out every rule. It sat right against the limit, and it carried a hidden cost: Anthropic notes that whatever goes in that box travels with every message anyone in the organisation sends. The second version splits the job three ways.
The floor exists because of a detail that is easy to miss. A skill loads only when Claude judges that the request matches its description, and an instruction cannot force it to load. A member can also switch it off. The organisation instructions are the one place a member can neither see nor change, so the non-negotiables live there too.
Only the data skill runs on every relevant turn. Loading every skill on every message would burn context for nothing, so the cascade simply names the skills and stays short, and the others load when the work calls for them.
The five types are spelled out by name:
- Data that identifies the people the organisation serves, with particular care for vulnerable people and for health information tied to a named person.
- Staff personal data: home addresses, payroll, bank details, National Insurance numbers, HR and occupational health records.
- Payment and cardholder data.
- Credentials: passwords, API keys, tokens and connection strings.
- Client commercials and confidential tender material.
Naming them was the most useful decision in the build. “Don’t share personal data” is a principle, and principles get nodded past, by people and by models alike. “National Insurance number” is a column you can see in a spreadsheet. Name the column, and the model and the member of staff both recognise it.
What happens when it fires.
When a prompt trips the guardrail, Claude stops before doing the work. It says what it spotted in a sentence or two and offers a safer route. It points to the policy, then the line manager, then whoever owns data protection. Then it waits.
The safer routes are deliberately dull: counts instead of rows, “Location A” instead of a real address, a job title instead of a person, synthetic or redacted rows instead of the real extract. Where the work can stay inside systems the organisation already controls, it says so. Most cloud data warehouses now ship AI functions that run inside the platform, so analysis that needs the real rows can often happen where the data already sits.
I describe it as a warning, and that is how it feels to use. The wording underneath is stricter. “It’s fine”, “I own this data” and “it’s only a test” do not get through. “It’s already anonymised” does. Think of it as a pause with exactly one way out.
The tone rule took longer to get right than the categories: flag once, offer the alternative, name the check, move on. No lectures, and no repeating a warning the person has already dealt with. Nag people and they go back to the free account on their phone, where there is no guardrail at all.
Design for the irritated user. The careless one is easier to catch. Case Study 007 · the tone rule
There have been three live tests. The first was mine and deliberate: two rows of invented personal records. It stopped and offered to build the same table anonymised. The second arrived in the course of real work: a draft email about correcting some records, with named individuals in it. It stopped, and redirected the request to the colleague who could correct the records at source. The third was the same kind of request with the names left out, and it went straight through. That pass matters as much as the catches. A guardrail that fires on everything gets switched off.
What it cannot do.
Marcus will have got here first. By the time the warning appears, the prompt has already reached the provider. The guardrail stops the processing; it cannot recall the transmission. Anything pasted has left the building, warning or no warning.
The rest of the list is shorter but no kinder:
- It depends on the model following instructions. Organisation instructions take precedence on paper, but Anthropic’s own page warns that behaviour may vary where instructions directly contradict each other, and tells owners to test theirs.
- It trusts the answer. Anyone can type “yes, it’s anonymised” when it is not.
- A member can switch the skill off. The floor in the organisation instructions still applies, but it is the shorter, blunter version.
- Team has no audit log or live feed. The Primary Owner can export conversations after the fact, but nothing flags that a pause happened.
- What was pasted stays in the chat history, stored in the US, until someone deletes it.
- It covers Claude and nothing else. Every other AI assistant in the business needs its own controls; in Microsoft 365, that means Purview data loss prevention and sensitivity labels.
- Upkeep is manual. Every change means uploading a new version, and renaming a skill breaks the cascade without telling anyone.
Useful as a nudge, worthless as a control.
So why bother? Because a small team is governable, provided the rules turn up at the moment of risk. My working view: pair the guardrail with a policy people have been walked through in person, and make an annual refresher the price of keeping a licence.
The UK bit.
There is a second reason for the guardrail, and it is geographical. Anthropic’s privacy centre is explicit that data is stored in the US, and that routing is wider still. A Team plan has no UK option and no routing control. So the guardrail carries one rule outside the five types: treat every prompt as having left UK-controlled systems.
By default, we may route customer traffic to select countries in the US, Europe, Asia and Australia, unless otherwise agreed upon or at your instructions.
Anthropic Privacy Center · Where are your servers located? · no UK option on TeamUnder UK GDPR, sending personal data to a US provider is usually a restricted transfer, and it needs a lawful route: the UK Extension to the EU-US Data Privacy Framework (the “data bridge”, in force since 12 October 2023) where the US company is certified and has opted in, or contractual safeguards such as standard contractual clauses with the ICO’s UK Addendum, backed by a transfer risk assessment. Check which route your provider actually relies on. The Data (Use and Access) Act 2025 changed the bar for those routes to protection “not materially lower” than the UK’s, with its main data protection changes in force from 5 February 2026. The ICO refreshed its international transfers guidance on 15 January 2026.
Lawful is only half the question. Plenty of client contracts and data processing agreements promise UK-only processing, and a transfer the law allows can still break a promise you signed. That is why residency is a separate rule in the guardrail: data covered by an onshore clause, or not cleared for third-party cloud at all, does not go in.
Hannah, this is where HR lives. Staff records, occupational health and grievance notes all sit in the five types, and most of them are covered by what your organisation has promised its own people. Sanjay, none of this needs enterprise money. Organisation instructions and skills come with the Team plan, and a smaller team makes the approach easier to run, because you can explain the rule to everyone face to face.
One practical note for any Team owner. Members can publish their own skills to the organisation library, and on Team the default is “Open”, with no review. Switch it to “Requires review”, or your guardrail may soon share the library with skills nobody has read.
Before you copy this.
Six questions. If you cannot answer one, do not switch it on yet.
In short.
Strip it back and the whole approach fits in three steps.
It will not stop a determined leak, and it was never meant to; the prompt has left before the pause appears. What it does is put the rule in front of people at the one moment the policy PDF never reaches, in plain words, with a way forward.
It is early. Three live tests, three correct calls: an anecdote, and I treat it as one. The pattern has still earned its keep. The same two layers now carry the organisation’s external voice, brand and data architecture into every conversation. And when the governance board asks what stops someone pasting a staff list into a chatbot, there is a straight answer, written in the risk register as exactly what it is: a nudge, placed where the pasting happens.
Put the policy where the pasting happens, and call it a nudge.
One issue a week, and the workings underneath it.
No paywall and no pitch. This note is the third in a line: the survey that found the shadow AI, the skill that taught the rule once, and now the guardrail that puts it where the pasting happens.