Shadow AI Is Already Inside Your Organisation. Here's What To Do About It.
Shadow AI isn't a hypothetical risk sitting in a future risk register — it's the default state of most workplaces today. Employees aren't being reckless; they're trying to work faster, using tools that feel no different from search. But every prompt is a potential export of data you don't control once it leaves.
Within weeks of engineers gaining ChatGPT access in 2023, Samsung Electronics suffered three separate confidential data exposures — proprietary source code, equipment-defect detection code, and a full internal meeting transcript, all pasted in with no malicious intent whatsoever. Samsung's response was an emergency company-wide ban, the same blunt instrument most organisations reach for once a leak is already public, because a written policy alone couldn't stop it in time.
The scale of the problem
Recent industry survey data puts this in perspective: two-thirds of office professionals report using AI tools at work despite believing it wasn't permitted under company policy, and more than a third admit to entering customer data into public models. Meanwhile, breach investigations show that AI-related incidents overwhelmingly involve systems with no governance policy in place at all. The ICO has published its own guidance on AI and data protection for organisations working through exactly this question.
"A blanket ban teaches nothing and pushes people toward personal accounts you can't see at all."
— Praeferre AI Governance AdvisoryWhat actually works: five steps
- Detect before you decide. You can't govern what you can't see. Real-time visibility into which AI tools are being used, by whom, comes before any policy decision.
- Classify data sensitivity by type, not by team. A blanket policy treats your legal department and your marketing interns identically. It shouldn't.
- Choose redact, obfuscate or block — deliberately, per category. Not every sensitive value needs the same response.
- Make every intervention a teaching moment. When something is blocked, explain why, in plain language, at the point it happens.
- Feed the data back into your compliance evidence base. Every interception should strengthen your GDPR and DPDP audit trail, not just sit in a log nobody reads.
How the numbers compare across regions
| Region | Using unsanctioned AI | Have a written AI policy | Policy enforced technically |
|---|---|---|---|
| UK & Ireland | 63% | 41% | 18% |
| EU | 59% | 47% | 22% |
| North America | 68% | 38% | 16% |
| APAC | 71% | 33% | 12% |
- Shadow AI
- Any AI tool used for work purposes without formal organisational approval or oversight.
- Redaction
- Stripping a sensitive value from a prompt while letting the rest of the content through.
- Obfuscation
- Replacing a sensitive value with realistic synthetic data so the prompt still functions.
Names, emails and phone numbers are typically redacted — the rest of the prompt still works perfectly well without them, so there's no reason to block the whole request.
Proprietary source code is usually obfuscated rather than blocked outright — engineers can still get a useful answer from a synthetic but structurally similar snippet.
Live contract and legal text is typically blocked entirely. There's rarely a version of "let some of this through" that doesn't risk a confidentiality breach.
See what your organisation is exposing right now. A short discovery call shows you what a live scan of your AI usage would surface in week one.
Explore AI Data Leak ProtectionDetection combines pattern-based matching (structured identifiers like IBANs, national insurance numbers, and email formats) with a lightweight classification model trained specifically for compliance-relevant entities — not a general-purpose model repurposed for the task. Documents are additionally classified by type before any content-level scan runs, so a contract is treated differently from a public brochure before a single sentence is read.
All of this happens client-side in the browser extension or desktop agent wherever possible, meaning the sensitive content itself never needs to leave the user's machine just to be classified — only the redacted or approved version continues on to the destination AI tool.
Common questions
Yes — which is exactly why a ban rarely works as a standalone control. Detection needs to be paired with a genuinely usable, sanctioned path to AI, not just a wall.
Traditional DLP wasn't built to understand what a prompt to an AI model actually is, or to distinguish a harmless request from one carrying a live contract clause. Purpose-built AI DLP reads intent and document type, not just keyword matches.
Redaction and obfuscation both happen in milliseconds, inline with typing. Most users don't notice a delay at all — they notice the sensitive value quietly missing from what they see get sent.
Comments
Want to talk this through with our team?
Tell us what you're working on and we'll reply within one business day.
Get in touch