Data security and privacy with AI — what you actually give away when you feed company data to someone else’s model
Most conversations about AI revolve around what the model can do. Far less often comes the question that, in a company, should come first: where does the data you feed it go, and who besides you can access it. This text is about exactly that — not to scare anyone, but to show, on hard numbers, where the real risk sits and what can be done about it.
One caveat up front. The point is not that AI is dangerous in itself. The point is that when an employee pastes a contract, a customer database or a fragment of code into a cloud model, that data leaves the company’s control — it lands on someone else’s server, under someone else’s policy, sometimes outside the jurisdiction where your GDPR applies. And this happens more often than anyone in the company assumes.
Shadow AI — the most expensive silent leak
The most telling data came from IBM’s “Cost of a Data Breach 2025” report — based not on forecasts, but on real breaches across hundreds of organizations. One thing from it is worth saying plainly:
“Shadow AI” — employees using AI tools without the company’s knowledge or approval — turned out to be one of the three most expensive factors in data breaches. An incident involving it costs on average 670,000 USD more than the already sizable average (4.44 million USD per breach). 20% of breached organizations were compromised precisely through shadow AI.
This is not a “tech company” problem. It is a problem for any company where someone — in good faith, to finish the job faster — drops data into a free chatbot, because “it’s just a bit of help.”
What actually leaks
The numbers get even more concrete when you look at what flows out:
- In 65% of breaches involving shadow AI, customers’ personal data leaked — more than the average for all breaches (53%). In 40% of cases, the company’s intellectual property leaked.
- ChatGPT alone recorded 410 million attempts to violate data-loss-prevention (DLP) policy in 2025 — attempts to send social security numbers, medical data, source code. These are not isolated slips; this is industrial scale.
- The volume of company data sent to AI tools reached 18,000 TB, growing 93% year over year.
In other words: the data that is the core of a company’s value — customers, contracts, know-how — is flowing outward at a pace no one measures, because no one is looking.
Why nobody notices
Because a shadow-AI leak is silent. No alarm, no break-in, no ransom on the screen. Just an employee who pasted something into a chat window. That is why such breaches are detected on average only after 247 days — over half a year during which the data has long been outside the company. A classic breach is visible; this one you have to find yourself, and most companies have no way to.
Where the hole comes from — missing control, not missing technology
The most important number in the whole report is not about attacks, but about our own carelessness:
- 97% of organizations that had an AI-related incident lacked proper access controls for those systems.
- 63% of companies have no AI governance policy at all — or are only just building one.
That means the problem is not that AI is “leaky.” It is that we let it into the company without a lock on the door: without a rule on who may use it, for what, and what data it may be given. Separately, it is worth noting that 13% of organizations have already reported direct breaches of the AI models or applications themselves — so it is not only the data-entry path that gets attacked, but the system too.
What to do about it
The good news is that this risk is manageable — and it does not require giving up AI, only setting it up consciously. A few things that genuinely reduce exposure:
- Policy before tool. Decide in writing what data may be given to a model at all, and what never (customer data, sensitive data, code). The policy alone — which 63% of companies lack — closes the biggest hole.
- Access control. Who, to what, with what permission. This is the missing layer in 97% of incidents.
- Team awareness. Shadow AI comes not from bad will but from haste. People who understand what they are risking stop pasting customer data into a free chat.
- Keeping data at home where possible. And here is a difference worth knowing: some AI use cases can be run locally/privately — so that the data never leaves the company’s infrastructure at all. Then “shadow AI” and “leaks to someone else’s cloud” simply have no way to happen, because there is nowhere for it to leak. This doesn’t solve everything, but it removes an entire category of risk from the statistics above.
Summary
AI in a company is not dangerous because it is AI. It becomes dangerous when we let it in without rules, without access control, and without awareness of where the data travels. IBM’s hard data shows the price of that carelessness: shadow AI is hundreds of thousands of dollars per incident, a leak of customer data in two-thirds of cases, and over half a year before anyone notices — and at the root there is almost always missing control, not missing technology.
The question worth asking in your own company is not “do we use AI” — because almost certainly someone already does. It is: “do we know what data we feed it and where that data goes?” If the answer is uncertain, that is exactly the hole worth closing — ideally before someone else finds it.
MafiaAI — a team of people and AI agents building tools, websites and solutions, including AI deployments run locally and privately, so your data stays with you. More: t8.pl
Data source: IBM “Cost of a Data Breach Report 2025” — IBM Newsroom, 2025-07-30.