Reading time: about 9 minutes

Shadow AI in the factory is not a question of whether, it is a question of how much. In most plants nobody can put a number on it: how many BOMs, NDA drawings and commercial offers have already been pasted into public AI models. A firewall ban does not know that number, it only hides it. The good news is that the scale can be measured, and it is better to measure it yourself before a NIS2 auditor does it for you. This piece shows how to estimate how much of your documentation is already with external AI vendors: which methods to use, where each one goes blind, and what the number means for obligations that, in Poland, start running in 2026. The fastest starting point is a 10-minute readiness mini-audit that leaves no data behind.

Why a ban does not answer the real question

The standard reaction to shadow AI is a line in the security policy and a domain block on the firewall. The trouble is that a block changes behaviour, not the underlying need. An engineer on a deadline who used to paste a spec into a public chat from a work laptop will, after the block, do the same thing from a phone on a private network. The task still gets done, it just disappears from any log. A ban does not remove the risk, it removes visibility, and that is a worse starting point for an audit than the state before the ban.

Two well documented cases show the mechanism. In April 2023 Samsung found three leak incidents in a single month: engineers had pasted source code and chip yield test sequences into public ChatGPT. In August 2025 CISA's automated security sensors flagged four documents marked "For Official Use Only" being uploaded to public ChatGPT, on the account of the agency's acting director, under a use that had been formally approved. The uncomfortable takeaway: even sanctioned use of a public model moved the documents outside the perimeter, because the approval covered the tool, not the classification of what went into it. You do not need a hacker to exfiltrate technical documentation. You need an employee with a deadline.

Since a ban cannot be enforced tightly, the only thing that gives you an edge is knowing the scale. A number, even a rough one, turns the conversation from "we have a ban in policy" into "we know how many documents, and of what class, are genuinely exposed, and here is what we are doing about it".

The scale: where to start your estimate

Before you compute your own figure, it helps to know the benchmark. The 2025 and 2026 data is consistent across measurement methods, and it all says the same thing: zero is unlikely.

  • 47% of enterprise AI conversations run through personal accounts rather than corporate-managed ones, outside organisational control (LayerX State of AI Usage Report 2026).
  • Only 18% of employees use AI on a weekly basis, yet nearly half use it over a longer window, so a measure based only on "active this week" understates the picture (LayerX, 2026).
  • More than 6% of AI conversations contain sensitive data, rising to over 8% in ChatGPT itself (LayerX, 2026).
  • 63% of organisations have no AI governance policy, and 97% of firms that suffered an AI-related incident lacked proper access controls on those tools (IBM Cost of a Data Breach Report 2025).
  • A shadow AI incident adds an average of USD 670,000 to breach cost (IBM, 2025), and per UpGuard's 2025 data 80% of employees use AI tools despite bans.

These are not intent surveys, they are largely telemetry from systems monitoring real traffic. The takeaway for a factory is simple: if your current answer is "this does not happen here", it almost certainly means "we do not measure it here".

What actually leaks from a factory

In our 2026 conversations with manufacturers the same four document categories keep coming up, and these are the ones to count first, because each has a different leak path:

  • BOMs and technical specifications, pasted in for summarisation or comparison against a competitor's offer,
  • customer drawings under NDA, sent for OCR or translation,
  • commercial correspondence with prices, costs and margins, condensed into briefs,
  • service history with serial numbers and end-customer data, mined for failure patterns.

Every one of these happens in European factories today, and usually not out of ill will, but because it is faster than any alternative the company offers. That matters for measurement: you are counting rational behaviour, not sabotage, so people describe it more readily when the question does not read like an interrogation.

How to measure shadow AI in your own plant

There is no single counter that returns the answer. You estimate the scale from several sources at once, because each sees a different slice. Four methods, from cheapest to most labour-intensive:

1. Proxy, DNS and firewall logs. You usually already have this data. Filtering for the domains of public AI tools shows how much traffic reaches them, from how many corporate devices, when and how often. It is the cheapest first measurement, and usually the first to break the assumption that "nobody here uses it".

2. DLP, CASB and browser telemetry. The layer that sees not just the connection but the content: pastes, file uploads, the categories of data leaving the perimeter. A well configured DLP can tell a pasted offer fragment from a full CAD file upload. This is the closest you get to a real answer to "what leaves, and how much".

3. A structured survey and short interviews. Logs do not see the private phone, and that is where the largest share of shadow AI escapes. An anonymous survey ("which tasks do you speed up with AI", "which tools", "on what device") plus a few conversations with design, sales and service fills the blind spot in the logs. One condition: the question must not threaten consequences, or the result will be understated.

4. Sampling document categories. Instead of counting everything, take the four classes of sensitive documents (BOMs, NDA drawings, offers, service history) and, for each, check how it is actually processed and where it could have leaked. This will not give an exact volume, but it shows which document classes are most exposed, and that is where prioritisation starts.

No single method gives the full picture. The practical minimum is logs (method 1) combined with a survey (method 3): the first measures what is visible, the second reaches what is deliberately hidden. The table below sets out what each method catches, and what it misses.

Measurement method What it catches What it misses Effort
Proxy, DNS, firewall logs Traffic to public AI tools from corporate devices, scale and frequency Use from phones and private networks, the content of pastes Low, you usually have the data
DLP, CASB, browser telemetry Pastes and uploads, data categories, specific files Tools outside the managed device Medium, needs configuration
Survey and department interviews Motivations, tasks, workarounds, private hardware Exact volume, risk of under-reporting Low, needs trust
Document-category sampling Which document classes are most exposed and how they leak The full volume of the phenomenon Medium, manual work
Schematic of measuring shadow AI: data leaving the plant to external AI models, captured as metrics. You estimate the scale of shadow AI from several sources at once, because each sees a different slice of the traffic.

What the number means under NIS2

The figure you measure is not a curiosity for IT. Under the amended NIS2 directive it becomes part of an obligation. Article 21 requires essential and important entities to manage supply-chain risk, including the security of relationships with direct service providers. Every public model your documentation flows into is, in this framing, a supply-chain provider: unregistered, with no processing agreement, no SLA and no audit trail.

In Poland this is no longer distant theory. The amended National Cybersecurity System Act entered into force on 3 April 2026, and essential and important entities must file for entry in the register by 3 October 2026, with full compliance due by 3 April 2027. That makes August 2026 a month of preparation, not a buffer. In Germany the December 2025 BSIG amendment added personal liability of management towards their own company (§38) alongside administrative fines on the entity (§65, up to EUR 10 million or 2% of turnover). Personal liability changes who in the company reads this paragraph, and how quickly they ask for a meeting with compliance.

In an audit the question will not be "do you have a ban in policy", it will be "how much documentation, and of what class, do external models process, and how do you know". A company that knows its figure from logs and a survey answers from data. A company with only a policy line answers "we don't know", which is the worst possible answer.

From measurement to decision

Measurement is not the goal, it is the basis for a decision. Once you know the scale, three positions remain in practice:

  1. Ban and nothing more. Cheap, fast, ineffective. 80% of employees use AI anyway (UpGuard, 2025), just out of reach of the logs. A ban without an alternative does not lower the figure, it hides it.
  2. Enterprise tier of a public vendor. ChatGPT Enterprise, Copilot, Gemini Workspace. Better than nothing, but data still leaves the perimeter, and under Article 21 NIS2 it is still a supply-chain provider to map and audit, whatever the licence costs.
  3. Approved private AI on your own infrastructure. An alternative faster than pasting into a public chat, because the model sees your documents, not the internet, and the data never leaves the plant. It is the only one of the three that actually lowers the measured figure instead of hiding it. In one sector, introducing an approved alternative cut unauthorised use by 89% (Healthcare Brew, 2026).

A concrete order of steps for August: first, run a first measurement from the logs, however rough; second, supplement it with an anonymous survey in design, sales and service; third, map the result against the list of document classes that cannot leave; fourth, only then choose a tool, because now you know which problem it has to solve. If you want a starting point without pulling in the team, the readiness mini-audit takes 10 minutes and leaves no data behind. For the wider context of when private AI actually beats public AI, see private AI for manufacturing: when it wins.

If you would rather talk it through, 30 minutes with a founder is no pitch, just concrete questions and what the first week of a shadow AI audit usually reveals.

Frequently asked questions

What is shadow AI in a factory?

It is the use of public AI tools (most often ChatGPT and similar) for company tasks without IT's knowledge or approval. In manufacturing it typically covers summarising specs, translating drawings, comparing offers and analysing service history, that is, operations on documents that should not leave the plant.

How do I find out how much shadow AI happens here?

Start with proxy, DNS and firewall logs filtered for the domains of public AI tools, then add a DLP layer or browser telemetry and an anonymous survey in the key departments. Logs alone understate the result, because they cannot see the private phone.

Does banning ChatGPT solve the problem?

No. A ban changes the channel, not the need: the work moves to private devices and disappears from the logs. 80% of employees use AI despite bans. What works is offering an approved, faster alternative alongside measurement.

What does shadow AI mean for a NIS2 audit?

Every public model your documentation reaches is a supply-chain provider under Article 21. An auditor will ask how much documentation, and of what class, external models process. In Poland the application for entry in the register is due by 3 October 2026, so this is a now topic.

How long does a first measurement take?

A rough measurement from existing logs is a matter of hours, not weeks. The readiness mini-audit gives a starting point in 10 minutes, and a fuller picture from DLP and a survey can be gathered within one or two weeks.