A board asks two questions — what AI are we using, and where is our data going — and the room goes quiet. Not because nobody did the work. Someone sent a questionnaire, collected the answers, and built a register. The problem is that a questionnaire is a survey of memory, and an AI inventory is a stocktake. Those are different instruments, and only one of them counts what is actually on the shelf.

Four reasons the count comes back short

People report what they remember. Teams remember the systems they chose on purpose — the assistant that went through procurement, the pilot with its own budget line. The script a colleague wrote nine months ago to triage incoming tickets does not come to mind, because it never felt like an AI system. It felt like a script.

AI arrives inside things you already approved. This is the largest gap by volume. A vendor ships a summarisation feature in a minor release of a tool that cleared procurement two years ago. Nobody declared it, because nobody was asked to — the tool did not change its name when it grew a model. Your approved-software list is accurate and your AI inventory is wrong at the same time.

The categories don’t match reality. A form that asks “do you use AI?” collects answers to the question people assume you mean. Ask about machine learning models and you will not hear about the chatbot. Ask about chatbots and you will not hear about the classifier sorting claims.

A snapshot ages badly. A questionnaire is true on the day it closes. Then trials start, extensions get installed, and a team wires up an agent over a long weekend. Six weeks later you are governing last quarter’s count.

A questionnaire counts what people remember. A stocktake counts what is on the shelf.

Name it by what it does, not by what it’s called

The most useful move is to stop hunting for the word “AI” and start looking for the behaviour. A system belongs in the inventory if it takes your data and produces a generated, inferred or scored output that somebody then acts on. That definition catches the vendor feature nobody declared, the internal script and the spreadsheet plugin, and it does not depend on anyone using the same vocabulary you do.

It also gives you a test the whole organisation can apply without a glossary: what data goes in, what comes out, and who acts on it.

What to count instead

Discovery is several partial counts, aimed at systems rather than at people. Which of them you can run depends on how the AI arrived — some organisations build, most only buy — but everyone can run more than one.

Whoever owns each of these — your developers, your IT function, or your vendor — can produce their part in an afternoon. The work is asking each of them for a list rather than for an assurance.

No single aisle gives you the count

Every channel above is incomplete on its own, and they overlap in confusing ways. Egress sees traffic but not purpose. OAuth sees a grant but not whether anyone uses it. Code sees intent but not what shipped. Spend sees a licence but not the tenant configuration.

So reconcile them. Put the channels side by side and treat every disagreement as a finding rather than as noise: a grant with no traffic, traffic with no owner, a licence with no grant. The questionnaire is still worth sending — as one channel among several, and best sent without blame, because a survey that reads like an audit returns the answers people believe are safe.

Where to start

Pull the OAuth grants and one month of egress, and lay them beside whatever register you already have. That comparison takes a morning and it tells you the size of your gap before you commit to a method. Then record the two facts per system that make an inventory usable later: what data flows in, and who acts on the output.

And put a date on it. Discovery is not a project you finish, it is a count you repeat, because the estate keeps taking deliveries. An inventory is accurate the day you close it and decaying the day after.