Shadow AI discovery and AI vendor risk
Nobody signed a decision to adopt AI. It arrived in a browser extension, in a note-taker that joined a meeting, in a feature a supplier switched on inside a product you already pay for, and in a personal account someone uses to get a draft finished before five. The question is not whether to allow it. It is what is already there, and what reaches it.
Two halves of one question
An AI estate has an inside and an outside, and most organizations can see neither. Inside are the tools staff use directly. Outside are the suppliers who have quietly put a model behind a product you have used for years. Both put your content somewhere it was not before.
The exposure is not hypothetical. Canada’s Centre for Cyber Security states it plainly in ITSAP.00.041:
Users may unknowingly provide sensitive corporate data or personally identifiable information (PII) in their AI queries and prompts.
Unknowingly is the operative word. This is a visibility problem before it is a policy problem, which is why discovery comes first.
Shadow AI: what is actually in the estate
Discovery uses the telemetry you already have rather than new agents on endpoints: DNS and egress logs, the identity provider’s application and consent records, OAuth grants issued against your tenant, SaaS-management or CASB inventories, expense data for tools bought on a card, and browser-extension inventories where your management tooling collects them.
Those sources answer different questions and none answers all of them, so the inventory records how each entry was found and how confident that finding is. An inventory that cannot say where a line came from does not survive its first challenge from a business unit.
Will discovery mean monitoring employees?
Discovery is an inventory of tools and data categories, not surveillance of people. It runs against the telemetry your organization already collects, and the method is scoped against your own employee-privacy obligations before it runs. That scoping is done before collection begins, not justified afterwards, and it is written into the method: which sources are in, at what granularity, aggregated or per-user, retained for how long, and who may see the output.
This matters in Canada beyond good manners. Employee personal information is personal information, and an organization looking for shadow AI can create its own privacy problem while doing it. In British Columbia the Personal Information Protection Act governs private-sector handling of employee personal information; federally, PIPEDA applies to organizations in the course of commercial activity. The discovery plan is written to sit inside whichever applies to you, and where the answer is genuinely contested it is a question for your counsel rather than for us.
What the inventory shows
- Which tools, named, with how they entered — procured, self-signup, bundled into an existing product, or embedded in a browser.
- Which data categories reach each one: customer personal information, source code, contracts and pricing, health or financial records, internal strategy, or nothing sensitive at all.
- Under which account — a corporate identity you can govern, or a personal one you cannot.
- What it is used for, in the words of the people using it. This is the part that makes the policy work.
The acceptable-use policy people actually follow
A policy that only prohibits gets routed around, and the routing is invisible to you. The version that holds does three things: it names what may never be pasted into a general-purpose AI tool, it says which tool to use instead for each common task, and it gives a route to ask for something new that returns an answer in days.
If the finding is that most of the workforce already uses AI, that is the normal result and it is useful: it tells you a prohibition would be ignored, and it names the tasks a sanctioned alternative has to cover to be adopted. Standing up that alternative is secure AI environment establishment; the policy is what points at it.
AI vendor risk: the questions that matter
Generic supplier questionnaires miss what is different about an AI supplier. These six are the ones whose answers change a decision, and OWASP LLM03 Supply Chain is the category they sit under.
| Question | What the answer has to settle |
|---|---|
| Training on customer data | Is customer content used to train or improve models, by default or by opt-out? Does the answer differ between the consumer tier your staff signed up for and the business tier you are buying? |
| Retention | How long are prompts, outputs, attachments and logs kept? Is there a zero-retention or short-retention mode, and does using it disable features you depend on? |
| Sub-processors | Which model providers, hosting providers and evaluation vendors sit behind the service, and how is the list published and changed? |
| Model provenance | Which models serve the endpoint, whether they are the vendor’s own or licensed, and what notice is given when they change under you. |
| Data residency | Where content is processed and stored, whether a region can be pinned, and whether support and abuse review can reach it from elsewhere. |
| Incident notification | What the vendor commits to tell you, how fast, and whether that commitment is in the contract or only in a trust-centre page it can edit. |
Proportionate review
Reviewing every AI vendor to the same depth guarantees the important ones get the same fifteen minutes as the trivial ones. Depth is set by what the vendor holds and what it can do:
- Light. No regulated or confidential content reaches it. Record the tool, the owner, the data categories, and move on.
- Standard. Confidential business content reaches it. Full questionnaire, published terms read, sub-processor list captured, renewal date tracked.
- Deep. Personal, health, financial or regulated content reaches it, or it can act in your systems. Add architecture questions, residency evidence, incident-notification terms, and a named position for your counsel to negotiate.
Accountability does not transfer with the data. Under PIPEDA’s accountability principle, an organization remains responsible for personal information transferred to a third party for processing and must use contractual or other means to provide a comparable level of protection while it is there (Schedule 1, clause 4.1.3). Sending it to a model provider is a transfer like any other.
Contract positions — and where we stop
We read a vendor’s published terms as a security and privacy reviewer and set out what they say about training on customer data, retention, sub-processors, residency and incident notification. What those terms mean as a contract, and what to negotiate, is work for your counsel. We recommend positions for your counsel to draft: no training on customer content, retention bounded and stated, sub-processor change notice with a right to object, residency committed in the agreement rather than on a web page, incident notification with a clock on it, and deletion on termination evidenced rather than asserted.
We do not draft contracts and we do not give an opinion on what one means. That line is not modesty; it is the difference between a security reviewer and a lawyer, and blurring it would leave you with neither.
What you receive
- An AI tool inventory with owner, data categories, account type and discovery source for every entry.
- An acceptable-use policy that names the sanctioned alternative for each common task and gives a request route for new ones.
- A vendor questionnaire set, tiered light, standard and deep, that your procurement team can run without us.
- A supplier risk register recording each vendor’s tier, answers, gaps, renewal date and the positions handed to your counsel.
The register is the source the governance work draws on (AI governance and regulatory readiness), the evidence a customer questionnaire draws on (answering the AI security questionnaire), and the same discipline applied to suppliers generally in third-party and M&A due diligence. Where governance rather than engineering is the gap, it maps to the GOVERN function of NIST AI RMF 1.0.
How the work is bounded
The scope is agreed in writing before work starts, and the engagement is quoted in writing with it.
SecHB does not issue certifications, attestations or audit opinions: those come from accredited certification bodies, CPA firms and QSAs. The work here is what an organization does to be ready for them.
Nothing here is legal advice. Where a question turns on the law, the work is done alongside the client’s counsel, not instead of them.
Questions we are asked
Will discovery mean monitoring employees?
Discovery is an inventory of tools and data categories, not surveillance of people. It runs against the telemetry your organization already collects, and the method is scoped against your own employee-privacy obligations before it runs.
What if the answer is that everyone is already using it?
If the finding is that most of the workforce already uses AI, that is the normal result and it is useful: it tells you a prohibition would be ignored, and it names the tasks a sanctioned alternative has to cover to be adopted.
Can you review a specific vendor’s terms for us?
We read a vendor’s published terms as a security and privacy reviewer and set out what they say about training on customer data, retention, sub-processors, residency and incident notification. What those terms mean as a contract, and what to negotiate, is work for your counsel.
How often should the AI vendor review repeat?
Re-review a vendor when its terms change, when the data you send it changes category, at renewal, and at a fixed interval set by the vendor’s tier — annually for the ones holding regulated data, less often for the rest.
Find out what is already there
Describe your identity provider, your SaaS-management tooling and the suppliers you already worry about. The reply says what discovery would use and what it would produce.