Copilot does not create a governance problem. It makes the one you already had visible to everyone at once.
The single most predictable moment in a Copilot rollout arrives about three weeks in. Someone in finance asks a routine question and gets back a salary band. Or a project manager asks about a supplier and gets a draft contract that nobody meant to share. The reaction in the room is usually that the AI has leaked something.
It has not. Copilot respects existing permissions. What it did was ask the question nobody had asked before: given everything this person is technically allowed to open, what is in there?
For fifteen years, oversharing was survivable because nobody could find anything. Search was bad enough to act as an accidental control.
That control has now gone. An assistant with natural language retrieval is very good at finding the one document in a site of forty thousand that answers a question, including the ones that were only ever safe because they were buried.
The order of work
Data readiness gets presented as a maturity model with five levels and a radar chart. In practice it is a ladder, and the rungs have to come in order. Skipping one does not accelerate anything; it just delays the failure.

Most organisations want to start at the top. Semantic models and business glossaries are satisfying work; they produce artefacts you can show people, and they feel strategic. Meanwhile the second rung, permissions, is unglamorous, politically awkward and involves telling site owners that their sharing habits are a liability.
Start at the bottom anyway. The bottom two rungs are where Copilot programmes actually fail, and they are the only ones with a hard external deadline, because the day you switch on an assistant is the day the exposure becomes real.
What good looks like on the Microsoft stack
The architecture Microsoft is pushing towards is straightforward and worth understanding, even if you are years away from implementing all of it. Land data once, govern it once, let everything consume the same copy.

The part people underestimate is that Purview sits underneath all of it rather than beside it. Sensitivity labels applied to a document in SharePoint are honoured when Copilot retrieves it. Labels flow into content that Copilot generates. Audit captures the prompts and responses. Data Loss Prevention can prevent specific files from being used in responses altogether. This is one governance model, not three.
For an assistant grounded in Microsoft 365 content, none of this requires Fabric. The Fabric conversation matters when the question spans systems, which is where most genuinely valuable analytical use cases live.
The remediation nobody budgets for
There is a body of work between deciding to deploy an assistant and safely deploying one. It is boring, it takes a quarter or two, and it rarely appears in the business case.
- Finding sites shared with Everyone, or with Everyone except external users, and closing them. This is usually a four-figure number of sites in a mid-sized organisation, and a five-figure number in a large one.
- Archiving or deleting stale content. Old versions of policies are not just ineffective; they can lead an assistant to cite them with total confidence.
- Identifying ownerless sites. Every site with no owner is a site with no one to ask about access, and there are always more than anyone expects.
- Applying sensitivity labels to the categories that matter, which in most organisations means HR, legal, commercial terms and anything customer identifiable.
- Deciding what to do about the shared drive that everyone knows about and nobody will claim.
None of this is AI work. All of it is prerequisite to AI work. My honest view is that this is the single largest hidden cost in enterprise AI adoption, and the organisations that treat it as a separate, properly funded programme do far better than the ones that treat it as a blocker to be worked around.
For the practitioners
- Run SharePoint data access governance reports first. They give you the oversharing picture without a discovery project, and they let you send site access reviews directly to owners.
- Restricted content discovery lets you exclude a site from Copilot and agent retrieval while you remediate it, rather than blocking the rollout entirely. It is the pragmatic middle path.
- Enable sensitivity labels for Office files in SharePoint and OneDrive before rollout. Without this, encrypted file protections are limited to content in use in Office apps on Windows.
- Where a label applies encryption, users need the EXTRACT usage right as well as VIEW for Copilot to return the content. This catches people out.
- Use Purview data risk assessments to monitor oversharing rather than as a one-off audit continuously. Remediation decays.
Reframing the conversation
If you are trying to get funding for this work, do not sell it as data governance. Nobody has ever been enthusiastic about funding data governance.
Sell it as the thing that determines whether the AI investment yields any returns. That is not a rhetorical trick; it is accurate. An assistant grounded in a messy estate produces confident, plausible, wrong answers, and the fastest way to kill adoption is to have that happen to a senior person in front of colleagues. People forgive a tool that says it does not know. They do not come back to one that misled them.
The organisations that did this work before rolling out are, without exception, the ones whose users still trust the tool a year later.
How long this actually takes
The honest answer is longer than anyone wants to hear and shorter than the fear suggests, provided you scope it as remediation of the sites that matter rather than as a general cleanup of everything.
| Phase | Work | Typical duration |
|---|---|---|
| Assess | Access governance and data risk reports. Produce the list of overshared, ownerless and stale sites. | Two to four weeks |
| Contain | Restrict discovery on the worst sites so the rollout can proceed safely while remediation continues. | One to two weeks |
| Remediate | Close broad sharing, assign owners, archive stale content, apply labels to the categories that matter. | One to two quarters |
| Sustain | Scheduled reassessment, site lifecycle management, label policy enforcement. | Ongoing |
The contain phase is the one people miss, and it is what allows the programme to keep moving. You do not have to fix everything before deploying anything. Please identify what is broken, prevent the assistant from reaching it, and then work through the list on a timeline someone owns.
Two arguments you will have
This is not an AI problem; it is an IT problem. Correct, and it does not help. The work is still on the critical path, and framing it as somebody else’s problem will keep it unfunded for another year. The useful framing is that the AI investment has surfaced a pre-existing liability with a now-visible cost, which is a more honest and considerably more effective description in front of a board.
We cannot ask site owners to review thousands of sites. Also true, which is why the tooling supports sending access reviews directly to owners rather than routing everything through a central team. The central team’s role is to produce the list, set the deadline, and escalate items that are not completed. Distributing the decisions is the only version of this that finishes.
Neither argument is unreasonable. Both are usually deployed to avoid starting, and the cost of not starting compounds because content volume does not stand still while you deliberate.
What to take away
- Copilot exposes your permissions model. It does not change it. Audit before you deploy, not after.
- Work the ladder from the bottom. Findable, then permissioned, then labelled, then current, then modelled.
- Budget the remediation as a named workstream with its own owner and timeline.
- Use restricted content discovery to keep the rollout moving while you fix the worst sites.
- Stale content poses a greater quality risk than missing content because an assistant will cite it confidently.
