How do you run a sovereign AI pilot in 90 days?
Pick one measurable workflow, run it on operator-owned hardware inside your perimeter, and let the sealed evidence decide the scale decision.
A sovereign AI pilot runs in three phases of thirty days. Days 1 to 30: choose one workflow with measurable volume, fix the success metric, and stand up operator-owned hardware inside the existing perimeter with zero egress. Days 31 to 60: run the workflow in parallel with the human process, logging every AI action to a sealed ledger. Days 61 to 90: review the evidence with compliance, decide scale or stop against the metric, and file the audit trail the pilot produced.
The shape matters in 2026 because most stalled AI programmes stalled at exactly this point: a promising demonstration that could not answer a compliance question. A pilot designed to produce evidence rather than enthusiasm is the difference between a third year of experiments and a deployment.
Which workflow should you pilot first?
One workflow, chosen for volume and measurability rather than ambition. Drafting, extraction and triage are the reliable candidates: correspondence drafting, data extraction from documents, and inbound case triage all have countable throughput and an existing human baseline to compare against. Avoid workflows whose output cannot be scored, and avoid piloting three things at once, because a pilot that measures nothing cleanly proves nothing at all. The test for a good candidate is simple: can you state today how many items the team processed last month and how long each one took? If not, pick a workflow where you can.
What has to happen in days 1 to 30?
The first month is scoping and installation, and the discipline is to finish both before any output is generated.
- Fix the success metric in writing: throughput, turnaround time or error rate against the human baseline.
- Stand up operator-owned hardware inside the existing network perimeter, configured for zero egress so nothing leaves the boundary.
- Load the sovereign models, version and hash the weights, and record the deployment in the audit ledger.
- Agree with compliance, before day 30, what evidence would justify scaling.
The last item is the one most pilots skip, and it is the one that makes day 90 a decision instead of a debate.
What has to happen in days 31 to 60?
Parallel running. The AI performs the workflow alongside the human process, not instead of it, and every action it takes is written to the sealed ledger: input received, output produced, model version used, and the human decision to accept or override. Staff correct the system openly, and the overrides are data. Measurement is continuous, against the metric fixed in month one, and nobody moves the metric mid-pilot. By day 60 the pilot holds a complete, signed record of what the system did on live work, which is precisely the record no slide deck can imitate.
What has to happen in days 61 to 90?
Review and decision. Compliance and the workflow owner sit with the evidence: throughput against baseline, override rates, and every case traceable in the ledger. The decision is scale or stop, taken against the metric agreed in month one, and both outcomes are successes, because a clean stop on evidence costs one quarter while a drifting pilot costs years. If the decision is scale, the pilot's own logs become the first artefact in the compliance case for production, because they demonstrate on live work that the system operates inside the boundary and that every action is accounted for.
Why does a public cloud trial prove nothing about a sovereign deployment?
Because the constraint being tested is the perimeter, not the model. A trial on a shared cloud service demonstrates that a model can perform a task, which was not in doubt. It demonstrates nothing about the deployment a regulated buyer would actually approve: data remaining on operator-owned hardware, zero egress, identity bound to the audit chain, and evidence a regulator can verify. A cloud trial can also create the governance problem the sovereign deployment exists to avoid, by moving live records outside the boundary during the experiment itself. Pilot inside the perimeter from day one, on the architecture you would run at scale.
What evidence should the pilot leave behind?
A sealed, verifiable record, not a summary written afterwards. Inside Mickai, our Sovereign Intelligence Operating System, every pilot action is signed into a post-quantum audit ledger under FIPS 204, with hardware-attested identity bound to each entry, and with model weights versioned and hashed so the record shows which model produced which output on whose authority. The evidence is verifiable offline: checking it requires no connection to any external service. That record is the pilot's real deliverable. Throughput numbers persuade a budget holder; a sealed ledger persuades an auditor, and scaling requires both.
“A pilot that cannot show a sealed record of what the system did was a demonstration, not a pilot.”
How the pilot architecture fits the wider system is set out at /sovereign-ai, and the film at /film shows the interface in operation.
Frequently asked questions
How long should an AI pilot run before we commit?
Ninety days is enough when the pilot is scoped to one measurable workflow. Thirty days to scope and install, thirty to run in parallel with the human process, and thirty to review the evidence and decide. Pilots that run longer without a decision date tend to drift into permanent experiments, which is the most expensive outcome of all.
What is a good first use case for a sovereign AI pilot?
Drafting, extraction or triage, because all three have countable volume and a human baseline. Correspondence drafting, document data extraction and inbound case triage are the patterns that succeed most often. The common factor is measurability: the workflow must produce a number that compliance and the budget holder both accept.
Can we pilot on a cloud service first and move on-premise later?
The two pilots test different things. A cloud trial tests the model, while a sovereign pilot tests the perimeter, the audit trail and the operating discipline, which are the constraints that decide whether regulated data can be used at all. Evidence from a cloud trial does not transfer to the sovereign compliance case, so the practical route is to pilot on the architecture you intend to run.
What should the success metric for an AI pilot be?
One number fixed before the pilot starts, drawn from the workflow itself: items processed per week, turnaround time per case, or error rate against the human baseline. Override rate is the essential companion measure, because it shows how often staff had to correct the system. Fix the metric in writing by day 30 and do not move it mid-pilot.
Who needs to be involved in a 90 day AI pilot?
Three roles: the workflow owner who runs the parallel process, a compliance reviewer who agrees the evidence standard up front and judges it at the end, and an operator for the hardware inside the perimeter. The decision at day 90 belongs jointly to the workflow owner and compliance, taken against the metric rather than against sentiment.