MICKAI®ArticlesMeta MCI leak of 45,000 tables sh…
Article · 5 August 2026

Meta MCI leak of 45,000 tables shows centralised AI training is a blast radius

A single misconfiguration exposed 45,000 tables of raw employee data to the whole company. The architecture, not the control, is what failed.

Author
Micky Irons
Published
5 August 2026
Follow Micky Irons
LinkedInX
ai governancemeta mcisev 2 incidentgdpr article 32on-premise ai training
Meta MCI leak of 45,000 tables shows centralised AI training is a blast radius

In late June 2026 Meta paused its Model Capability Initiative after a SEV 2 left 45,000 internal tables, including private messages, meeting transcriptions, performance records and employee prompts, open to essentially every employee. The lesson is structural. Centralised training pools every sensitive input into one blast radius. Regulated buyers should train on the tenant's hardware, with no data leaving the perimeter.

What Meta paused, and why

The Model Capability Initiative, or MCI, was Meta's internal programme to train models on the day-to-day work of Meta employees. Tracking software installed on US employee laptops from April 2026 captured keystrokes, mouse movements, click locations, screen content and internal chat, feeding the collected material into a training corpus. The programme did not offer an opt-out on company devices. In May 2026 more than 1,600 Meta employees signed an internal petition demanding it be cancelled outright. In late June a security review found that 45,000 internal database tables containing the collected data were accessible to essentially every Meta employee. Meta classified the incident SEV 2 on its internal severity scale, announced a pause of MCI on 23 June 2026, and continued to disclose further exposures through the first half of July.

Why a SEV 2 at that scale is worse than it sounds

SEV 2 in Meta's internal rating system is a high-priority incident, in the tier reserved for events that materially affect the whole company. What made this one structurally different from a typical SEV 2 is the content of the exposed tables. The material was not user data collected under a public terms of service. It was raw internal work, including performance discussions, unpublished plans, private conversations, and, according to employees who spoke publicly, personal tax and medical information. The audience with visibility was not a small circle of authorised engineers. It was the general employee population. A regulator reading the incident brief will note that both the sensitivity of the data and the size of the exposed population were unusually high, and that the failure mode was a single permissions misconfiguration.

The pattern the incident exposed

Centralised, vendor-owned AI training is not a Meta-specific idea. It is the default posture across the industry. A vendor collects data from customers, aggregates it into a training corpus, applies coarse access controls to the corpus, and trains a shared model on top. The pattern has an attractive economic story, because a single shared model is cheaper to run than one model per customer. It also has a single, well-understood structural weakness. The training corpus is a concentration of every sensitive input the vendor has ever seen. When the access control on that corpus fails, the failure is uniform across every contributor. That is the story of MCI. One misconfiguration exposed 45,000 tables that had been aggregated from every part of the company at once.

Why the regulator reads this as a structural failure

Under GDPR Article 32, controllers and processors must implement measures appropriate to the risk, including confidentiality, integrity and availability of processing systems. The UK GDPR carries the same test. Neither text tolerates the argument that a single misconfiguration is an acceptable failure mode when the underlying architecture concentrates every sensitive input into one accessible pool. The Information Commissioner's Office guidance on data protection impact assessments requires controllers to assess the nature, scope and risks of the processing, which in practice means describing the exposure created by the chosen architecture, not only the strength of the controls applied to it. A centralised training pool, by construction, sits badly against that test. The exposed population is the whole pool. A DPIA that recommends a centralised training pool for regulated data has to explain, in writing, why an on-premise per-tenant alternative was not adopted.

The on-premise alternative

The alternative is simple to describe. The training pipeline runs on hardware the customer owns. The data does not leave the customer's perimeter to reach the training runtime. The resulting model is the customer's model, held on the customer's silicon, not a vendor asset. If a misconfiguration exposes a table, the exposure is bounded to that one customer's perimeter, not to every customer sharing a vendor pool. This is the shape we ship with MICKAI. Training runs inside a studio inside the operating system, on the customer's own servers, with the Open Audit Record signing every action that touches a training input, so a regulator can verify offline exactly what was ingested, when, by which agent and under what policy.

Two training postures, side by side

PropertyCentralised vendor poolOn-premise per-tenant
Where data lives during trainingVendor infrastructureCustomer infrastructure
Aggregation across customersYes, into one poolNone
Exposure from one misconfigurationEvery contributorOne tenant
DPIA position on residual riskRequires justificationContained by construction
Audit trail of what was ingestedVendor logs, revocablePost-quantum signed, offline verifiable
Model ownershipVendorCustomer
Cross-border transfer riskPresent by defaultNone if the appliance stays in country

What comes next in the industry

MCI will not be the last incident of its shape. Any vendor that trains a shared model on customer-supplied data is running the same architecture, with the same single-misconfiguration failure mode. The June and July 2026 disclosures will be cited in every serious procurement conversation for the rest of the year, and they will be cited in DPIAs, in ICO correspondence and in board-level risk registers. The point regulated buyers should take is not that Meta was uniquely careless. It is that the architecture chosen for MCI is the same architecture chosen by the mainstream AI training industry. That architecture concentrates a company-scale exposure into a vendor's misconfiguration budget. On-premise per-tenant training removes the concentration.

What data was in the 45,000 exposed tables?

According to disclosures across late June and early July 2026, the tables held collected keystrokes, mouse movements, screen content, private internal chat, meeting transcriptions, performance records and AI prompts submitted by Meta employees, captured through the MCI tracking software installed on US company laptops from April 2026. Employees speaking publicly said personal tax and medical information was also present in some records.

Did any customer or user data leak in the Meta MCI incident?

The exposed material was internal Meta employee data, not customer or advertiser data. The point for external regulated buyers is architectural rather than personal. The failure mode that exposed 45,000 tables of employee data is the same failure mode a vendor would have if it applied the same aggregate-then-share approach to customer data.

Is an on-premise training deployment enough by itself?

No. Removing the vendor pool eliminates the cross-customer exposure, but the on-premise pipeline still needs signed ingestion, an offline-verifiable audit record of every training input, a documented purpose limitation and a defined retention schedule. Without those, an on-premise pipeline is still a DPIA exposure. It is a smaller one.

How should a regulated enterprise train an AI on employee data without creating a single blast radius?

Train on hardware inside the enterprise perimeter, per tenant or per legal entity. Do not aggregate the training corpus across entities. Sign every ingestion event into a tamper-evident record an auditor can verify offline. Hold the model weights and the training corpus under the enterprise's own keys, not the vendor's. Publish a DPIA that names the exposure of the chosen architecture in plain language. That posture keeps the failure mode bounded to a single enterprise, and it survives an ICO enquiry in a way that a centralised vendor pool does not.

What is MICKAI?

MICKAI is a British-built Sovereign Intelligence Operating System. It runs entirely on hardware the customer owns, on premise and air-gapped, with no data egress. Every consequential action is signed into the Open Audit Record, a post-quantum, tamper-evident ledger any outside party can verify offline in a browser. MICKAI ships 63 studios on one operating system, with 10 production-ready at launch and 53 in development, and is protected by 104 filed UK patent applications across 2,340 claims.

Subscribe
Get every new Mickai article by email.

Long-form essays on sovereign AI from Micky Irons. One email per article. No tracking, no marketing, no third parties. Every email includes a one-click unsubscribe link.

Prefer RSS? Subscribe at /articles/feed.xml.

Originally published at https://mickai.co.uk/articles/meta-mci-45000-tables-centralised-ai-blast-radius. If you operate in a regulated sector or want sovereign AI on your own hardware, the audit form on mickai.co.uk is the entry point.
More articles