Does DORA require threat-led penetration testing of your AI systems?
DORA does not name AI, but AI supporting critical functions falls inside the resilience testing programme, including TLPT for designated entities.
Not by name, but often in effect. DORA, in force since 17 January 2025, requires every financial entity in scope to run a digital operational resilience testing programme covering the ICT systems that support critical or important functions, and it mandates advanced threat-led penetration testing (TLPT), built on the TIBER-EU model, for entities designated by their authorities. DORA never mentions AI. But when an AI system supports a critical or important function, it sits inside the testing perimeter and must be exercised, observed and evidenced like any other ICT asset.
The question matters in 2026 because AI has left the pilot phase. Financial entities now run models inside fraud detection, credit workflows, client communications and operational monitoring, and supervisors have started asking how those systems are tested rather than whether they exist.
What does DORA actually require for resilience testing?
DORA sets out a two-tier testing regime. Every financial entity in scope must maintain a digital operational resilience testing programme, proportionate to its size and risk profile, covering the ICT systems that support critical or important functions. The programme includes vulnerability assessments, scenario-based testing and penetration testing as appropriate. On top of that baseline, entities designated by their competent authorities must carry out threat-led penetration testing at least every three years, exercising live production systems against realistic adversary behaviour modelled on current threat intelligence.
Does DORA ever mention AI systems?
No. The regulation is deliberately technology-neutral. Its unit of analysis is the function, not the technology: if an ICT system supports a critical or important function, it is in scope whether it is a database, a payments switch or a model serving inference. That neutrality cuts both ways. A firm cannot argue its AI layer is exempt because the text never names it, and a supervisor does not need new rules to ask how the AI inside a critical function was tested. As AI moves deeper into decisioning and client-facing workflows, more of it crosses the critical or important threshold every year.
Who actually has to run TLPT?
Not every firm. TLPT applies to financial entities designated by their authorities, typically the larger and more systemically significant institutions, following the TIBER-EU model of intelligence-led red teaming. Smaller entities remain subject to the baseline testing programme under the proportionality principle. The distinction matters in both directions: presenting TLPT as a universal duty overstates the regulation, but treating the baseline as optional understates it. Every in-scope entity must be able to test the ICT behind its critical functions, and that baseline is where AI deployments commonly struggle.
Why is a cloud AI endpoint hard to test threat-led?
Threat-led testing assumes you can instrument the target. A red team exercising a critical function needs to observe the system from inside: its processes, its logs, its failure modes under attack. A cloud AI service ends at an API. The tester can probe the endpoint, but cannot observe the model runtime, cannot inspect the serving infrastructure, and usually cannot attack it at all without breaching the provider's terms of service. The provider's own assurance reports describe the provider's controls, not your deployment's behaviour under your threat scenarios. What remains is testing the wrapper around the AI rather than the AI itself.
“A financial entity that cannot instrument its own AI layer cannot honestly claim to have tested it.”
What does a testable AI layer look like?
We suggest a simple instrumentation test with three questions. Can your red team observe the AI system's inputs, outputs and resource behaviour during an exercise without asking a third party? Can you rehearse the failure of the AI component and measure the effect on the critical function it supports? Can you hand the tester signed, timestamped evidence of what the system did once the exercise ends? A deployment that passes all three can be brought inside a DORA testing programme without special pleading. A deployment that fails them is being attested on trust, and trust is precisely what resilience testing exists to replace.
How does a sovereign deployment change the testing question?
Mickai is a Sovereign Intelligence Operating System, a SIOS, that runs offline on operator-owned hardware. Because the entire AI layer sits inside the operator's own perimeter, it can be exercised like any other internal system: red teams can instrument it, scenario tests can degrade it, and recovery can be rehearsed against it. A zero-egress inbound perimeter means there is no external dependency to carve out of the test scope. Every action is recorded in an audit ledger signed under FIPS 204 (ML-DSA), so the evidence a tester needs after the exercise already exists as sealed records, with hardware-attested identity binding each entry to the machine and operator that produced it.
How the whole system fits together is set out at /sovereign-ai, and the film at /film shows the interface in operation.
Frequently asked questions
Does DORA apply to AI systems even though it never mentions them?
Yes, by function rather than by name. DORA covers ICT systems supporting critical or important functions, and an AI system inside such a function is in scope on the same basis as any other ICT asset. The technology-neutral drafting means no exemption exists for models or inference services.
Do I have to run TLPT if my firm has not been designated?
No. Threat-led penetration testing applies to entities designated by their competent authorities. Undesignated firms remain subject to the baseline digital operational resilience testing programme, which still requires proportionate testing of the ICT behind critical or important functions, including any AI components within them.
Can I run threat-led testing against a cloud AI API?
Only in a limited sense. A tester can probe the endpoint your firm consumes, but the model runtime and serving infrastructure belong to the provider and normally sit outside your authorised test scope. That leaves the AI layer of a critical function attested on the provider's own reports rather than exercised under your threat scenarios, which is a weak position in a supervisory review.
How does on-premise AI make DORA testing easier?
Because the operator owns the runtime, the AI layer can be instrumented, degraded and recovered inside the entity's own testing programme. No third-party permission is needed, no scope carve-out is required, and the evidence of system behaviour is generated locally. Mickai adds a sealed audit ledger so the record of what the system did during a test survives as signed evidence.
What evidence will supervisors ask for on AI resilience testing?
Expect requests for the testing programme itself, the mapping of AI systems to critical or important functions, test results with remediation tracking, and proof the tests exercised realistic failure of the AI component. Evidence held in your own records, rather than promised by a vendor, is the difference between answering in days and answering in weeks.