Short answer. Monitor a production AI system on 4 fronts at once: model behaviour (accuracy, data drift, override rate), operational health (latency, errors, cost per call), compliance evidence (the EU AI Act Article 72 post-market monitoring plan), and incident readiness (a defined path to report a serious incident within 15 days). The EU AI Act makes this a legal obligation for high-risk banking AI from 2 August 2026, and DORA already requires that the underlying ICT service stay operationally resilient. Monitoring is not a dashboard you check when something feels wrong, it is a standing control with named owners and logged evidence.
An AI model that passed validation on launch day does not stay compliant on its own. Input data shifts, customer behaviour changes, an upstream system starts sending a new field format, and the model that scored 94% accuracy in testing quietly drops to 81% in production. Nobody notices until a regulator, an auditor, or a customer complaint forces the question. Production monitoring exists to catch that drift before it becomes a finding.
In practice, a bank runs monitoring on 4 layers that report into one owner:
Ableneo shipped 34 production AI projects in 2025, and roughly 4 of 5 reached production rather than stalling as pilots. The projects that stay in production are the ones where monitoring was designed in before go-live, not bolted on after the first incident.
For a bank in the EU, monitoring a live AI system is no longer a matter of engineering hygiene. From 2 August 2026, most AI used in credit scoring, fraud detection, and customer risk assessment counts as high-risk under the EU AI Act, and Article 72 obliges the provider to run a documented post-market monitoring system that actively collects and analyses performance data across the system’s lifetime. The monitoring plan is part of the technical documentation an auditor can demand.
DORA raises the second wall. Since January 2025, the ICT services that host a bank’s AI models must meet operational resilience obligations: continuous monitoring, defined incident classification, and strict reporting timelines to the competent authority. An AI system in a bank sits inside both regimes at once. The model has to be monitored for accuracy and bias under the AI Act, and the service running it has to be monitored for availability and resilience under DORA. A bank that treats these as one integrated control, rather than two disconnected teams, is the one that passes the audit.
Monitor production AI on 4 layers at once: model behaviour, operational health, compliance evidence, and incident readiness.
Article 72 requires a post-market monitoring system that is proportionate to the risk of the AI system and documented in a written plan. The system must actively and systematically collect, document, and analyse data on how the model performs across its whole operational life, so the provider can prove continuous compliance with the high-risk requirements. Passive logging is not enough. The obligation is to analyse the data and act on it, and the data can come from the bank’s own telemetry or from the deployers using the system.
In concrete terms, a bank keeps a running record of model performance, detected drift, and the corrective actions taken. When the model drifts from its validated baseline, the monitoring system surfaces it, dates it, and shows the auditor what happened next.
Track fewer metrics, but track them against a defined threshold with an owner. For a credit or fraud model, the core set is accuracy or precision against realised outcomes, input data drift, output distribution drift, human override rate, and false-positive rate. A rising override rate is often the earliest honest signal that people on the floor no longer trust the model, and it shows up before the accuracy numbers move.
Alongside model quality, track the operational signals that keep the service resilient under DORA: latency, error rate, availability, and cost per inference. Cost belongs on the same board as accuracy, because a model that quietly triples its token cost per decision is a governance problem, not just a finance one. Every metric needs a threshold, an alert, and a named person who acts when it breaches.
Accountability sits with the bank, not the vendor, and it is a board-level responsibility under DORA. The management body owns the ICT risk framework, which now covers the AI services running in production. Day to day, monitoring is usually shared: a model owner in the business line watches model behaviour, an ICT operations team watches service health, and a second-line risk function validates that the controls work. The mistake banks make is leaving the AI model in a gap between the data science team that built it and the operations team that runs everything else.
A clear operating model closes that gap. One accountable owner per model, a defined escalation path, and a single monitoring view that both the business and ICT operations read from the same source. This is the same accountability question a DORA auditor asks about sign-off, extended into the running life of the system.
Two clocks run in parallel. Under EU AI Act Article 73, the provider of a high-risk system reports a serious incident to the market surveillance authority immediately after establishing the causal link, and no later than 15 days after becoming aware of it, with a shorter window where the incident is severe. Under DORA, a major ICT-related incident triggers its own reporting chain: an initial notification within hours of classifying the incident as major, an intermediate report within 72 hours, and a final report within one month.
The practical consequence is that a bank needs the classification decision to happen fast and consistently. If nobody can say within the first hours whether an AI failure is a “serious incident” or a “major ICT incident”, the reporting clock is already running down. Pre-agreed thresholds and a rehearsed escalation path are what turn a 2 a.m. anomaly into a controlled report rather than a missed deadline.
Validation happens once, before deployment, and asks whether the model is fit to go live. Monitoring runs continuously, after deployment, and asks whether the model is still fit today. A model can pass independent validation cleanly and then drift into non-compliance 3 months later because the world it was trained on has moved. The EU AI Act treats both as obligations: validation proves the model met the bar at launch, and post-market monitoring proves it keeps meeting the bar. A bank that funds validation but underfunds monitoring has bought a certificate, not a control.
Ableneo builds monitoring into the architecture of an AI system from the first design session, because a model without observability is a liability the moment it reaches production. Across 34 production AI projects in 2025 for clients including banks and insurers in the ČSOB, Erste, and UNIQA orbit, the pattern holds: the systems that survive audits and stay in production are the ones where drift, cost, and compliance evidence are monitored as one control with a named owner. Ableneo pairs that engineering discipline with the governance layer a regulated buyer needs, so the answer to “how do you know your live AI is still compliant” is a running record, not a hopeful assumption. See how this fits an end-to-end approach on the Ableneo AI transformation page.
Key takeaways
Planning AI in a regulated business? Ableneo takes systems from classification to governed production.