Featured image: conceptual AI-generated illustration created for this article; it is not documentary evidence.
An AI system passes its evaluation, meets the launch threshold and enters production. That sounds like the end of testing. In reality, it begins a different experiment—one involving changing users, new data, unexpected behaviour, software updates, human workarounds and consequences that a test set could not fully reproduce.
A pre-deployment score answers a bounded question about a system under specified conditions. Production asks a harsher question: does this whole sociotechnical system continue to work, for the people actually using it, as its environment changes?
That is why post-deployment monitoring is not an optional analytics dashboard. It is part of the evidence required to keep using an AI system responsibly.
Passing the test establishes a baseline, not a lifetime guarantee
Models are evaluated on data selected before launch. The future may not resemble that sample. Customer language changes. Fraud tactics adapt. Seasonal demand moves. A camera or sensor is replaced. Staff begin interpreting an assistant’s suggestions differently. A third-party model is silently updated. A policy changes the meaning of an acceptable outcome.
Any of these changes can separate the deployed system from the assumptions under which it was approved.
The US National Institute of Standards and Technology’s March 2026 report, Challenges to the Monitoring of Deployed AI Systems, describes post-deployment monitoring as necessary for validating whether an AI system operates reliably and as expected in real-world conditions. It also makes an important research point: the field still contains gaps, barriers and open questions. Monitoring is necessary, but there is no universal panel of metrics that makes every AI use safe.
That uncertainty should produce better experimental discipline, not paralysis. A team can define what it believes should remain stable, watch for evidence that contradicts those beliefs, and decide in advance what action different signals will trigger.
Start with a monitoring claim
“We monitor the model” is not a testable claim. A monitoring plan should state what continued success means in context.
For example:
The system should continue to route ordinary customer enquiries accurately enough to reduce waiting time, without increasing unresolved cases or producing materially worse outcomes for users who write in less common dialects.
That sentence names a function, a benefit, a failure condition and an uneven-impact risk. It is more useful than a generic target such as “maintain 90% accuracy”, because aggregate accuracy can remain stable while one group, task or harm category deteriorates.
Turn the claim into a small set of questions:
- Is the input population changing?
- Is task performance changing?
- Are errors becoming more severe or concentrated?
- Are humans relying on, correcting or bypassing the system differently?
- Are costs, latency and availability still acceptable?
- Are complaints or incidents revealing harms the original evaluation missed?
Monitor the system, not only the model
A model metric can look healthy while the service around it fails. Imagine a document assistant whose factuality score is unchanged, but staff now paste confidential material into prompts because the workflow is inconvenient. Or a classifier whose technical accuracy is stable, while employees increasingly accept its recommendation without checking because workloads have risen.
Post-deployment monitoring therefore needs several layers.
1. Inputs and context
Track whether production inputs still resemble the conditions used for testing. Useful signals may include missing fields, language mix, device type, image quality, task length, topic distribution or the frequency of unusual cases. The aim is not to collect everything. It is to identify changes relevant to the system’s documented limits.
2. Outputs and performance
Measure the outcomes that matter for the task: false positives and false negatives, unsupported statements, successful completions, calibration, abstentions, latency or human correction rates. Where ground truth arrives late, use carefully labelled leading indicators rather than pretending that immediate certainty exists.
3. Uneven performance
Break down results across relevant populations, languages, locations, task types or operating conditions where it is lawful and proportionate to do so. An overall average can conceal a sharp decline in a smaller segment. The UK Information Commissioner’s Office advises regular monitoring or testing of AI systems to keep outputs within statistical-accuracy tolerances and detect unfair or discriminatory outcomes. Its statistical-accuracy audit framework also warns that missing model drift can lead to unfair outputs.
4. Human interaction
Observe when users override, ignore, escalate or blindly accept the system. A high override rate may show that the tool is wrong, poorly explained or badly placed in the workflow. A very low override rate is not automatically reassuring: it may reflect genuine reliability, automation bias, weak training or a user interface that makes disagreement difficult.
5. Incidents and near misses
Record confirmed harms, complaints, security events and regulatory breaches, but also capture near misses: an incorrect answer caught before it reached a customer, a harmful recommendation blocked by a reviewer, or a data leak prevented by an access control. Near misses reveal weak defences before consequences grow.
6. Dependencies
Monitor model versions, prompts, retrieval sources, external APIs, policies and human staffing. When one component changes, the effective system changes. Without version records, teams may notice an outcome shift but be unable to reconstruct its cause.
Drift is not one phenomenon
The word drift often compresses several different changes:
- Data drift: the distribution of inputs changes—for example, a new customer population uses the service.
- Concept drift: the relationship between inputs and correct outcomes changes—for example, fraud patterns evolve.
- Performance drift: measured results worsen, whether or not the cause is known.
- Behavioural drift: people change how they use or respond to the system.
- Dependency drift: an upstream model, database, policy or interface changes.
A shift is not automatically harmful. A model may encounter a broader population while performance stays acceptable. Conversely, outputs may become unsafe without a dramatic statistical change in inputs. Treat drift alerts as prompts for investigation, not automatic proof of failure.
NIST’s AI Risk Management Framework playbook recommends comparing production metrics with pre-deployment measurements and monitoring performance in real time where appropriate. It describes drift as a condition in which systems no longer meet original assumptions and limitations. See the playbook’s Measure guidance.
Set thresholds with actions attached
A dashboard without a response plan can turn monitoring into theatre. For every important signal, define who receives it, how quickly it is reviewed and what happens next.
Use graduated responses:
- Observe: annotate a small movement and collect more evidence.
- Investigate: sample cases, compare segments and inspect recent changes.
- Constrain: reduce scope, require stronger human review or disable a vulnerable feature.
- Pause: stop the affected workflow while the issue is assessed.
- Withdraw: retire the system when risks cannot be controlled or the original benefit no longer justifies them.
Thresholds should reflect impact, not only frequency. One low-severity formatting error and one unsafe clinical recommendation should not enter the same queue merely because both count as “one failure”.
Do not let the vendor own all the evidence
When an organisation buys an AI service, it may lack access to training data or internal model details. That makes contractual monitoring more important, not less.
Before deployment, agree what information the provider will supply: version-change notices, service incidents, evaluation summaries, known limitations, deletion arrangements, security responsibilities and performance-review support. Separately measure the outcomes visible in your own workflow. A vendor’s global benchmark cannot tell you whether the tool works for your users, staff and risk tolerance.
The ICO’s guidance on accuracy and statistical accuracy notes that regular updates and reviews should be agreed when development is outsourced or an AI solution is purchased, including protection against changing population data and concept or model drift.
Preserve the distinction between signals and conclusions
A monitoring report should clearly label:
- Observation: “Human overrides rose from 8% to 14% after the interface update.”
- Interpretation: “Reviewers may have less confidence in the revised explanation panel.”
- Alternative explanation: “The new interface may simply make overrides easier to record.”
- Decision: “Sample 100 overridden cases and interview ten reviewers before changing the model.”
This discipline prevents a moving graph from becoming a dramatic but unsupported story. Monitoring creates evidence for inquiry; it does not remove the need for causal reasoning.
A practical post-deployment monitoring card
- Purpose: What job is the AI system meant to perform, and for whom?
- Baseline: Which pre-deployment results and conditions will production be compared with?
- Limits: Which uses, populations or inputs were excluded or weakly tested?
- Signals: Which input, output, impact, human-behaviour, incident and dependency measures will be collected?
- Segments: Which breakdowns are needed to detect uneven performance?
- Cadence: What is watched continuously, weekly, monthly and after a change?
- Ownership: Who investigates each alert and who can pause the system?
- Response: What thresholds trigger observation, investigation, constraint, pause or withdrawal?
- Feedback: How can affected people challenge outcomes or report problems?
- Change control: How are versions, prompts, data sources and policy changes recorded?
- Review: When is the whole monitoring plan reassessed?
- Exit: How can the organisation stop safely and preserve necessary records?
Monitoring is part of the product
The UK Government’s Introduction to AI Assurance describes assurance as measuring and communicating whether systems meet relevant criteria in context. That framing matters. A model is not trustworthy merely because a launch document used the word. Trustworthiness must be supported by continuing evidence.
The research lesson is simple but demanding: deployment changes the object being studied. The model now interacts with institutions, incentives, people and evolving environments. A responsible team does not ask only, “Did it pass?” It asks, “What would tell us that it is no longer working as intended—and are we prepared to act when that evidence appears?”
Primary guidance and further reading
- NIST, March 2026: Challenges to the Monitoring of Deployed AI Systems.
- NIST AI Risk Management Framework: Measure Playbook.
- Information Commissioner’s Office: Statistical accuracy audit framework.
- Information Commissioner’s Office: Accuracy and statistical accuracy in AI.
- UK Government: Introduction to AI Assurance.
Discover more from Marychuks.com AI, Psychology, Business & CreativeVerse
Subscribe to get the latest posts sent to your email.