Why an AI workflow ages differently from ordinary software
Conventional software either runs or fails. A workflow with a language model has a third state: it runs and gets worse while doing so. The outputs are still complete, in the right format and on time, but they meet the rules less often, become longer, vaguer or wrong in one place. Monitoring that only asks whether the workflow is running sees none of this.
The causes are almost always the same five:
- The model: providers improve their models, change the behaviour of a version or retire older versions on announced dates.
- The input data: new product groups, a different data sheet format, more enquiries in a language that used to be rare.
- The connected systems: a new version of the interface, a renamed field in the CRM, a changed request limit.
- The rules: a new legal situation, a new brand voice, a new mandatory term.
- Volume and costs: more tasks, longer inputs, a changed price per request.
None of these causes is an error in the strict sense. They are part of operation, just as an oil change is part of running a machine. That is why every workflow needs monitoring at two levels: is it running technically? And is the result still as good in substance as at acceptance?
The metrics of an AI workflow
The following table is the checklist we use when handing a workflow over for live operation. Not every workflow needs every row, but every row that is missing should be a deliberate decision. You set the thresholds for each workflow, based on the values from the parallel run.
| Metric | What it shows | Response to a deviation |
|---|---|---|
| Throughput: tasks per run or day | Whether input arrives and the workflow processes it completely | Check the source: empty mailbox, expired access, changed export |
| Technical error rate | Aborts, timeouts, interface errors | Retry, throttle, check the provider's status; if persistent, stop the workflow |
| Rejection rate of the checks | Share of outputs that fail automatic checks | If it rises, the model or the input has changed. Look at a sample, tighten the rule or instruction |
| Referral rate | Share of cases the workflow puts before a person | If it rises, rules are no longer working. If it falls sharply, check too: perhaps the workflow is deciding what it should put forward |
| Correction rate at approval | Share of outputs that people change before approving | The best quality signal. Collect only for the workflow as a whole, not per person |
| Format characteristics | Length, language, mandatory terms, character errors over time | If the distribution shifts, the model's behaviour has changed, often before any other metric |
| Deviation from the reference set | How the workflow processes a fixed set of already checked cases today | Recalculate regularly and at every model change; look at deviations individually |
| Sample | Human review of a small number of random outputs each week | Catches what no rule captures. Add the results as new rules or reference cases |
| Run time per task | Whether the workflow finishes within its time window | Check load distribution, caching, choice of model |
| Cost per task | Model costs in relation to volume | Longer inputs, retry loops or a new price; clarify the cause before the invoice arrives |
| Quotas in connected systems | Consumption of request limits in CRM, ERP or interfaces | Adjust caching, throttling, time windows |
| Provider announcements | Retirement dates for model versions, changes to interfaces | Not a measurement but a date in the calendar, with lead time for the switch |
The scene at the beginning would have shown up in the format characteristics row after a few days: the average length of the meta descriptions rises, and the share above the limit grows. A check that rejects overly long texts would never have let them go online in the first place.
Report what deviates, not what is running
Monitoring that sends twenty alerts every morning stops being read after two weeks. Good monitoring stays silent as long as everything is within bounds, and reports the deviation once, to the right place. Three rules have proven effective:
- Throttle alerts: at a spare parts dealer, error alerts are throttled to one per hour per type of error, a dashboard in the back end shows the status, and a weekly report arrives on Fridays.
- Report changes, not states: in our own operations, a nightly check at a hosting provider reported the same known error every night. Since it was switched to reporting changes rather than states, it only reports what is new. You can read about it in the case study on our own operations.
- Provide a stop switch: anyone who sees a deviation must be able to stop the workflow without destroying anything. At the same client a stop switch halts the weekly rollout; if checks fail after the data is inserted, there is a bit-identical rollback instead of publication.
When the model changes
Model changes are not an exception but something you can plan for. Providers usually announce the end of a model version in advance, and new models are often better or cheaper. The switch itself is a small rebuild with a fixed procedure:
- Pin the version: the workflow uses a named model version, not "the latest". That way nothing changes without your decision.
- Keep a reference set ready: a few hundred cases with checked, approved results from the parallel operation, typical cases and the difficult ones.
- Run the new model against the reference set: with the same rules and checks, metrics side by side.
- Adjust rules and instructions: new models often respond differently to the same instruction. Every change is versioned.
- Sample check by the business department: the people who approve today look at a selection before the switch.
- Switch with a way back: the previous version remains operational for as long as the provider offers it.
- Document: the model version is in the log for each task. Anyone who later asks why a text reads as it does can also see which model produced it.
The same procedure applies if you change provider, for example because a model becomes available in an EU data centre or a provider changes its terms. Models are interchangeable; what remains are rules, checks and the reference set. If the operating mode changes in the process, the switch must first be cleared with your data protection officer, because the data processing agreement and data location change with it. Which operating mode suits which data is described in the article On your own servers, in Europe or in the cloud. Whether a model change has to be submitted to the works council again is best settled in advance in the works agreement; how is explained in the article Introducing AI with the works council.
Who is responsible for what in operation
Monitoring without responsibility produces alerts that nobody deals with. Every workflow needs three named roles:
- Business owner on your side: decides on rules, approvals and exceptions, receives alerts about deviations in content and quality.
- Technical operation: receives technical alerts, keeps access, interfaces and model versions up to date. That is your IT or a service provider.
- Escalation: who may stop the workflow when something goes wrong, and who then decides when it runs again.
This includes a short operating manual for each workflow: what each alert means and what to do then. It is part of our handover, together with code, rules and documentation, which belong to you. With it, your IT can run the workflow itself. If you would rather not, you can book ongoing operation: monitoring of all workflows with a report, adjustment when models, systems or rules change, response on the next working day and the same day in the event of an outage, cancellable monthly. Scope and terms are under Approach and prices, and the connections to CRM, ERP and PIM under Systems.
What the law requires for monitoring
For most workflows, such as texts, data and enquiries, the EU AI Act, Regulation (EU) 2024/1689, does not prescribe any particular monitoring; they count as minimal-risk systems. It is different for high-risk systems under Annex III, for example in HR: here Article 26 obliges the deployer to monitor operation and to keep the automatically generated logs for at least six months, insofar as they are under its control. Article 72 obliges the provider to carry out post-market monitoring. The GDPR applies to the logs themselves: they should contain identifiers, not content, and under the principle of storage limitation in Article 5(1)(e) be kept only as long as they are needed. This article reflects the position as of October 2026 and is not legal advice.
What this means for you
An AI workflow is not finished at acceptance; it is in operation. Ask before every build, not afterwards: which metrics are monitored, who receives which alert, is there a reference set and a stop switch, and who handles the switch to a new model when the provider retires a version? If you already have an AI workflow in your company and cannot answer these questions, or are planning one and want to get them right from the start, talk to us in a 30-minute call: book a call.
Further reading
- Case study: weekly rollout with eleven checks, stop switch and rollback
- Case study: our own operations with nightly checks
- Service: texts and content at scale, with checks and a log
- AI data protection in companies: where models should run for which data, on your own servers, in Europe or in the cloud
- Industry: AI for professional services, with nightly checks of customer systems
