Skip to content
0800 374 775

Automation that survives contact with the exception

The happy path is the easy part. What decides whether an automation is still running in a year is what it does when the input is wrong, and whether anybody finds out.

Automating a process is straightforward while every input looks like the example you built it against. Then a supplier sends a PDF with two invoices in it, a quantity arrives as “approx 40”, and somebody enters a date in the American order.

How the automation behaves at that moment determines whether it saves the business time for years or gets quietly abandoned in month four.

The three ways they die

It fails silently. The job stops, or skips the record, and nothing tells anybody. Three weeks later someone notices the numbers are wrong and the trust is gone permanently. This is the worst outcome, and it is the most common.

It guesses. Rather than stopping, it makes a decision about ambiguous input and carries on. Now there is bad data downstream, generated confidently, and finding it means checking everything.

It needs a person to babysit it. Someone has to check it ran, re-run it when it did not, and handle the rejects. That is still worth having, but it is a smaller win than was promised and it usually ends up on the desk of whoever is least able to refuse.

What to design instead

Decide the failure behaviour before the happy path. For each step: if the input is not what we expect, do we stop, park it, or proceed with a default? All three are valid answers and the wrong one is not having chosen.

Park, do not drop. Anything that cannot be processed goes to an exception queue that a named person looks at, with the reason attached. The queue is the interface between the automation and human judgement, and it is the part that makes the whole thing trustworthy.

Be loud about silence. The alert that matters is not “the job failed”, it is “the job has not run”. A process that should produce forty records a day and produced none should raise something, and a process that produced four hundred should too.

Make it idempotent. Running it twice must not create two invoices. This sounds obvious and it is the most common defect we find, because it only bites during a recovery, which is exactly when things get run twice.

Log enough to answer “why did it do that”. Six months on, somebody will ask why a particular record was handled the way it was. If the answer requires reading the code and guessing, the automation cannot be maintained.

The measurement worth taking first

Before automating anything, find out where the hours actually go. Not by asking, by watching, because people describe the process as designed rather than as performed.

Frequently the answer is not the task at all. It is the waiting: the handover between two systems, the approval sitting in an inbox, the report that has to be requested. Automating the typing when the delay is the approval saves ten minutes of a two day cycle.

That is why we would rather map the process before quoting the build. The most valuable automation is often much smaller and further upstream than the one the business asked for.

And the unglamorous truth

The best automations are boring. They handle one well-understood, high-volume, low-judgement task, they park anything unusual, and they tell someone when they have not run.

Ambitious ones that make decisions across several systems demo brilliantly and are the ones nobody trusts by winter. If you have to choose, choose the boring one and let it earn the case for the next.

Next step

Recognise any of this? Let's talk.

We respond within one business day.