When It Goes Wrong
Agents fail in four predictable ways. Learn to spot each one in sixty seconds, fix it with one line, and know when to just do the job yourself.
Earlier this year I handed an AI agent (software that uses your computer to finish a task you describe) a job I knew inside out: sort a batch of old Collab365 articles into the right categories, ready for the move to our new site. I had done this kind of sorting myself for years. I wrote what I thought was a decent brief and let it run.
It came back fast. And the receipt, the end-of-run report an agent writes to show you what it did, read beautifully. Every article accounted for. Every category named. A tidy closing summary, quietly pleased with itself.
If confidence were the measure, it was the best work I had seen all week.
Then I did the boring bit. I opened the folders and spot-checked the results against the originals.
Most of it was genuinely good. But dotted through the batch were articles filed somewhere plausible and wrong. The agent had hit content that did not fit any category neatly and, instead of stopping to ask me, it had invented a rule of its own and applied it with total conviction.
Nothing broke. Nothing was lost. Review caught it before anything touched the live site, and putting it right cost me part of an afternoon.
But here is the sentence I want you to take from this whole chapter: the receipt gave no hint that anything was wrong. The failure was invisible until I looked.

The price of finished work
Chapters one to six taught you to delegate. This chapter teaches the other half of the job: supervising. Because the shift this guide is about, from answers to finished work, comes with a bill attached.
A chat answer you can shrug at. If it is wrong, you notice while you are doing the typing yourself. Finished work is different. The typing already happened, and any mistakes are baked into a file that looks done.
Remember the test from chapter 1: who does the typing? When the agent types, you inherit a new job title. Not typist. Reviewer.
Every manager who has ever signed off a report knows that trade, and most will tell you it is a good one. But only if you actually read what you sign.
The good news: agent failures are not random. After months of running our whole business this way, I can tell you that nearly every failure I have seen falls into one of four patterns. Learn the four, and a scary black box turns into a checklist.
Misunderstood brief. Wrong assumption. Confident nonsense. Half-finished job. Nearly every agent failure you will ever meet is one of these, and each has a sixty-second check.
The four ways it goes wrong
1. The misunderstood brief
The most common failure, and the least mysterious: the agent did what you said, not what you meant.
Say you have the competitor scan from chapter 2: twelve competitor websites turned into a short briefing note. You asked for “a summary of these sites”. What you wanted was two pages your manager could read on the train. What came back is a wall of bullet points, one indistinguishable block per site. Technically, a summary.
In the receipt: everything ticked off, no errors, no questions, and a deliverable that is the wrong shape. The work is complete. It is just not the work.
The sixty-second check: open the result and ask one question. Could I send this, exactly as it stands, to the person who needs it? Not “is it impressive”. Could I send it.
The fix: don't argue, show. “Not the format I need. Here's an example of what good looks like. Rework yours to match it: same facts, that shape.” One example beats three paragraphs of description. You met this move in chapter 5; this is where it earns its keep.
2. The wrong assumption
This was my mis-sorted archive. Halfway through a task, the agent hits a question the brief does not answer. Instead of asking, it guesses. Sensibly, silently, and sometimes wrongly.
In the receipt: a decision you never made. A folder name you never gave it. A phrase like “I assumed” or “I treated X as Y”. Decent agents narrate their guesses, which is exactly why receipts are worth reading.
The sixty-second check: scan the receipt for choices you do not recognise. Anything your brief did not specify, the agent decided. Find those decisions and hold each one up against what you actually wanted.
The fix: “You assumed X. It's actually Y. Redo only the parts that assumption touched and leave everything else alone.”
And the prevention, which you already know from chapter 5: end every brief with the standing escalation rule. “Ask me before you delete, send or spend. When unsure, stop and ask.” One rule, and this whole failure mode mostly disappears.
3. Confident nonsense
The third mode is the one that unsettles people most, so let's describe it plainly. Sometimes an AI states something flatly wrong with total confidence. Not a typo. Not a rounding slip. A figure that appears in no source document, a name nobody at the company has, a “fact” assembled out of thin air because it sounded like the kind of thing that should be true.
Numbers and names are where it stings. A made-up sentence in a paragraph reads oddly and gets caught. A made-up number in a spreadsheet looks exactly like a real one.
The industry word for this is hallucination. You will see it on every AI page you ever read, so now you know what it means: wrong, and sure of itself. The word matters less than the habit that defends against it.
Why does it happen? Because underneath, these tools work by producing the most plausible next words, not by looking truth up in a ledger. Plausible and true overlap most of the time. When they do not, you get a confident wrong answer delivered in exactly the same tone as every right one.
Which is the practical point: the tone carries no signal. Only checking does.
The sixty-second check: pick three numbers or names from the output and trace each one back to the source file it supposedly came from. All three trace cleanly? Spot-check a couple more and move on. Even one fails to trace? Stop trusting the document and check everything.
One rule with no exceptions: for anything involving money, invoices, budgets, payroll, quotes, every number gets checked against its source before it goes anywhere. Not because agents are usually wrong. Because “usually right” is not a standard anyone accepts in finance.
The fix: “Where in the source does X come from? Show me the exact line. Correct anything you cannot trace to a source, and mark anything uncertain as unverified rather than guessing.”
4. The half-finished job
The quietest failure. The agent runs out of steam, or quietly narrows the scope, and reports success anyway. You asked for all of them. It did most of them.
In the receipt: the maths does not add up. Fewer files out than in. A summary that says “processed the invoices” without saying how many. A section that starts strong and trails off. The word “sample” anywhere you asked for everything.
The sixty-second check: count. Items in against items out. If you gave it a folder of files and got back a spreadsheet of rows, the two numbers should match, and the receipt should state both. Any decent agent can report them; make that a standard line in your briefs.
The fix: “The brief covered every file in the folder. List anything you skipped or shortened, and why. Then finish the rest.”
| What the receipt shows | The 60-second check | The fix phrase | |
|---|---|---|---|
| Misunderstood brief | All done, no errors, but the deliverable is the wrong shape | Could I send this, as it stands, to the person who needs it? | “Here's an example of what good looks like. Rework yours to match it.” |
| Wrong assumption | A decision you never made: “I assumed...”, a name you never gave | Scan the receipt for choices you do not recognise | “You assumed X. It's actually Y. Redo only the parts that touched.” |
| Confident nonsense | A figure or name that appears in no source document | Trace three numbers or names back to their sources | “Show me where X comes from. Correct anything you cannot trace.” |
| Half-finished job | Items out do not match items in; “sample” where you asked for all | Count. In against out. The numbers should match. | “List what you skipped and why, then finish the rest.” |
Blame the brief before the tool
Here is the default posture I want to hand you, because it took me an embarrassing number of failed runs to learn it: when an agent gets it wrong, re-read your brief before you blame the tool.
When a temp fumbles a job on day one, a decent manager does not ring the agency first. They re-read the handover note, and most of the time they find the hole. Same here. Most failures I see trace back to a brief that a sharp human assistant would also have fumbled, just more slowly and with more apologising.
So the standard move after a failed run is not starting over, and it is not typing an essay. It is the re-brief: one message that names what is wrong, precisely, and tells the agent to fix only that.
Notice what the template refuses to do. It does not restate the whole brief, and it does not invite the agent to “improve” things you were happy with. Fix only that is what stops a small correction turning into a fresh set of surprises.
Will the first run succeed without any of this? Honestly: it depends on the task. A clear brief on simple files usually lands first time. Fuzzier work, judgement-heavy sorting, anything with messy sources, often needs a round of feedback. What I can tell you from daily use is that one precise re-brief fixes most of what a spot-check finds.
And two failed re-briefs is a signal. The signal is not “try a third”.
When to just do it yourself
Sometimes the honest answer is: take the task back. Anyone who tells you agents always win is selling something.
Three signs it is time. One: explaining the task is taking longer than doing it. Delegation has overheads. A two-minute job that needs ten minutes of context is not a delegation candidate, this week or possibly ever.
Two: it has failed twice after honest re-briefs. Not vague ones; real ones that named the problem. Some tasks sit past the edge of what today's tools do well, and the edge is real.
Three: a single wrong item would be catastrophic, and checking properly would mean redoing the work anyway. If verifying every line costs the same as writing every line, the agent is not saving you anything. Chapter 6's amber list lives here.
Taking a task back is not defeat. Deciding what not to delegate is a management skill too, and people who are good at it are unsentimental in both directions.
The two-line habit that makes you the supervisor
One last habit, and it is the one that turns this chapter into a professional edge.
Whenever agent-made work leaves your hands and travels to another human, a manager, a client, a supplier, keep a two-line note for yourself: what I spot-checked, and what I didn't. That's it. Two lines, in whatever notes app you already use.
Something like: “Checked every total against the PDFs and all supplier names. Didn't check the address column.”
The note does two jobs. It keeps you honest about how much checking you actually did before hitting send. And on the day someone asks “are these figures right?”, you have a specific answer instead of a hopeful one.
That question is coming to every workplace that adopts these tools. The person with the two-line note is the person trusted with the next, bigger delegation. That is the quiet career case hiding in this chapter: agents do not make checking obsolete. They make the people who check well more valuable.
That is the loop. Read the receipt, spot-check, re-brief or accept, note what you checked. Sixty seconds on most days. Part of an afternoon on a bad one.
Next: the tools themselves, mapped. You now know how to supervise every one of them, including the ones that have not shipped yet.