AI for Work
The beginner's guide

The Autonomy Ladder

From one watched task to a Monday report that writes itself. Four rungs of autonomy, what to check at each, and the rule for climbing safely.

AI for Work
Chat gives answers. Agents give finished work.
1Today
You type
AI advises, you do the work
2The shift
It types
Agents finish tasks on your computer
3Chapter 4
First task
Delegated, checked, done
Plain English, real receiptsno vendor theatre, no recycled AI hype
Collab365
Mark Jones
Mark Jones · Collab365

Say there is a report you write every Monday morning. The same project tracker, the same team notes, the same headings as last week. Say it costs you forty minutes before your week has properly started.

By chapter 4 you had delegated one task and watched it run. That was one rung. This chapter is the rest of the ladder: from a task you watch, to a task that finishes while your laptop is shut, to a Monday report that is waiting for you at 8:05, already written.

The ladder ends somewhere genuinely strange: agents that never switch off. We will climb up and look at that rung too, and I will tell you honestly why you should not live there yet.

One promise before we start climbing. Nothing in this chapter asks you for more blind trust. Every rung up demands a stronger review habit, never a weaker one. That principle carries the whole chapter, so let us pin it to the wall now.

Never less review

Each rung of autonomy you climb demands a better receipt and a sharper check, not a blinder eye. Autonomy without review is how the horror stories happen.

A four-rung autonomy ladder from one watched task to triggered always-on work
Four rungs. Each one changes exactly one thing: how much happens without you in the room.

The four rungs at a glance

Rung 1: one task, watched. Rung 2: one task, unattended. Rung 3: scheduled and recurring. Rung 4: triggered and always-on.

Everything else you have learned comes up the ladder with you. The briefing skills from chapter 5, the traffic lights from chapter 6, the failure-spotting from chapter 7. The rungs only change how much happens while you are not looking.

Rung 1: one task, watched

This is where chapter 4 left you. You paste the brief, the agent works, and you sit there like a slightly nervous manager on someone's first day.

You see every permission prompt, the tool asking “may I?” before it touches anything. You answer its questions live. If it wanders off course, you stop it mid-stride.

What goes wrong at this rung is almost always a brief problem, which is why chapters 5 and 7 exist. The stakes stay low precisely because you are watching.

The check is the one you already know: read the receipt, the agent's own written account of what it did, then spot-check three items against the source.

Do not rush off this rung. It is where you learn your tool's personality: what it asks about, what it assumes, where it stumbles. That knowledge is the foundation for everything above.

Rung 2: close the laptop

The second rung feels like nothing and changes everything. You hand over the brief, and then you leave.

Both of the big beginner tools do this today. Claude Cowork can run a task unattended inside the folders you have given it. ChatGPT Work runs its tasks in the cloud, so the work carries on after your laptop lid clicks shut, and it holds sensitive actions until you come back to approve them.

Here is what actually changes: nobody is there to answer questions. At rung 1, the agent hits an ambiguity and asks you. At rung 2 it must either wait, stop, or guess.

The guess is the danger. Say it finds two spreadsheets with similar names, picks the wrong one, and builds twenty minutes of tidy, confident work on top of the wrong file.

The safety rail is the standing escalation rule from chapter 5: “Ask me before you delete, send or spend. When unsure, stop and ask.” At this rung it stops being a nice-to-have.

Plus, for an unattended run, one task addition: when you stop, leave a note, because nobody is there to hear the question.

And the review grows to match. When you come back, read the whole receipt, not the skim you could get away with while watching. Look specifically for decisions it made without you. Every fork it met alone is a place to check.

Rung 3: the Monday report writes itself

Rung 3 is the one this chapter is named for. You stop starting the task at all.

Back to that Monday report. You write the brief once, set a weekly schedule against it, and retire from the assembling business. Monday, 8:00, the agent opens the tracker, the notes and last week's report, and drafts the new one. By the time you sit down with your coffee, it is waiting.

Most of the tools from chapter 8 do this today. ChatGPT Work has Scheduled Tasks, which can run a task once, on a schedule, on a trigger, or as a monitor. Claude Cowork runs scheduled tasks in the folders you have chosen. The Gemini app has Scheduled Actions on its Pro and Ultra tiers, and Claude in Chrome can repeat a browser task daily, weekly or monthly while Chrome is open.

Copilot splits the job in two. Agent Mode works on the file with you, inside Word, Excel or PowerPoint, and has no scheduler; scheduling lives in Copilot Chat as Scheduled Prompts, where any prompt can run on a recurrence, up to ten per person.

Here is the Monday report brief in full: the chapter 5 brief, now on a schedule. Notice how much of it exists purely to handle the fact that you will not be there.

monday-report-weekly.txt
Produce my weekly status report. (The Monday 8:00am schedule is set in the tool, not in this brief.) OUTCOME A one-page status document named 'Status - week of [Monday's date]', saved into the Reports subfolder of this folder. INPUTS Only these three files in this folder: the project tracker spreadsheet, the shared team notes document, and last week's status report. Use last week's report as the layout template and match its headings exactly. CONSTRAINTS Read-only on the tracker and the notes: never edit them. Create the new report as a new file and never overwrite last week's. Do not email, share or send anything to anyone. Work only inside this folder. DEFINITION OF DONE Four sections completed: Done, In progress, Blocked, Next week. Risks belong under Blocked. Every line traceable to a row in the tracker or a dated entry in the notes. The date range covered is shown under the title. One page, no padding. IF YOU ARE NOT SURE If the tracker has not been updated since the last run, if a number in the tracker disagrees with the notes, or if a section would be empty: do not guess. Write what you could not verify in bold at the top of the draft, and stop there.

Read that last section again. At rung 2 the escalation rule gained “stop and leave a note”. At rung 3 it is protecting every future Monday too, because a scheduled task will happily repeat a mistake weekly, on time, forever.

What changes at this rung is the enemy. Rungs 1 and 2 fail loudly, in front of you or in a receipt you read the same evening. Rung 3 fails by drift.

Drift is the world changing underneath a standing brief. Say a colleague renames the tracker. Say someone adds a column, or the notes stop being updated over the holidays. The agent, doing exactly what it was told, produces a confident report built on three-week-old numbers.

So the rung 3 ritual has three parts. Never use a scheduled output unread. Compare this week's report with last week's, because drift shows up as a strange difference or a suspicious sameness. And check any number you are about to repeat to another human back to its source.

Five minutes of review instead of forty minutes of assembly. Still a bargain. But those five minutes are not optional, and they never become optional.

Here is what the first scheduled Monday actually feels like. You open the Reports folder and there are two things waiting: the draft, and the receipt describing where every line came from.

You read the receipt first, then the draft. You check the three numbers you would be most embarrassed to get wrong. You fix one clumsy sentence, note that the agent flagged an empty Blocked section rather than padding it, and send the report.

That flagged gap is the system working. A scheduled agent that says “I could not verify this” is worth ten that always come back confident. Treasure the tools, and the briefs, that produce honest gaps.

Illustrative weekly task schedule filled in for Monday at 08:00 and not yet saved
Illustrative reconstruction, not a product screenshot. The brief is set to repeat every Monday at 08:00 but has not been saved yet.

Rung 4: triggered and always-on

The top rung removes the schedule too. Nothing runs on a timer. Instead, something happens in the world and an agent reacts to it. A trigger is a standing instruction: when this arrives, do that.

You can watch the market building this rung right now. ChatGPT Work's Scheduled Tasks already include trigger and monitor modes. Grok Bot, in beta since August 2026, gives each task its own persistent computer in the cloud, with its own logins, working whether you are at your desk or asleep. Google announced Gemini Spark in May 2026, an always-on agent that works across its apps, currently in beta and US-only.

I will be straight with you. Rung 4 is where the market is heading in 2026. It is not where you should live yet. Most readers of this guide belong on rungs 2 and 3 for now, and that is not a consolation prize.

Two new things go wrong at this height. The first is volume: an always-on agent produces work faster than a busy person reviews it, and unreviewed output quietly hardens into trusted output.

The second is subtler. Whatever triggers the run was written by someone else. An email arriving in an inbox is a stranger's text, not your instruction. A triggered agent must treat what arrives as material to summarise, never as orders to obey, and nothing it produces should reach another human without your eyes on it first.

So the rung 4 rule is the strictest on the ladder: everything rung 3 asks, plus a hard stop between the agent and other people. If that sounds like a lot of supervision for something sold as “autonomous”, you have understood this chapter.

The rule that never changes

Notice what happened as we climbed. The agent did more and more on its own, and your review ritual got more structured, not less.

Watched runs get a skim and a spot-check. Unattended runs get the full receipt. Scheduled runs get a comparison against last week and a numbers check. Always-on agents get all of that plus a locked door between them and your colleagues.

This is the opposite of how most people imagine automation, and it is the difference between the people quietly saving hours every week and the horror stories. More autonomy earns more scrutiny. There is no rung where you stop reading the receipts.

What it looks likeTools that do it todayReview ritual required
Rung 1: watchedYou start it, watch it work, and answer its questions live.Any chapter 8 tool: Claude Cowork, ChatGPT Work, Copilot Agent Mode, Gemini, CometRead the receipt, spot-check three items against the source.
Rung 2: unattendedYou start it, then close the laptop. It finishes alone, or stops and leaves a note.Claude Cowork (unattended runs); ChatGPT Work (cloud-run, holds sensitive actions for approval)Read the full receipt on return. Check every decision it made without you.
Rung 3: scheduledIt starts itself on a schedule you set once. The Monday report writes itself.ChatGPT Work Scheduled Tasks; Claude Cowork schedules; Gemini app Scheduled Actions (Pro and Ultra tiers); Copilot Chat Scheduled Prompts; Claude in Chrome recurring tasksNever use an output unread. Compare against last run, check numbers back to source.
Rung 4: triggered, always-onNo schedule. Something arrives or changes, and the agent acts on it.ChatGPT Work triggers and monitors; Grok Bot always-on cloud agents ($200/mo); Gemini Spark (beta, US-only)Everything rung 3 asks, plus a hard rule: nothing reaches another human unreviewed.
The autonomy ladder. Climb one rung at a time, and only after ten clean runs on the rung you are standing on.Tool facts checked: 24 August 2026

What the rungs cost

Cheerful news first. Rungs 1 to 3 are almost certainly inside a plan you already pay for if you followed chapter 8. Claude Pro at $20 a month includes Cowork, with its unattended and scheduled runs. ChatGPT Plus at $20 includes ChatGPT Work and Scheduled Tasks. Gemini's Scheduled Actions come with its AI Pro tier at around $20.

Rung 4 is where the bill changes. Grok Bot costs $200 a month standalone after a 14-day trial. Gemini Spark needs an AI Ultra plan, and those start at $100 a month.

Watch for usage caps in the small print too. Vendors meter the heavier agent features by tier: Gemini's Auto Browse in Chrome, for example, allows 20 agent tasks a day on the Pro tier and 200 on Ultra. A weekly report will never notice a cap. An always-on agent might, which tells you who these limits are really for.

That pricing gap is the market telling you the same thing this chapter is: the always-on rung is early, expensive, and aimed at people who already review like professionals.

10 clean runs

The toll for the next rung: ten runs at your current rung with nothing surprising in the receipt. Not roughly fine. Clean.

The climbing rule

How do you know when to climb? Not by the calendar, and not by how confident you feel. By receipts.

A clean run means you read the receipt, your spot-checks passed, and nothing in it surprised you. “It was basically fine” does not count. Clean counts.

Ten clean runs earns the next rung. Ten watched runs of the folder task before you close the laptop on it. Ten unattended runs before it goes on a schedule. Ten scheduled Mondays, reviewed and clean, before you so much as read about triggers.

And a surprise sends you back down a rung. Not as punishment. A surprise means your brief and your tool still have things to work out, and the rung below is where you can watch them work it out.

Say the scheduled Monday report one day lists a project nobody recognises. You do not shrug and delete the line. You run the same brief watched, find where the wrong assumption crept in, fix the brief, and re-earn the schedule.

Climbing down is not failure. Managers do the equivalent with people every week: new kinds of work get closer supervision until trust is earned again. You are simply managing.

Where this leaves your Monday

Picture three months from now. Monday, 8:05. The report is drafted. You spend five minutes reviewing it, correct one line, and send it. The forty minutes are yours again.

Notice who is still essential in that picture. The judgement, the correction, the decision that it is fit to send: that was always the actual job. The assembling never was.

One chapter left. It is about making all of this stick: the small weekly habit that separates people who tried an agent once from people whose Mondays genuinely look different.