Saying “I use AI” is weak evidence at work.
Your manager has probably heard it from half the team. They may have seen clever prompts, polished demos and generated reports. What they still do not know is who can use the tool without creating a bigger checking problem for everybody else.
That is where you can become more useful.
To prove your value with AI, show one real work improvement that another person can inspect. Include the old baseline, the part AI assisted, the data and action boundaries, normal and failure tests, the full checking effort, what did not improve, and the judgement that remained human-owned. Evidence beats enthusiasm; a negative result still counts.
At Collab365, saying “AI helped build this” stopped being useful evidence very quickly. Helen does not need another screenshot of a model typing. She needs to know what changed, where she can see it, what was checked and what is still unproven.
That is why our useful record is usually boring: the exact file, the exact route, the tests, the remaining boundary and the human decision. It is much more credible than a perfect demo with no receipt.
What does “proving my value with AI” actually mean?
It does not mean proving that AI is brilliant. The vendor can do that demonstration.
It means showing that you can make a responsible decision about work:
- choose a problem worth solving;
- understand the task before changing it;
- give the tool enough authorised context;
- keep the risky decisions with a person;
- test the cases a demo avoids;
- measure the whole method;
- report weak results as well as good ones; and
- make the work repeatable without hiding the judgement in your head.
That is useful even if your organisation changes tools next year.
What evidence should I collect?
Think of the proof as a chain. If one link is missing, the story becomes much harder to trust.
The result does not have to be “keep”. A well-supported decision to change or stop the workflow is evidence of judgement too.
1. The baseline
Show how the task works before AI touches it.
Record the trigger, inputs, steps, people involved, hands-on time, waiting time, common corrections and approval. Use representative examples rather than a heroic best case.
Without a baseline, “faster” means “it felt faster”.
2. The assistance boundary
Write one sentence that says what the tool may do and what remains with a person.
For example:
“The tool drafts a weekly status update from the five approved team submissions. The project manager checks every date, number, risk and commitment against the originals and decides the final status before circulation.”
That sentence is more valuable than a long system prompt because it identifies ownership.
3. The tests
Use ordinary, awkward and failure cases. Include missing information, conflicting sources and a case where the correct behaviour is to stop and ask.
Keep the source, output and correction record together so somebody else can inspect what happened.
4. The results
Count the whole task:
New total work = preparation + tool time + checking + correction + approval + failed runs.
Compare that with the old total. Also record quality, omissions, risk, user confidence and work moved to somebody else.
If drafting got quicker but checking became harder, say so.
5. The repeatable method
Give another capable person the written pack without coaching them through it. Ask them to attempt the task and note every place they have to come back to you.
Those questions are useful evidence. Before a wider rollout, turn them into the scope, ownership and stop rules that prevent the creator becoming the AI helpdesk.
That is the cover test. Every return question is missing context, a vague rule or judgement that has not travelled yet.
What does not prove value?
Theatre is easy to create because generative tools are very good at producing something that looks finished.
A polished output can be part of the evidence. It cannot replace the chain behind it.
Be cautious of these claims:
- “It saved me hours.” Compared with what, across how many cases, including review?
- “The answer was perfect.” According to which source and acceptance criteria?
- “I built an agent.” What can it access or change, and who approves the action?
- “Everybody can use my prompt.” Do they have the same context, permissions and judgement?
- “We automated the process.” Does it run reliably, recover from failure and stop before a consequential action?
- “The team loved it.” What did people actually do, and what support burden appeared later?
This is not cynicism. It is how you stop a useful experiment being dismissed as another AI demo.
What is a good first AI project at work?
Choose a recurring task with a visible input and output, modest consequences and a result you can check against an authoritative source.
Good candidates might include:
- a first draft of a weekly internal update;
- extracting actions from approved meeting notes for the meeting owner to correct;
- checking a document against a published internal standard;
- classifying routine requests for a person to review;
- turning approved source material into a standard internal summary; or
- generating test cases for a process that a subject expert then runs.
Avoid an impressive but unrepeatable one-off. Also avoid a task where a mistake could affect employment, health, finance, legal rights, safeguarding or physical safety without stronger governance and specialist oversight.
Microsoft's guidance on choosing when Copilot or an agent is the right tool uses four practical checks: repeatability, impact, error detectability and time sensitivity. It also says delegation does not transfer accountability.
Those checks are a sensible minimum whatever tool you use.
A worked example: the weekly project update
This is a worked example, not a customer result. The numbers are deliberately blank because you must measure them in your own setting.
Suppose a project manager receives five team updates every Thursday and produces a report on Friday.
Define the old method
| Field | What to record |
|---|---|
| Trigger | Five named team updates received by Thursday afternoon |
| Output | Approved Friday status report |
| Sources | The five submissions, plan, risk log and decision log |
| Hands-on time | Measure it across several real weeks |
| Common rework | Missing dates, inconsistent status, unsupported claims, duplicated actions |
| Approver | Name the project owner who decides the final status |
Set the AI boundary
The tool may:
- assemble a first draft using the approved source pack;
- flag missing fields and contradictions;
- use the agreed report structure; and
- cite the source update beside each factual claim.
The tool may not:
- invent progress for a missing update;
- decide whether a risk is red, amber or green;
- change an owner or deadline;
- send the report; or
- use material outside the approved pack.
Build the test set
Use at least these cases:
- All five updates are complete and consistent.
- One update is missing.
- Two updates give different milestone dates.
- A risk is described politely but is clearly getting worse.
- One source contains a note that must not appear in the wider report.
Record the result without polishing it
For each case, record:
- whether the draft followed the structure;
- unsupported statements;
- missed contradictions;
- source-citation errors;
- information that should have been withheld;
- hands-on checking and correction time; and
- whether the named approver would accept it.
Only after that do you decide whether the workflow is Keep, Change, Stop or Not tested yet.
Notice what the example demonstrates. The valuable part is not typing the prompt. It is knowing which sources count, where the decision sits, which failure cases matter and when the report must stop.
That wider operating route is what distinguishes a dependable AI workflow from one clever prompt.
What can I show my manager?
Give them one page, not a thirty-slide AI strategy.
Use this structure:
The problem
One sentence describing the recurring work and why it matters.
The old method
The baseline, including checking and approval.
The proposed assistance
What AI does, the approved tool and data, and the human-owned decision.
The evidence
Which cases you tested, what passed, what failed and the measured result.
The boundary
What the method must never do and when it must ask a person.
The decision
Keep, Change, Stop or Not tested yet.
The ask
What you need from the manager: approval, a named owner, access to safe test material, time for a second test or agreement to stop.
The page should make sense to somebody who never saw the demo.
How should I talk about the result without overclaiming?
Try this:
“I found one recurring task where AI might help. I measured the existing method, set a boundary, tested representative and failure cases, and recorded the checking effort. The current result is [status]. Here is what improved, what did not, and the part that still needs a named person. I would like your agreement on the next test and the owner.”
That language is calm because the evidence is carrying the weight.
Do not promise a promotion. Ask your manager what evidence the next role actually requires. Your experiment is relevant only if it demonstrates part of that requirement.
For one role, the useful evidence may be process improvement. For another, it may be risk judgement, stakeholder communication, coaching, commercial impact or service quality. Local expectations matter more than a generic list of “future skills”.
How do I stop the workflow depending on me?
A saved prompt is not a handover.
Another person also needs:
- the task trigger and finish line;
- approved inputs and sources;
- an example of a satisfactory result;
- privacy, permission and action boundaries;
- quality checks against original material;
- stop-and-ask rules;
- failure examples;
- the named approval step; and
- the current Keep, Change, Stop or Not tested yet status.
Run the cover test. If the task comes back, do not blame the colleague. Find the missing context.
Can this make my career safe or guarantee promotion?
No.
The ILO's task-level exposure work says transformation is more likely than full replacement across the labour market, but it does not forecast your employer or personal outcome. Futureproof makes the same boundary explicit: its scores describe task exposure, not the probability that a person loses a job.
A completed experiment proves something narrower: you can inspect one piece of work, set a responsible boundary, test it and report the result honestly.
That may be valuable evidence in a promotion or development conversation. The decision still belongs to your employer.
Build one visible example
If you already have one suitable recurring task and want a guided route through the task map, trust checks, cover test and next-shape decision, use the Build My First Repeatable AI Workflow Board.
It is a paid Board for non-technical professionals, founders, consultants and team leads who already use an approved AI tool but cannot yet hand one workflow to somebody else with confidence. It helps you document and test one example. It does not do the work for you, certify the result, guarantee a saving or make a career safe.
Do not start there if you still cannot name the task. Look up which parts of your work may be changing first, or use the 30-day plan to choose safely.
The aim is not to look clever with AI. It is to leave somebody else able to see what you improved, what you protected and what you were wise enough not to automate.
Sources and attribution
- Microsoft Support: Decide when Copilot or an agent is the right tool for your work.
- Microsoft Learn: Data, privacy and security for Microsoft 365 Copilot.
- NIST AI Risk Management Framework and NIST AI 600-1: Generative Artificial Intelligence Profile.
- International Labour Organization: Generative AI and Jobs, a Refined Global Index of Occupational Exposure, ILO Working Paper 140, 2025.
- Collab365 Futureproof method and limitations and dated release 2026-q4.1, licensed CC BY 4.0 with full upstream attribution in the release.
