Behind the Scenes

AI Made the Content Look Finished. Our Allotment Test Said Otherwise.

How an allotment experiment exposed the limits of AI scoring rubrics, and the checks Collab365 now uses to distinguish polished output from useful guidance.

AI can produce content that is clear, complete and completely wrong for the person who needs it.

We found that by asking our content system to help with an allotment problem.

Helen and I had added scoring gates to the AI-assisted workflow inside Collab365 Spaces. A draft could be checked for clarity, accuracy, completeness, relevance and safety, then returned for repair.

It looked sensible on paper.

Then we tried it outside our familiar Microsoft 365 territory.

The generic answer passed

Our question was not simply "How do you grow vegetables?"

We both work. We cannot be at the plot every day. We wanted low-maintenance choices, manageable watering, less weeding and a plan that did not turn the allotment into another full-time job.

The first material looked polished. It covered planting, crop rotation, watering and pests.

It could score well against a generic gardening rubric while still failing our Tuesday afternoon reality.

That exposed a weakness in the scoring system. A rubric can reward the qualities it names and miss the ones it does not.

"Good" needs an observable test

We now push past vague scores with questions such as:

  1. Who will use this?
  2. What are they trying to do?
  3. What constraints shape the job?
  4. What source proves the factual advice?
  5. What would the person produce, decide or change?
  6. How could they tell whether it worked?
  7. What could go wrong if the advice is followed?

For the allotment, a useful output might be a weekly plan that survives missed visits and dry weather. For a Power Automate article, it might be a flow that handles a failed action without losing data.

The test must match the work.

A model should not grade its own confidence into truth

AI critique is useful for finding duplication, missing sections and unclear instructions. It is weaker when the same model writes the answer, invents the standard and awards itself a high score.

Our safer pattern separates jobs:

  • a writer creates or repairs the draft;
  • current primary sources support factual claims;
  • a separate verifier checks the draft against explicit criteria;
  • a human reviews the final evidence and practical fit;
  • unresolved uncertainty remains visible.

We also cap revision loops. Repeated automated polishing can increase cost and make prose smoother without making the advice more true.

The allotment experiment did not prove that our system produces perfect content. It did the opposite. It showed us how easily a tidy process can approve the wrong thing.

That was a successful test because it failed before a reader had to discover the weakness for us.

If you need an evidence-led way to test AI output before it reaches customers, join The AI Authority.

Sources

  • This article is a first-person account. Personal experience is not presented as independently audited external evidence.