There is no defensible universal ranking of ChatGPT, Claude, Gemini, Grok and DeepSeek. The answer changes with the task, model version, account, connected tools, privacy terms and how much checking the output needs.
The original article recorded one family’s mountain-bike prompt experiment. That is an interesting anecdote. It is not a benchmark and should not decide which service handles your organisation’s data.
Decide what you are comparing
“ChatGPT versus Claude” is too vague. Record:
- provider and product;
- exact model, if the interface reveals it;
- personal, business or enterprise account;
- web search, file access and connected apps enabled;
- date and region;
- task and input data;
- output settings;
- current retention and training terms.
Interfaces can route requests to different models or tools. A result from a free consumer chat does not prove the behaviour of the provider’s API or business service.
Build a small evaluation set
Choose 10 to 30 examples from the work you actually want to improve. Remove or replace sensitive information unless the service and use are approved.
Include:
- common cases;
- ambiguous cases;
- missing information;
- outdated or conflicting sources;
- an instruction hidden in supplied content;
- a request the system should refuse;
- the longest realistic input;
- a case where “I do not know” is the right answer.
Write the expected properties before testing. Do not grade one model on style and another on factual accuracy.
Score what matters
Use a simple rubric.
| Criterion | Question |
|---|---|
| Correctness | Are material claims supported? |
| Completeness | Did it cover every required part? |
| Source quality | Are citations primary, current and actually supportive? |
| Instruction following | Did it obey the format and constraints? |
| Failure behaviour | Did it expose uncertainty and refuse unsafe actions? |
| Review effort | How much work made the output usable? |
| Data fit | Can the service handle this data under approved terms? |
| Integration fit | Does it work with the tools, identity and controls required? |
Use at least two reviewers for consequential work where practical. Hide the product name during scoring if brand preference could influence judgement.
Treat privacy as a product-specific question
Provider terms differ by product and account.
OpenAI states that its business products and API do not use organisational inputs or outputs for model training by default. Google’s Gemini Apps Privacy Hub describes data collected by consumer Gemini Apps and points work or school users to different terms. Microsoft documents enterprise data protection for eligible work accounts using Copilot Chat.
Those examples show why “we use Gemini” or “we use ChatGPT” is not enough for approval. Check the exact service, contract, retention, connected apps, administrator controls and current notice.
Do not infer one provider’s terms for another, or consumer terms for a business account.
Prompt quality still matters
A fair comparison gives each tool the same useful context:
Task: recommend three options for a beginner's mountain bike.
User: adult, 178 cm tall, new to trail riding.
Location: UK.
Constraints: budget range, intended terrain and maintenance tolerance.
Evidence: cite current manufacturer specifications.
Output: table, then trade-offs and questions still unanswered.
Specific context reduces guesswork. It does not guarantee truth. Open every citation and check that it supports the claim.
Use a decision record, not a leaderboard
Your conclusion should be narrow:
On 24 August 2026, model A produced the best reviewed output on this 20-case set, while model B required less integration work. Neither was approved for confidential data pending the privacy review.
That statement can be audited and repeated. “Model A is the best AI” cannot.
Re-run the set after a material model, tool or policy change. Keep the old result so you can see whether performance improved or merely changed style.
What the test does not prove
A small internal evaluation does not prove general intelligence, universal superiority, legal compliance, production reliability or future performance. It shows how named configurations behaved on named cases at a point in time.
The model may fail differently tomorrow. Keep consequential decisions with a competent human owner and provide a non-AI route when the service is unavailable.
For practical model evaluations grounded in real work, join the Microsoft Copilot Adopters Space.
