What do I need to set up?
Compare equivalent ways of working. Start with normal chat unless your task genuinely needs an agent to work across files.
Normal ChatGPT or Claude chat
You need an account for each tool and one non-sensitive source file or block of text. Open a fresh conversation in both tools, paste the comparison brief below, and give each system the same material. This route tests answer quality. It does not test whether an agent can inspect files, run tools or correct a finished artifact.
ChatGPT Work or Codex
Use this route for a task with a finished deliverable, such as a page, spreadsheet or working file. ChatGPT Work can use approved files and tools to create work you review. Codex can work inside a project and verify changes. Keep the source files, constraints and finish line identical to the Claude route. If your plan or workspace does not offer these tools, use normal chat and score only the output you can actually test.
Claude Projects or Claude Code
Claude Projects can hold uploaded source material and project instructions on eligible paid plans. Claude Code is the closer comparison when the task requires work across project files. Keep project knowledge and permissions limited to this test. If you do not use either feature, create a fresh Claude chat and use the manual route.
The question that gives bad answers
"Which model is better?" sounds sensible, but it hides the work. A polished paragraph, a faithful translation and a tested website are different jobs. The winner can change when the task changes.
I write in English, Portuguese and French, and I also use AI to build real projects. That combination changed my own answer. ChatGPT has become my default, especially when I need to preserve meaning across languages or continue from an idea to a working artifact with Codex. That is personal evidence, not a universal benchmark.
Compare chat with chat, project with project, and coding agent with coding agent. Otherwise you may be measuring the surface instead of the model or workflow.
The five-part test
- Use a real task. Choose something you already need to finish, not a puzzle designed to make one tool look clever.
- Give both tools the same source material. Use the same file, notes, examples and facts.
- Ask for the same finished outcome. Define the artifact, audience, constraints and finish line before either tool starts.
- Allow one identical correction. This tests whether the system can understand and repair the miss instead of merely producing a strong first draft.
- Score the work that remains. Judge the finished usefulness, not the charm of the answer.
Copy this comparison brief
Replace the text in square brackets, then paste the same completed brief into both systems.
I am comparing two AI systems on one real task. Do not optimize for a clever first answer. Optimize for a finished result I can use and verify. FINISHED OUTCOME [Describe the exact artifact or decision you need.] AUDIENCE AND USE [Who will use it, where, and for what?] SOURCE MATERIAL [Attach or paste the same sources in both systems.] CONSTRAINTS - Use only the supplied sources for factual claims. - Preserve [meaning, tone, terminology, language, format, length, or brand rules]. - Do not invent missing facts. - State what you can and cannot do in this workspace. - Ask before using another source, tool, connector, or file. FINISH LINE [Explain what must be true before the task counts as complete.] Before starting, identify any missing information that would materially change the result. After producing the result, list: 1. What you verified 2. What remains uncertain 3. What still needs human judgment 4. What manual work remains I will then give both systems the same correction and score the revised result.
Score the revised result
Rate each category from 1 to 5 after the same correction. Add one plain-language note for anything that would matter in real use.
| What to score | Tool A | Tool B | What you are looking for |
|---|---|---|---|
| Accuracy and source discipline | /5 | /5 | Claims match the supplied evidence and uncertainty is visible. |
| Meaning and intent | /5 | /5 | The result preserves what you meant, not only the topic. |
| Natural voice or target language | /5 | /5 | The wording sounds right for the audience and language you actually use. |
| Quality after one correction | /5 | /5 | The system understands the miss and repairs it without breaking something else. |
| Finished usefulness | /5 | /5 | You can use the artifact with little cleanup. |
| Manual work left | Low / medium / high | Low / medium / high | Count the checking, formatting, copying, testing and repair still left to you. |
If you work in more than one language
Repeat the test in every language you genuinely use. Do not translate one tool's answer and compare it with the other tool's original. Give both systems the same source and ask for the same target-language outcome.
Score preservation of meaning separately from grammar. A sentence can be perfectly grammatical and still flatten your point of view.
If the task is a build
Do not stop at the first generated file. Add three checks to the finish line: inspect the relevant files, test the artifact, and correct what fails. Record whether the system completed those steps or merely told you what to do next.
For this build task, completion means more than generating a first draft. Before you finish: 1. Inspect the relevant source files. 2. Run the available validation or test. 3. Open or render the artifact when possible. 4. Correct the problems you find. 5. Tell me what you could not verify. Stop before publishing, deploying, sending, deleting, or changing access.
What a good result looks like
- You compared one real outcome rather than a vague prompt
- Both systems received equivalent context and constraints
- You scored the revised artifact, not only the first answer
- You can explain why one result fits this task better
- You kept the human decisions and factual checks visible
- You know what to retest when the products change
Three problems you may hit
One tool had much more context
Start again with the same source material in fresh workspaces. If one product has access to a connector or project knowledge the other does not, record that as part of the workflow advantage instead of pretending the model alone caused the difference.
The prettier answer keeps winning
Hide the tool names and score the work against your finish line. Then inspect the sources, unsupported claims and manual cleanup before revealing which tool made it.
The result changes every time
Run the task more than once when the decision matters. One test is personal evidence, not a permanent leaderboard. Record the date, product surface and plan you used.
Using this inside a company
A team comparison needs shared test cases, approved data, named reviewers and a definition of good that reflects the actual work. Keep price, permissions, security, integration and maintenance in the decision. Output quality is only one part of adoption.
Start with one contained workflow and a small group of real users. A company-wide model decision based on one demo will age badly.
Need a fair AI comparison for your team?
I help teams turn real work into test cases, evaluation criteria and adoption rules they can use after the demo is over.
Tell me what you need to compareSources and technical notes
Product capabilities, plans and workspace access change. Check the tools available in your own account before running the test. Technical details reviewed 2 August 2026.