Skip to main content
AI ComparisonChatGPTClaudeGeminidata analysisAI comparison

ChatGPT vs Claude vs Gemini: I Tested Them on a Messy Excel File

I fed the same chaotic spreadsheet to all three AI models. Only one didn't hallucinate the numbers.

D
Davide DeMango
ยทยท7 min

ChatGPT vs Claude vs Gemini: I Tested Them on a Messy Excel File

I fed the same chaotic spreadsheet to ChatGPT, Claude, and Gemini โ€” 847 rows of inconsistent dates, merged cells, duplicate entries, and formulas that referenced sheets that didn't exist anymore. Only one model gave me numbers I could actually trust. The other two confidently made things up, and if I hadn't double-checked, I would've sent wrong data to a client.

This matters because millions of people are now using AI to "clean up" spreadsheets without verifying the output. You're about to see exactly which model hallucinates under pressure, why it happens, and how to catch it before it costs you. Let's get into what actually happened.

The Test: One Broken Spreadsheet, Three AI Models

I used a real file โ€” a small business's sales tracker with 14 months of data, three renamed columns, a few #REF! errors, and rows where someone had manually typed "N/A" instead of leaving cells blank. This is the kind of file every small business owner actually has sitting in their Google Drive right now.

I gave all three models the identical prompt: "Clean this data, calculate total revenue by month, and flag any inconsistencies you find." Same file, same instructions, same starting conditions.

ChatGPT (GPT-4o) cleaned the formatting beautifully and gave me a polished monthly breakdown. The problem: it silently averaged two months that had missing data instead of flagging them, which inflated my totals by roughly $8,200. It looked confident. It was wrong.

Gemini 1.5 Pro handled the file upload smoothly and caught more of the formatting issues than ChatGPT did. But when it hit the #REF! errors, it invented plausible-looking numbers to "fill the gaps" instead of telling me the data was missing. That's the scariest kind of mistake โ€” the output looks completely normal.

Claude 3.5 Sonnet was the only one that stopped and said: "Rows 112, 340, and 601 contain missing or invalid values. I did not include these in the revenue total โ€” here's what changed and why." It gave me a smaller, less impressive-looking number. It was also the only correct one.

Why This Happens: The "Confidence Trap" in AI Data Work

Here's the thing nobody tells you: AI models are trained to sound helpful, and "helpful" often gets confused with "complete." When a model hits a gap in your data, it has two choices โ€” flag the gap and look less polished, or fill the gap and look impressive. Most models default to filling.

This is called hallucination under ambiguity, and it gets worse the more "structured" your task looks. Spreadsheets feel like math to us, so we assume the AI is doing math. But the model isn't running your formulas โ€” it's predicting what a completed spreadsheet should look like based on patterns. That's a completely different process, and it's why the errors are so easy to miss.

The mental model that changes everything: treat AI spreadsheet work like a junior analyst's first draft, not a calculator. A calculator can't lie to you. A junior analyst absolutely can, especially if they think giving you a clean answer matters more than telling you they're stuck.

Claude performed better here specifically because Anthropic has trained it to be more conservative when data is incomplete โ€” it's built into how the model is tuned, not a lucky guess. That's not a permanent advantage (models update constantly), but as of right now, it's a measurable behavioral difference you can test yourself in five minutes.

The workflow that catches this every time: after any AI cleans your data, ask "What assumptions did you make, and what data did you have to estimate or skip?" If the model can't answer that clearly, don't trust the output.

How to Actually Use AI on Your Messy Spreadsheets Today

Open whichever tool you already pay for โ€” ChatGPT, Claude, or Gemini all work for this. Upload your spreadsheet directly rather than pasting data as text; all three handle .xlsx and .csv files now, and direct upload preserves structure better.

Use this exact prompt structure instead of a vague "clean this up" request: "Review this spreadsheet for missing values, duplicate rows, and formatting inconsistencies. Do not calculate any totals yet โ€” first, give me a list of every problem you found and where it is." This forces the model to show its work before it starts making decisions for you.

Once you have that list, verify it manually against 5-10 rows yourself. This takes two minutes and it's the step everyone skips. If the AI's list matches what you spot-check, move to step three.

Now ask for the calculations, but add this line every single time: "Flag any row where you had to estimate, assume, or skip data instead of using an exact value." This one sentence is the difference between catching an $8,200 error and sending it to your client.

If you're using Claude, you can push further with "Walk me through your calculation for [specific month] step by step before giving me the final number." Making the model show its reasoning path exposes hallucinations almost immediately, because errors that look fine in a final answer often fall apart when explained step by step.

The Part Most People Get Wrong

Most people assume a polished-looking output means an accurate one. That's wrong, and it's the single biggest reason AI mistakes make it into real business decisions. A clean table with perfectly formatted numbers feels trustworthy โ€” but formatting and accuracy are two completely unrelated things.

The second mistake: people treat all three AI models as interchangeable for data work. They're not. ChatGPT and Gemini are excellent for brainstorming, writing, and general research, but this test shows a real, current gap in how conservative each model is with ambiguous numerical data.

The fix isn't "always use Claude." Tools update every few months and today's gap could close tomorrow. The fix is building the habit of asking every model "what did you assume?" โ€” every time, regardless of which one you're using.

Key Takeaways

  • Confidence isn't accuracy: A clean, polished AI output can still contain silently fabricated numbers.
  • Claude flagged errors, the others filled them: In this specific test, Claude 3.5 Sonnet was the only model that refused to guess missing data.
  • Ask before you calculate: Get the model to list problems in your data before requesting any totals or math.
  • Force transparency with one sentence: "Flag any row where you had to estimate or assume data" catches most hallucinations instantly.
  • Spot-check manually every time: Two minutes of manual verification is the cheapest insurance against a costly AI mistake.

What to Do Right Now

Open your own messiest spreadsheet โ€” the one you've been avoiding โ€” and upload it to Claude, ChatGPT, or Gemini right now. Use the prompt: "Review this spreadsheet for missing values, duplicate rows, and inconsistencies. Do not calculate anything yet โ€” just list every problem you find and where it is." Compare what it finds against five rows you check yourself, and you'll immediately see how much you can trust it.

ChatGPTClaudeGeminidata analysisAI comparison

Get Weekly AI Insights Delivered Free

Join 5,000+ subscribers getting the latest AI tool breakdowns, prompts, and strategies every week. No spam, ever.

  • โœฆ Weekly AI tool reviews
  • โœฆ Exclusive prompt packs
  • โœฆ Early resource access
  • โœฆ No spam, unsubscribe anytime

By subscribing, you agree to our Privacy Policy. Unsubscribe anytime.