Skip to main content
AI ComparisonChatGPTClaudeGeminidata analysisAI comparison

ChatGPT vs Claude vs Gemini: I Gave Them the Same Messy Spreadsheet

One spreadsheet, three AI models, wildly different results. Here's which one actually saved me hours of cleanup work.

D
Davide DeMango
ยทยท8 min

ChatGPT vs Claude vs Gemini: I Gave Them the Same Messy Spreadsheet

I took one genuinely disgusting spreadsheet โ€” 1,200 rows of customer orders with inconsistent date formats, duplicate entries, and typos in half the company names โ€” and fed it to ChatGPT, Claude, and Gemini with the exact same prompt. The results weren't close. One model saved me two hours of manual cleanup, one got the job half-right and lied about it, and one just refused to finish the task.

If you're using AI to clean data, analyze reports, or make sense of exports from your CRM, which model you pick actually matters. This isn't a "they're all basically the same" situation. Here's exactly what happened, why it happened, and which one you should be using for this kind of work.

The Test: 1,200 Rows, One Prompt, Three Very Different Outcomes

I used a real spreadsheet โ€” anonymized customer order data pulled from a Shopify export. It had the classic messy-data problems: dates written as "3/4/24," "March 4, 2024," and "2024-03-04" all in the same column. Company names like "Acme Corp," "ACME Corp.," and "acme corporation" that were clearly the same customer but wouldn't match in a simple filter.

I gave all three the same prompt: "Clean this spreadsheet. Standardize all dates to YYYY-MM-DD format, merge duplicate customer entries even if the names are slightly different, and flag any rows with missing data. Show me what you changed and why."

Claude handled it best. It processed the whole file, standardized every date correctly, and โ€” this is the part that impressed me โ€” it caught 34 duplicate customers that weren't exact matches, like "J. Smith" and "Jonathan Smith" with the same email domain. It gave me a summary table of every merge decision so I could double-check its logic.

ChatGPT (using GPT-4o) did a solid job on the dates but only caught the obvious duplicates โ€” exact or near-exact name matches. It missed about 15 of the same fuzzy-duplicate cases Claude caught, and when I asked why, it admitted it "didn't apply fuzzy matching by default."

Gemini started strong but choked on the file size. Past row 800 or so, it started skipping data silently โ€” no error message, it just stopped processing and returned incomplete results as if the job were done. That's the dangerous one, because you might not notice until you've already used the bad data.

Why This Happens: It's Not About "Which AI Is Smarter"

Here's what nobody explains clearly: these three models don't fail or succeed because one is generally "better." They fail differently because of how each one handles context and reasoning under ambiguity, and messy spreadsheets are basically a stress test for that.

Claude's advantage here comes from how Anthropic trained it to reason step-by-step through structured tasks before responding โ€” it's less likely to pattern-match a "good enough" answer and more likely to actually work through the logic of "is Acme Corp the same as ACME Corporation?" That's why it caught the fuzzy duplicates. It wasn't smarter in some abstract sense โ€” it was more thorough by design.

ChatGPT is excellent at fast, confident answers, which is exactly why it's good for writing, brainstorming, and quick analysis. But that same speed-oriented training means it's more likely to apply the obvious rule (exact match) and call it done, unless you explicitly force it to go deeper.

Gemini's silent failure is the scariest pattern of the three. It's not that Gemini is bad at data โ€” it's that when it hits a context or processing limit, it doesn't tell you. It just gives you a confident, incomplete answer that looks finished. If you're not spot-checking output, you'll never catch it.

The mental model to take from this: the model that "sounds" most confident isn't necessarily the one that did the most complete work. You have to test, not assume.

How to Actually Do This Yourself Today

Here's the exact workflow to replicate this with your own messy data, whether it's a CRM export, a client list, or expense reports.

Step 1: Upload the file directly, don't copy-paste. All three tools now accept file uploads (CSV, Excel). Copy-pasting messy data into the chat box loses formatting and increases errors โ€” always upload the actual file.

Step 2: Use a prompt that forces transparency, not just cleanup. Instead of "clean this data," use something like: "Clean and standardize this spreadsheet. For every change you make โ€” merged duplicates, corrected formats, flagged errors โ€” list it in a summary table so I can verify your work before I use it." This single addition is what exposes whether the model is doing shallow or deep work.

Step 3: Spot-check the middle and end of the file, not just the top. Gemini's failure only showed up past row 800. Always scroll to a random chunk in the middle of your output and the very last rows โ€” that's where silent failures hide.

Step 4: For anything over 500 rows with genuine messiness (typos, inconsistent formatting, fuzzy duplicates), start with Claude. For quick, straightforward cleanup on smaller, cleaner files, ChatGPT is faster and perfectly reliable. Save Gemini for tasks that don't require processing the entire file in one pass.

The Part Most People Get Wrong

Most people assume if an AI gives you an answer, it did the entire job. That's wrong, and it's the single most expensive mistake in AI-assisted data work.

When Gemini stopped processing at row 800, it didn't say "I stopped." It presented the output as complete. If you're not the kind of person who checks the row count of your output against your input, you'll ship broken data into a report, a client deliverable, or a decision โ€” and you won't know until someone downstream catches the error.

The fix isn't "don't trust AI with data." The fix is always ask for a verification step in your prompt, and always spot-check output against the original file count and structure. Treat every AI data-cleanup job the way you'd treat an intern's first draft โ€” good enough to save you time, not good enough to skip review entirely.

Key Takeaways

  • Claude wins on messy, ambiguous data: Its step-by-step reasoning caught fuzzy duplicates the other two missed.
  • ChatGPT is fast and reliable for straightforward cleanup: Great for smaller files without complex fuzzy-matching needs.
  • Gemini can fail silently on large files: It may stop processing without telling you โ€” always verify row counts.
  • The prompt matters as much as the model: Asking for a "summary of changes" forces deeper, more transparent work.
  • Never trust AI-cleaned data without spot-checking: Check the middle and end of the file, not just the top.

What to Do Right Now

Pull up your messiest spreadsheet โ€” the one you've been avoiding โ€” and upload it to Claude right now. Use this exact prompt: "Clean and standardize this spreadsheet. List every change you make in a summary table so I can verify your work." Check the last 20 rows before you trust any of it.

ChatGPTClaudeGeminidata analysisAI comparison

Get Weekly AI Insights Delivered Free

Join 5,000+ subscribers getting the latest AI tool breakdowns, prompts, and strategies every week. No spam, ever.

  • โœฆ Weekly AI tool reviews
  • โœฆ Exclusive prompt packs
  • โœฆ Early resource access
  • โœฆ No spam, unsubscribe anytime

By subscribing, you agree to our Privacy Policy. Unsubscribe anytime.