I Fed the Same 50-Page Contract to ChatGPT, Claude, and Gemini
Only one of them caught the hidden liability clause. The other two either missed it entirely or buried it in vague summary language that made it easy to skim past. I ran this test because everyone keeps asking me "which AI is actually best," and honestly, the answer depends entirely on what you're using it for. So I took a real 50-page vendor contract, fed it to all three models with the exact same prompt, and compared what came back line by line. What I found changes how I'll use these tools for anything legal, financial, or high-stakes going forward โ and it should change how you do too.
The Test: One Contract, One Prompt, Three Very Different Results
Here's exactly what I did. I took a 50-page vendor services agreement (real contract, redacted company names) and uploaded it to ChatGPT-4, Claude 3.5 Sonnet, and Gemini 1.5 Pro. Same file, same prompt, same time of day: "Review this contract and flag any clauses that could expose us to unexpected liability, unfavorable terms, or risk. List them with page numbers."
Gemini gave me a clean, organized summary in about 15 seconds. It correctly identified the payment terms and termination clause issues, but it completely missed a clause on page 34 that shifted liability for third-party data breaches onto my side โ a clause buried in dense legal language under a section titled "General Provisions" (classic move by whoever drafted this).
ChatGPT caught the general shape of the risk but described it vaguely: "There may be some liability concerns around data handling responsibilities." That's technically true, but it's not useful. If you're a business owner reading that, you have no idea how bad it actually is or what to do next.
Claude was the only one that nailed it. It quoted the exact sentence, explained in plain English what it meant ("you would be responsible for costs even if the breach originated from the vendor's system"), and flagged it as high-priority risk. It also caught two smaller issues the other two missed entirely โ an auto-renewal clause with a narrow 15-day opt-out window, and a non-compete that was broader than industry standard.
This isn't a one-off. I ran the same test on two other contracts afterward, and the pattern held: Claude consistently pulled out specific, quotable risk language while the other two gave broader, safer-sounding summaries that looked complete but weren't.
Why This Happens: The Difference Between Summarizing and Actually Reading
Here's the insight most comparison articles miss: these models aren't all doing the same task, even when you give them the same prompt. ChatGPT and Gemini tend to summarize โ they compress the document into digestible chunks and tell you the general themes. Claude tends to analyze โ it treats the request more like a legal review than a summary task, which means it's more likely to flag specifics instead of generalities.
This comes down to how each model handles long-context reasoning. Claude was built with heavy emphasis on document analysis and has consistently scored higher on tasks requiring you to find a "needle in a haystack" buried in a long text. Anthropic has published benchmarks showing Claude's accuracy on retrieval tasks across 100K+ token documents outperforms competitors specifically when the information is subtle or contradicts the document's overall tone.
That's exactly what happened here. The liability clause wasn't hidden โ it was in plain sight on page 34. But it was written to sound like boilerplate, buried under a boring section header, and phrased in a way that a fast summarizer would skim right past. Claude didn't skim. It read every clause as if it mattered, because you told it to look for risk, and it took that instruction literally rather than generally.
The mental model to take from this: when you ask an AI to "review" something, you're not asking it to compress information โ you're asking it to interrogate it. Most models default to compression because that's what most users want most of the time. You have to explicitly signal that this is a scrutiny task, not a summary task, or you'll get the safer, shallower version even from a capable model.
This also explains why the same tool can feel brilliant on one task and mediocre on another. It's not that Gemini is "worse" โ it's fantastic for quick synthesis of long documents when you just need the gist. It's the wrong tool when you need someone to catch the one sentence that could cost you $50,000.
How to Actually Use This Today
Next time you need an AI to review a contract, agreement, or any high-stakes document, do this instead of just uploading and asking "summarize this":
Step 1: Use Claude for the first pass. Upload the document and use a scrutiny-specific prompt: "Act as a contract lawyer reviewing this on my behalf. Identify every clause that creates liability, financial risk, or unfavorable obligations for my side. Quote the exact language and cite the page number for each one." The word "quote" matters โ it forces the model to pull real text instead of paraphrasing, which is where details get lost.
Step 2: Cross-check with ChatGPT for plain-English translation. Once Claude flags the risky clauses, paste those specific sections into ChatGPT and ask: "Explain this clause like I'm not a lawyer. What's the worst-case scenario if this plays out badly for me?" ChatGPT is genuinely excellent at translating dense legal or technical language into something a normal person understands.
Step 3: Use Gemini for the executive summary you'll actually send to your team. Once you know what the real risks are, Gemini is great at turning that into a clean, organized document โ bullet points, priority levels, suggested next steps. This is where its strength in structure and formatting actually shines.
Step 4: Never treat any AI output as final. Run the same prompt twice, on different days if possible, and see if the answers stay consistent. If a model gives you different risk assessments on the same document, that's a signal to get human eyes on it โ a real lawyer for anything with real money attached.
This three-tool workflow takes maybe 20 extra minutes compared to just uploading a PDF and reading whatever comes back. Given that the clause Claude caught could have meant my business eating costs for a vendor's data breach, 20 minutes is nothing.
The Part Most People Get Wrong
Most people pick one AI tool and stick with it for everything โ legal review, writing, coding, research โ because switching feels like extra work. That's the mistake. These models have real, measurable differences in what they're good at, and treating them as interchangeable means you're leaving risk on the table without knowing it.
The other common error: assuming a confident-sounding summary means a complete summary. All three models in my test sounded confident. Gemini's output looked polished and thorough โ it just wasn't. Confidence in tone has nothing to do with accuracy in content, and that gap is exactly where expensive mistakes hide.
The fix isn't "use Claude for everything now." It's understanding that the prompt you write matters as much as the model you choose. A generic "summarize this" prompt will get you a generic result from any of these tools. A specific, adversarial prompt that tells the AI exactly what kind of scrutiny you want will get you dramatically better results โ even from the "weaker" model on this particular task.
Key Takeaways
- Claude wins for document scrutiny: In this test, it was the only model that caught a buried, high-cost liability clause the other two missed.
- Summarizing and analyzing are different tasks: Most AI models default to compression unless you explicitly ask for interrogation-level review.
- Your prompt determines your risk exposure: A vague "review this" prompt gets vague results โ specificity forces the model to dig deeper.
- Cross-checking beats single-tool trust: Using Claude to flag risk, ChatGPT to translate it, and Gemini to organize it produces a far more reliable result than any single pass.
- Confidence isn't accuracy: Every model in this test sounded sure of itself โ only one was actually thorough.
What to Do Right Now
Open Claude right now, upload any contract or agreement you've been meaning to review, and use this exact prompt: "Act as a contract lawyer reviewing this on my behalf. Identify every clause that creates liability, financial risk, or unfavorable obligations for my side. Quote the exact language and cite the page number for each one." Compare what it flags against what you remember reading โ you'll be surprised what you missed the first time.