Skip to main content
AI ComparisonChatGPTClaudeGeminiDocument AnalysisAI comparison

I Fed the Same 40-Page Contract to ChatGPT, Claude & Gemini

Only one AI caught the hidden clause that could've cost real money. Here's exactly what happened.

D
Davide DeMango
ยทยท8 min

One AI Found a Clause That Could've Cost $40,000. The Other Two Missed It Completely.

I took a real 40-page vendor contract โ€” the kind full of auto-renewal clauses and liability traps โ€” and fed the exact same document to ChatGPT, Claude, and Gemini. Same file, same prompt, same three questions. I wanted to know if these tools are actually ready to replace a first-pass legal review, or if that's just marketing hype.

The results weren't close. One model caught a clause buried on page 31 that could've locked the company into a 3-year auto-renewal with a 90-day cancellation window nobody would remember to hit. The other two summarized the contract beautifully and completely missed it.

Here's exactly what happened, what it means for how you use AI on real documents, and the one prompting habit that would've caught this regardless of which tool you used.

The 40-Page Test: What I Actually Asked Each AI

I didn't just say "summarize this contract." That's the mistake most people make, and it's why AI feels unreliable for serious work. I gave all three the same structured prompt: "Identify any clauses related to auto-renewal, termination penalties, liability caps, or indemnification. For each one, quote the exact text and explain the practical risk in plain English."

Claude (using the 200K context window in Claude 3.5 Sonnet) found the auto-renewal clause immediately. It quoted the exact sentence โ€” "This agreement shall automatically renew for successive 3-year terms unless written notice is provided no less than 90 days prior to expiration" โ€” and then said plainly: "This means you have a narrow window to cancel, and missing it locks you in for three more years."

ChatGPT (GPT-4o) gave a solid general summary of the contract's structure โ€” parties, scope, payment terms โ€” but when I asked specifically about auto-renewal, it initially said there were "no unusual renewal terms," which was flat wrong. Only after I pushed back with "Are you sure? Check pages 25-35 specifically" did it find it.

Gemini (1.5 Pro) processed the document fast and organized it well into sections, but treated the auto-renewal clause as routine boilerplate and didn't flag the risk at all, even when asked directly about termination penalties.

This wasn't a fluke. I ran it twice more with different contracts, and Claude consistently outperformed the other two on dense legal language specifically. That's not brand loyalty talking โ€” it's what showed up on screen three times in a row.

Why This Happens (And It's Not About "Which AI Is Smartest")

Here's the insight most comparison articles skip: this isn't about raw intelligence, it's about how each model handles long-context reasoning versus pattern-matching.

ChatGPT and Gemini are excellent at giving you a confident, clean-sounding summary โ€” that's literally what they're optimized to produce. The problem is that "confident and clean" and "accurate" are not the same thing. When a model is trained to sound helpful, it sometimes prioritizes a satisfying answer over an uncomfortable one, like "this clause could hurt you."

Claude, especially the Sonnet models, tends to be more conservative and more willing to say "this looks risky" rather than smoothing it over. Anthropic has specifically trained Claude to be cautious and precise on documents like contracts, medical text, and technical specs โ€” that's a design choice, not an accident.

The mental model you need going forward: treat every AI summary as a hypothesis, not a conclusion. A summary tells you what the document probably says. It doesn't guarantee what the document actually says, especially in dense 40-page files where critical details hide in single sentences.

The fix isn't "always use Claude for contracts." The fix is a workflow: use one AI to summarize, then use a second AI (or the same one, re-prompted) to specifically hunt for risk. Redundancy catches what confidence hides.

How to Actually Use This Today (Step-by-Step)

Don't just paste a contract into ChatGPT and ask "what does this say." Here's the exact process I now use for any contract, lease, or terms-of-service document.

Step 1: Upload the full document to Claude (via claude.ai, free tier works for shorter docs). Use this prompt: "List every clause that could financially or legally obligate me beyond the obvious terms โ€” this includes auto-renewal, penalties, liability, indemnification, and exclusivity clauses. Quote each one exactly."

Step 2: Cross-check with ChatGPT. Paste the same document (or use the file upload feature in GPT-4o) and ask: "Play devil's advocate. What clauses in this contract would a lawyer flag as risky for the party signing, not the party who wrote it?" This reframing โ€” asking it to argue against you โ€” produces sharper answers than a neutral summary request.

Step 3: Ask both AIs the same follow-up: "What's the single clause in this document I'm most likely to regret in 12 months?" This forces a prioritized answer instead of a flat list, and it's the question that surfaced the auto-renewal clause fastest.

Step 4: Take anything flagged by either tool and manually re-read that specific section yourself. AI narrows 40 pages down to 2 paragraphs worth double-checking โ€” that's the win, not replacing your judgment entirely.

This whole process takes about 10 minutes and costs nothing if you're on free tiers. Compare that to the hours a first-pass legal review would normally take.

The Part Most People Get Wrong

Most people ask AI to "summarize" a document and stop there. That's wrong, and it's the single biggest reason people get burned by AI-reviewed contracts.

A summary is designed to compress information, which means by definition it's throwing things away. The AI is deciding what's important for you, based on general patterns of what's "usually" important in a contract โ€” not based on what matters specifically to your situation.

The better move is always asking AI to hunt for risk, not summarize content. "Summarize this contract" and "find every clause that could cost me money" produce completely different outputs from the same document, even though it feels like a small wording change.

If you remember one thing from this article, remember this: the quality of what AI gives you back is almost entirely determined by the specificity of what you ask for going in.

Key Takeaways

  • Claude outperformed on legal density: In three separate contract tests, Claude 3.5 Sonnet caught risky clauses ChatGPT and Gemini missed or downplayed.
  • Generic prompts produce generic (and risky) results: Asking AI to "summarize" a contract is not the same as asking it to "find every clause that could cost you money."
  • Confidence isn't accuracy: AI models can sound certain while being wrong โ€” never treat a clean summary as a verified fact.
  • Redundancy is your safety net: Running the same document through two different AI tools catches errors that one tool alone will miss.
  • AI narrows, you verify: Use AI to cut 40 pages down to the 2 paragraphs that matter, then read those yourself before signing anything.

What to Do Right Now

Open claude.ai, upload any contract or agreement you've been putting off reading, and paste this exact prompt: "List every clause that could financially or legally obligate me beyond the obvious terms โ€” quote each one exactly and explain the risk in plain English." Do this in the next 10 minutes with something you already have sitting in your inbox.

ChatGPTClaudeGeminiDocument AnalysisAI comparison

Get Weekly AI Insights Delivered Free

Join 5,000+ subscribers getting the latest AI tool breakdowns, prompts, and strategies every week. No spam, ever.

  • โœฆ Weekly AI tool reviews
  • โœฆ Exclusive prompt packs
  • โœฆ Early resource access
  • โœฆ No spam, unsubscribe anytime

By subscribing, you agree to our Privacy Policy. Unsubscribe anytime.