Skip to main content
AI ComparisonAI comparisonDocument AnalysisChatGPT vs Claude vs Gemini

I Fed the Same 40-Page Contract to ChatGPT, Claude, and Gemini

One AI caught a liability clause the others missed entirely. Here's what actually happened when I tested all three on real legal fine print.

D
Davide
Β·Β·8 min

I Fed the Same 40-Page Contract to ChatGPT, Claude, and Gemini

Last week I ran an experiment that should worry anyone using AI to review legal documents. I took a real 40-page commercial lease agreement and fed it to ChatGPT-4, Claude 3.5 Sonnet, and Gemini 1.5 Pro β€” same document, same prompts, same conditions. Only one of them caught a liability clause buried on page 27 that could've cost a small business owner six figures if it went unnoticed.

The other two skimmed right past it. Here's exactly what happened, which AI actually deserves your trust with important documents, and the specific prompting technique that made all the difference.

The Clause That 2 Out of 3 AIs Missed Completely

The document was a standard-looking commercial lease with an indemnification clause hidden in a subsection about "tenant improvements." On the surface, it read like boilerplate legal language. It wasn't.

Buried in the middle of a 200-word paragraph was language that shifted unlimited liability for third-party injury claims onto the tenant β€” even for accidents caused by the landlord's own negligence. This is the kind of clause that sounds procedural until someone gets hurt and a lawyer points out you agreed to cover damages that weren't your fault.

I asked all three AIs the same question: "Review this lease and flag any clauses that create unusual risk or liability exposure for the tenant."

Claude 3.5 Sonnet caught it immediately. It didn't just flag the clause β€” it explained why it was dangerous, compared it to standard indemnification language, and suggested a specific redline: limiting liability to claims arising from the tenant's own actions.

ChatGPT-4 summarized the lease well overall but described that section as "standard indemnification language" β€” completely missing the unlimited liability shift. Gemini 1.5 Pro did something worse: it acknowledged the clause existed but categorized it as "low risk," which is arguably more dangerous than missing it entirely.

Why This Happened β€” And What It Reveals About How These AIs "Read"

Here's the part most comparisons skip: the AIs weren't reading the document the same way.

ChatGPT and Gemini both process long documents by chunking them into segments and summarizing each piece before stitching together an overall response. This works fine for structure and general summaries, but it means clauses that depend on context from other parts of the document can get evaluated in isolation.

Claude handled the lease differently. When I asked it to explain its reasoning, it referenced language from an earlier section (the personal guaranty clause) that changed how the liability clause should be interpreted. It connected dots across 15 pages instead of reading in isolated chunks.

This matters because legal documents are built on cross-references. A single word like "notwithstanding" on page 12 can completely change the meaning of a clause on page 30. If an AI treats each section as its own island, it will miss exactly the kind of nested risk that actually shows up in real contracts.

The mental model to take away: document length isn't the real challenge β€” document interconnection is. A 5-page contract with heavy cross-referencing can be harder for AI to parse correctly than a 40-page document that's mostly independent clauses.

How to Actually Use AI for Contract Review Starting Today

You don't need to abandon ChatGPT or Gemini β€” you need a better process. Here's the exact workflow I now use for any contract over 10 pages.

Step 1: Use Claude first for the initial risk scan. Upload the full document and prompt: "Act as a contract attorney. Identify every clause that creates financial risk, liability exposure, or unusual obligations for [my role β€” tenant/buyer/contractor]. For each one, explain the specific risk in plain English."

Step 2: Cross-check with ChatGPT using a different lens. ChatGPT is genuinely strong at comparing language against industry standards. Prompt it with: "Compare the indemnification and liability clauses in this document to standard commercial lease language. Flag anything that deviates from typical terms."

Step 3: Use Gemini for structural and financial summary only. Gemini's strength is organizing numbers, dates, and payment terms clearly β€” not nuanced risk analysis. Ask it: "Create a table of every deadline, payment obligation, and renewal date in this document."

Step 4: Never skip a human lawyer for anything with real financial stakes. This entire process took me 12 minutes and gave me a prioritized list of what to bring to an actual attorney β€” which cuts legal review time (and cost) significantly, but doesn't replace it.

The Part Most People Get Wrong

Most people treat AI contract review like a single yes/no filter β€” upload the document, ask "is this safe to sign," and trust whatever answer comes back. That's the mistake that almost cost the business owner in my example six figures.

AI models don't fail loudly. They fail by sounding confident while being wrong. Gemini didn't say "I'm not sure about this clause" β€” it said "low risk" with the same tone of certainty it used for everything else in the document.

The fix isn't finding the "best" AI and trusting it completely. It's using multiple models specifically because they fail differently, then treating any disagreement between them as a signal to look closer, not a coin flip to ignore.

Key Takeaways

  • Claude's advantage: It connects context across long documents instead of reading sections in isolation, making it stronger for cross-referenced legal language.
  • ChatGPT's advantage: It's better at comparing your document against standard industry language and norms.
  • Gemini's advantage: It excels at extracting structured data β€” dates, numbers, obligations β€” but is weaker at nuanced risk judgment.
  • The real risk factor: Document interconnection matters more than document length when it comes to AI accuracy.
  • The right process: Use multiple AI tools for different jobs, then treat disagreement between them as your cue to dig deeper β€” never rely on a single model's confidence.

What to Do Right Now

Open Claude.ai, upload any contract you're currently reviewing or about to sign, and run this exact prompt: "Act as a contract attorney. Identify every clause that creates financial risk or unusual obligations for me, and explain each risk in plain English." Do this in the next 10 minutes β€” before you sign anything else this month.

AI comparisonDocument AnalysisChatGPT vs Claude vs Gemini

Get Weekly AI Insights Delivered Free

Join 5,000+ subscribers getting the latest AI tool breakdowns, prompts, and strategies every week. No spam, ever.

  • ✦ Weekly AI tool reviews
  • ✦ Exclusive prompt packs
  • ✦ Early resource access
  • ✦ No spam, unsubscribe anytime

By subscribing, you agree to our Privacy Policy. Unsubscribe anytime.