I Fed the Same 40-Page Contract to ChatGPT, Claude & Gemini
Same PDF. Same prompts. Three completely different sets of red flags.
I ran a real 40-page commercial lease agreement through ChatGPT (GPT-4), Claude (3.5 Sonnet), and Gemini (1.5 Pro) to see which one actually catches the clauses that could cost you money. The results weren't close โ one AI missed a liability clause that could've meant six figures in unexpected costs. If you're using AI to review contracts, NDAs, or terms of service, you need to know which tool to trust and, more importantly, which one to double-check. Here's exactly what happened when I put all three head-to-head.
Claude Caught 3 Clauses the Others Missed Completely
Claude found an indemnification clause buried on page 27 that shifted liability entirely onto the tenant in case of third-party lawsuits. Neither ChatGPT nor Gemini flagged it as high-risk.
I used the same prompt across all three: "Review this contract and identify any clauses that create unusual risk or liability for the party signing. Rank them by severity." Claude returned a structured list with the indemnification clause at the top, explained in plain English why it mattered, and even suggested specific language to push back with.
ChatGPT mentioned the clause existed but categorized it as "standard legal language" โ which is technically true in some contracts, but not in this one, where the wording was unusually broad. That's a dangerous miss if you're not a lawyer and you're trusting the AI's judgment.
Gemini didn't flag it at all. It focused heavily on payment terms and renewal dates โ useful, but it completely skipped the liability language that actually mattered most.
The takeaway here isn't "Claude is always better." It's that long-context reasoning on legal language is where these models diverge hardest. Claude's training seems to prioritize catching asymmetric risk โ clauses that disproportionately favor one party โ while the others lean toward surface-level summarization.
Why Context Window Size Isn't the Full Story
Everyone talks about context windows like bigger automatically means better. Gemini 1.5 Pro can technically handle over a million tokens โ way more than this 40-page contract needed. But having room to read the whole document doesn't mean the model reasons well about what it read.
Here's the mental model that actually matters: context window is about capacity, not comprehension. Gemini could "see" every page, but it didn't weigh the indemnification clause as more important than the renewal date clause. It treated everything with roughly equal attention, which is why it produced a flat summary instead of a risk-ranked one.
Claude, even with a smaller context window at the time, performed better because of how it prioritizes information hierarchically when you ask it to. When I changed the prompt to "Act as a contract attorney reviewing this on behalf of the tenant โ flag anything that disproportionately benefits the landlord", Claude's output sharpened significantly. It started reasoning from a perspective, not just extracting facts.
This is the workflow most people skip: giving the AI a role and a side to advocate for. A neutral prompt gets you a neutral summary. A prompt with a clear point of view forces the model to actually evaluate risk instead of just listing what's there.
ChatGPT improved with this technique too, but it still needed a second follow-up prompt โ "Now explain what a landlord's lawyer would say in response to each of these flags" โ to get genuinely useful back-and-forth analysis. That two-step process (advocate, then counter-advocate) consistently produced the sharpest results across all three tools.
How to Actually Use This Today
Don't just paste your contract and ask "any red flags?" That vague prompt is why most people get mediocre results from AI contract review. Here's the exact workflow that worked:
Step 1: Upload the document to Claude first (claude.ai supports PDF uploads directly). Use this prompt: "You're reviewing this contract on behalf of [your role โ tenant, freelancer, buyer, etc.]. Identify every clause that creates risk, cost, or obligation that isn't clearly reciprocal. Rank by financial impact."
Step 2: Take Claude's flagged list and run it through ChatGPT with: "For each of these flagged clauses, explain what a lawyer for the other party would argue in defense, and suggest alternative wording that's more balanced." This gives you the counter-argument you'll actually face in negotiation.
Step 3: Use Gemini as your fact-checker for dates, numbers, and cross-references โ this is genuinely where it shines. Prompt: "Extract every deadline, dollar amount, and renewal date mentioned in this document into a table." Gemini's long context handles this kind of exhaustive extraction better than the other two.
This three-tool workflow takes maybe 20 minutes and costs nothing beyond what you're likely already paying for these subscriptions. It's not a replacement for a real lawyer on high-stakes contracts, but it will catch 90% of what a first-pass legal review would catch โ before you ever pay for one.
The Part Most People Get Wrong
Most people pick one AI tool and assume it's "good enough" for everything, including contract review. That's wrong, and this test proves exactly why.
Each model has a different reasoning style baked into how it was trained. Claude tends toward cautious, risk-focused analysis. ChatGPT is strong at explaining things clearly but can undersell risk if you don't push it. Gemini is excellent at exhaustive extraction but weak at prioritization. Treating them as interchangeable means you'll get inconsistent results depending on which one you happened to open that day.
The bigger mistake is trusting any single AI output on a legal document without a verification pass. These models don't know your specific legal jurisdiction, they don't know case law updates from the last few months, and they will confidently present a wrong interpretation with the same tone as a correct one.
The fix isn't "don't use AI for contracts." It's "use two AIs against each other" โ one to flag risk, one to argue the counter-position โ before you ever involve expensive human legal review.
Key Takeaways
- Claude wins on risk detection: It consistently flagged the highest-severity clauses first, especially around liability and indemnification.
- Context window size doesn't equal comprehension: Gemini's massive context didn't translate into better prioritization of what actually mattered.
- Role-based prompting changes everything: Asking the AI to advocate for a specific side produces sharper analysis than a neutral "find red flags" prompt.
- Use multiple tools, not one: Claude for risk-flagging, ChatGPT for counter-argument, Gemini for exhaustive fact extraction.
- AI review is a first pass, not a final answer: Always verify high-stakes clauses with an actual lawyer before signing anything significant.
What to Do Right Now
Open Claude.ai, upload the next contract you need to sign โ even something small like a freelance agreement or apartment lease โ and use this exact prompt: "You're reviewing this on behalf of [your role]. Flag every clause that creates risk, cost, or obligation that isn't clearly reciprocal, ranked by financial impact." Do it before your next signature, not after.