Skip to main content
AI ComparisonChatGPTClaudeGeminiemail writingAI comparison

ChatGPT vs Claude vs Gemini: Rewriting Bad Email Threads

We fed the same passive-aggressive email chain to all three AIs. The winner shocked us — and saved a client relationship.

D
Davide
··8 min

We Fed a Passive-Aggressive Email Chain to ChatGPT, Claude, and Gemini — Here's What Happened

Most AI comparisons test the easy stuff: summarize this, write me a poem, explain quantum physics. We went harder. We took a real, ugly email thread — the kind where a client is clearly furious but wrapping it in corporate politeness — and asked all three major AIs to rewrite it into something that could actually save the relationship. The results were not what we expected. One AI produced a response so good we genuinely used it with a client. One produced something that would have made things worse. And one gave us the most technically correct answer that somehow felt completely hollow. By the end of this article, you'll know exactly which AI to reach for when the stakes are high — and why using the wrong one in a professional context could cost you more than time.


The Actual Email Thread We Used (and Why It Was a Perfect Test)

The email chain we used was three messages deep. It started with a client asking about a delayed deliverable — politely. Then a second message, three days later, slightly sharper. Then a third: "I just want to make sure we're aligned on expectations going forward."

If you've ever worked in client services, you know exactly what that sentence means. It's not a question. It's a warning shot dressed in a blazer.

We gave all three AIs the same setup prompt: "Here is an email thread with a client. The client is clearly frustrated but hasn't said so directly. Rewrite the final reply from our side — acknowledge the delay, take ownership, rebuild trust, and keep the relationship intact. Don't be defensive. Don't over-apologize. Sound like a confident professional, not a groveling intern."

That last instruction matters. It's the difference between a reply that sounds human and one that sounds like it was written by someone who just completed a mandatory HR training module.

The test wasn't just about grammar or tone. It was about emotional intelligence in text form — which is genuinely one of the hardest things to get right, even for humans.


How Each AI Actually Responded (The Honest Breakdown)

ChatGPT (GPT-4o) went straight to confident. Its rewrite opened by acknowledging the delay without burying the lede, used a single clean apology instead of three nervous ones, and pivoted quickly to a specific action plan with a revised timeline. It read like someone who'd handled this before. The problem? It was slightly too polished. One phrase — "I want to assure you that your project is our top priority" — felt like it came from a template. It was good. It wasn't great.

Gemini (Google's Gemini 1.5 Pro) gave us the longest response by far. It restructured the email into labeled sections — almost like a business document — and included a proposed "check-in schedule" that nobody asked for. It wasn't wrong, exactly. But it completely missed the emotional temperature of the situation. The client didn't want a project management framework. They wanted to feel heard. Gemini treated the problem like a logistics issue, not a relationship issue.

Claude (Claude 3.5 Sonnet) did something neither of the others did: it opened with a line that named the situation without naming it. Something close to: "I can see how the lack of updates over the past week would be frustrating — that's on us, and I want to fix it." That sentence does three things at once. It validates the client's feeling, it takes clear ownership, and it moves forward. No drama. No over-explanation. Just signal.

The key difference wasn't vocabulary or length. It was that Claude seemed to understand the subtext of the email — not just what was said, but what the client actually needed to hear. ChatGPT answered the surface. Claude answered the room.


The Hidden Skill These AIs Are Actually Testing: Emotional Register

Here's the thing most AI comparison articles miss entirely: the challenge in rewriting a tense email isn't grammar. It's calibrating emotional register — matching the level of gravity the other person is feeling without escalating it or dismissing it.

Think of it like a dial. On one end: overly casual ("Hey, totally my bad on that one!"). On the other end: overly formal and defensive ("Please be advised that all deliverables are subject to revision timelines as outlined in Section 4 of our agreement"). The right response lives in the narrow band between them.

Claude's training — which heavily emphasizes nuanced, contextual language — gives it a structural advantage here. Anthropic explicitly trained Claude on tasks requiring careful, multi-layered communication, and you can feel it in outputs like this one. It's not a coincidence. It's a design choice.

This doesn't mean ChatGPT is useless for email. It's excellent for high-volume, lower-stakes communication — first outreach, follow-ups, status updates. But when there's real emotional weight in the thread, Claude's ability to read and match tone is genuinely ahead.

A useful mental model: think of ChatGPT as the confident generalist, Gemini as the thorough analyst, and Claude as the person in the room who actually reads people. Different strengths. Different moments. Knowing which to reach for is the whole game.


How to Use This in Your Own Work Right Now

Here's the exact workflow to try today — no setup required, no paid tools beyond what you probably already have.

Step 1: Copy any email thread where the relationship feels even slightly tense. It doesn't need to be a crisis. Even a slightly clipped reply from someone who's usually warm is worth running through this.

Step 2: Go to Claude (claude.ai — free tier works). Paste the thread and use this prompt:

"Here's an email thread. The tone from the other person has shifted — they seem frustrated, even if they haven't said so. Rewrite my last reply to: acknowledge the situation honestly, take appropriate ownership, avoid being defensive or over-apologetic, and move the conversation forward. Keep it under 150 words."

The word count constraint is important. It forces the AI to prioritize — and it usually forces out the filler phrases that make replies feel corporate.

Step 3: Read the output out loud. This sounds basic, but it's the fastest way to catch anything that feels off. If you stumble on a sentence, so will your client.

Step 4: Make one or two personal edits — add a specific detail only you would know, or adjust a phrase to match your natural voice. This step takes 90 seconds and makes the whole reply feel real instead of generated.

The total time from "I need to reply to this" to "reply sent" is under 10 minutes. You can do this for free, today, with any tense email sitting in your drafts right now.


The Part Most People Get Wrong

Most people paste an email into ChatGPT, read the first output, and send it. That's wrong — and not for the reason you think.

The problem isn't that the AI output is bad. The problem is that the first output is the average answer. It's what the AI produces when it's playing it safe. The second or third iteration — after you push back, add constraints, or ask for a different approach — is almost always sharper. Treat the first draft as a starting point, not a finished product.

The other big mistake is using the wrong AI for the context. Running a high-stakes client recovery email through Gemini because it's free and you're already in Google Workspace is like asking your most detail-oriented colleague to handle your most emotionally sensitive call. Technically capable. Wrong person for the moment.

Finally, people forget to edit for voice. AI-written emails have a fingerprint — slightly too balanced, slightly too structured. One genuine, specific detail breaks that pattern entirely. Mention the actual project name. Reference something from your last call. That specificity is what turns a good AI-assisted email into one that sounds like you wrote it on your best day.


Key Takeaways

  • Claude 3.5 Sonnet: Best choice for emotionally complex or high-stakes email rewrites — it reads subtext better than the competition right now.
  • ChatGPT (GPT-4o): Reliable for professional, high-volume communication where tone is neutral and speed matters.
  • Gemini 1.5 Pro: Strong at structure and thoroughness, but tends to miss emotional nuance — use it for analytical tasks, not relationship repair.
  • Emotional register: The real skill being tested in tense email rewrites — match the client's level of gravity without escalating or dismissing it.
  • The first AI draft: Always a starting point, not a finished product — push for a second iteration and add one personal detail before sending.

What to Do Right Now

Open Claude (claude.ai) and find one email in your drafts or sent folder that felt uncomfortable to write. Paste the thread and use this prompt: "Rewrite my reply to acknowledge the other person's frustration, take clear ownership, and move forward — no more than 150 words, no defensive language." Read it, make one personal edit, and you'll see exactly what we're talking about in under 10 minutes.

ChatGPTClaudeGeminiemail writingAI comparison

Get Weekly AI Insights Delivered Free

Join 5,000+ subscribers getting the latest AI tool breakdowns, prompts, and strategies every week. No spam, ever.

  • ✦ Weekly AI tool reviews
  • ✦ Exclusive prompt packs
  • ✦ Early resource access
  • ✦ No spam, unsubscribe anytime

By subscribing, you agree to our Privacy Policy. Unsubscribe anytime.