I Fed the Same Cringe-Worthy Cold Email to ChatGPT, Claude, and Gemini — Here's What Happened
Most cold emails are terrible. They're long, self-centered, and read like they were written by someone who just discovered the thrill of bullet points. I took one of the worst examples I could find — the kind that makes you wince and immediately hit delete — and ran it through all three major AI models to see who could save it.
The results were genuinely surprising. Not just because one model clearly dominated, but because each AI revealed something distinct about how it thinks about persuasion, tone, and what actually makes someone respond to an email. By the end of this article, you'll know exactly which model to use for sales copy, which one to avoid, and why the winner wasn't ChatGPT.
Here's the Terrible Email I Used — And Why It's Perfect for This Test
The original email was a real cold outreach message I found shared in a sales community as an example of what not to do. Here it is, unedited:
"Hi, my name is Jason and I work at TechSolve Pro. We are a leading provider of innovative B2B software solutions that help businesses like yours streamline operations and maximize ROI. I wanted to reach out because I believe our product could be a great fit for your company. We have helped hundreds of clients achieve incredible results. Would you be open to a quick 30-minute call this week to discuss how we can help you too? Looking forward to hearing from you. Best, Jason"
This email commits every sin in the book. It's all about Jason and his company, not the reader. The value proposition is completely vague — "incredible results" means nothing. There's no personalization, no specific pain point, and the call-to-action is a huge commitment (30-minute call) with zero warmup.
It's the perfect test case because fixing it requires real judgment. An AI can't just swap out words — it needs to understand why this email fails and restructure the entire logic of the pitch.
I gave each model the same prompt: "Rewrite this cold sales email so it's more likely to get a response. Make it feel personal, lead with value, and keep it under 100 words." No extra coaching, no system prompts, no second chances. First output only.
ChatGPT's Rewrite Was Good — Just Not Good Enough
ChatGPT (GPT-4o) produced a clean, competent rewrite. It shortened the email, removed the fluff, and added a clearer value hook. Here's the output:
"Hi [Name], I noticed [Company] recently expanded into [market] — congrats. A lot of teams in that position struggle with [specific pain point], and it usually costs them [X outcome]. We helped [Similar Company] solve it in 60 days. Worth a 10-minute chat? — Jason"
Honestly? That's solid. It flipped the focus from Jason to the reader, introduced specificity with bracketed placeholders, and cut the ask down to 10 minutes instead of 30. The structure follows a proven problem-agitate-solution arc.
But here's the issue: ChatGPT played it safe. The placeholders like [specific pain point] and [Similar Company] tell you where personalization should go, but they don't push you to do the harder work of finding a genuinely compelling angle. It's a template, not a transformation.
GPT-4o is excellent at structural rewrites — it knows the rules of good copywriting cold. But it tends to produce "correct" rather than "compelling." If you showed this to a seasoned sales trainer, they'd say it was a B+ effort. Technically right, but not the kind of email that makes someone stop scrolling.
For sales copy, ChatGPT is great when you already know your ICP (ideal customer profile) cold and just need clean execution. If you're still figuring out the angle, it won't help you find it.
Claude's Rewrite Did Something the Other Two Completely Missed
Claude 3.5 Sonnet was the one that genuinely shocked me. I gave it the exact same prompt with zero extra context, and it came back with this:
"Hi [Name] — quick question: is [specific bottleneck] slowing your team down right now? We helped [Peer Company] cut that time by 40% in their first month. No pitch, no 30-minute call — just a 3-question email back if it's relevant. Worth it? — Jason"
That last line — "just a 3-question email back" — is the detail that separates good copy from great copy. Claude didn't just lower the friction; it reframed the entire ask. Instead of a calendar commitment, it offered a conversation that costs the reader almost nothing. That's a fundamentally different persuasion strategy, and Claude invented it without being told to.
Claude (made by Anthropic) tends to reason about intent more deeply than its competitors. It didn't just clean up Jason's email — it questioned the underlying strategy of asking for a call at all and replaced it with something smarter. That's not template thinking. That's editorial judgment.
The reason this matters for your workflow: Claude is the better choice when the original content is strategically broken, not just poorly written. If someone hands you a deck, a landing page, or an email where the whole approach is wrong, Claude is more likely to identify the real problem rather than polish a flawed strategy.
One caveat — Claude's outputs can occasionally feel slightly more "writerly" than punchy. For B2B sales specifically, you may want to tighten Claude's draft slightly. But that's a 30-second edit, not a structural problem.
Gemini's Rewrite Tells You Exactly When to Use Google's Model
Gemini 1.5 Pro gave me the longest output of the three. It rewrote the email, then added a breakdown of what it changed and why, then offered two alternative versions with different tones — one more formal, one more casual.
Here's its primary rewrite:
"Hi [Name], I came across [Company]'s recent [specific event/content] and thought of a problem we see often in your space: [pain point]. We solved this for [Client Name], reducing their [metric] by [X%] in under 90 days. Would a quick email exchange make sense to see if there's a fit? — Jason"
It's fine. It hits most of the same notes as the ChatGPT version — reader-focused, specific, low-friction ask. But what's more interesting is what Gemini did around the rewrite.
The multi-version output is actually Gemini's superpower in this context. If you're A/B testing email sequences, or you're a marketing manager who needs to show options to a client, Gemini's tendency to offer variations without being asked is genuinely useful. You're not just getting one answer — you're getting a menu.
Gemini also pulled in a more data-driven framing than the other two, adding "reducing their [metric] by [X%]" even though my original email had no numbers at all. That's Gemini leaning into its strength: it's been trained heavily on structured, factual content, and it defaults to quantifiable proof points naturally.
Where Gemini falls short is depth of strategic insight. It's excellent at execution and variation — but it didn't question the strategy the way Claude did. It improved the email; it didn't reimagine it.
Use Gemini when you need multiple solid options fast, especially if you're working inside Google Workspace and want to pipe the output directly into Gmail or Docs.
The Part Most People Get Wrong
Most people use AI to fix the words in a bad email. That's wrong. The words are almost never the real problem.
Bad sales emails fail because of bad strategy — wrong ask, wrong timing, wrong focus. If you just ask an AI to "make this better," you'll get a polished version of the same mistake. You need to explicitly tell the AI to question the strategy, not just the copy. The prompt "Rewrite this so it's better" will get you cosmetic improvements. The prompt "What's fundamentally wrong with this email's approach, and how would you fix it from the ground up?" will get you a real rethink.
This is why Claude won this test — not because it's a better writer, but because it reasoned about the strategy before touching the words. You can actually force ChatGPT and Gemini to do the same thing by changing your prompt. Try: "Before you rewrite this email, tell me the three biggest strategic mistakes in it. Then rewrite it." That one shift will dramatically improve the output from any model.
The other mistake people make is running these tests without a rubric. "Which one sounds better?" is a feeling, not a metric. Judge AI sales copy on: open-rate potential (subject line hook), time-to-value (how fast does the reader understand what's in it for them), friction of the ask (is the CTA appropriate for a cold contact?), and specificity (does it sound like it was written for this person?). Score each rewrite on those four dimensions and the winner becomes obvious, not subjective.
Key Takeaways
- Claude 3.5 Sonnet: Best for strategic rewrites where the original content is fundamentally broken — it questions the approach, not just the execution.
- ChatGPT (GPT-4o): Best for clean, structured execution when you already know your angle — fast, reliable, and follows copywriting frameworks well.
- Gemini 1.5 Pro: Best when you need multiple variations quickly, especially inside the Google ecosystem — it defaults to options, not a single answer.
- Your prompt matters more than your model: Ask any AI to question the strategy before rewriting, and you'll get dramatically better output regardless of which tool you use.
- Judge AI copy with a rubric: Score outputs on hook strength, time-to-value, CTA friction, and specificity — gut feeling isn't a reliable filter.
What to Do Right Now
Grab the worst email in your drafts folder — the one you've been avoiding sending because it feels off — and paste it into Claude with this exact prompt: "Before you rewrite this email, tell me the three biggest strategic mistakes in it. Then rewrite it in under 100 words so a cold contact is actually likely to respond." Do the same thing in ChatGPT and Gemini. You'll have three completely different rewrites in under 10 minutes, and the comparison alone will teach you more about persuasion than most copywriting courses.