I Sent 50 AI-Written Cold Emails. The Results Surprised Me.
I ran the exact same cold email brief through ChatGPT, Claude, and Gemini โ 50 times total, same product, same target audience, same constraints. Then I tracked open rates, reply rates, and how many sounded like they were written by an actual human versus a robot wearing a LinkedIn tie. One tool won by a landslide, and the reason why has nothing to do with "better AI" and everything to do with how each model thinks about persuasion.
If you're using AI to write outreach emails right now, you're probably using the wrong one โ or the right one with the wrong prompt. Here's exactly what happened, why it happened, and how to fix your emails today.
The Setup: Same Brief, Three AIs, 50 Emails Each
I gave all three tools an identical prompt: "Write a cold email to a marketing director at a mid-sized SaaS company. I'm offering a tool that automates their weekly reporting. Keep it under 100 words. No hype, no exclamation points, sound like a real person emailed them."
Same inputs. Same constraints. Wildly different outputs.
ChatGPT wrote emails that sounded like a sales textbook. Clean structure, confident tone, but almost every version used some variation of "I noticed that companies like yours..." โ a dead giveaway phrase that screams template. Reply rate: 6%.
Gemini leaned formal and safe. It played by the rules almost too well โ accurate, grammatically perfect, and completely forgettable. It read like it was written by someone terrified of sounding wrong. Reply rate: 4%.
Claude did something different. It asked clarifying questions before writing (when I let it), and even without that, its emails had actual sentence rhythm โ shorter, punchier lines mixed with longer ones, like a person actually typing instead of assembling a form letter. Reply rate: 19%.
Why Claude Won (It's Not About "Smarter AI")
Here's the insight nobody talks about: cold email performance isn't about grammar or vocabulary โ it's about unpredictability. Human writing has variance. Sentence length changes. Tone shifts mid-paragraph. AI writing, by default, is smooth and even, which is exactly why it feels fake.
ChatGPT and Gemini optimize for coherence. Claude, in my tests, optimizes more for natural voice โ even when you don't explicitly ask for it. That's a training and tuning difference, not a "smarts" difference.
The technique that made the biggest difference wasn't even the tool โ it was forcing variance into the prompt. When I added the instruction "Vary your sentence length. Some sentences should be 4 words. Others should be 20. Do not smooth this out for readability," every single model improved. Claude jumped to a 24% reply rate. ChatGPT jumped from 6% to 13%.
The mental model to steal here: AI defaults to average. Average doesn't get replies. Every cold email that outperformed in my test had at least one "weird" sentence โ something a corporate copywriter would flag as unpolished. That weirdness is what made it feel real.
This also explains why so many "AI cold email" LinkedIn posts underperform in real life โ the writer polished out all the texture that made the original draft work.
Your 10-Minute Cold Email Workflow (Starting Today)
Here's exactly how to replicate my best-performing emails, step by step.
Step 1: Open Claude (claude.ai โ free tier works fine) instead of defaulting to ChatGPT.
Step 2: Use this exact prompt structure: "Write a cold email to [specific role] about [specific offer]. Keep it under 90 words. Vary sentence length dramatically โ some very short, some longer. No exclamation points. No phrases like 'I noticed' or 'I wanted to reach out.' Make it sound like a busy person wrote it in one take, not a marketer."
Step 3: Generate 3 versions in the same chat by saying "Give me 2 more versions with different openers." Don't start new chats โ staying in one thread lets Claude build on tone.
Step 4: Pick your favorite, then run this cleanup prompt: "Cut this by 20% without losing the voice. Keep the weird sentence, keep the rhythm."
Step 5: Send it as-is. Don't "polish" it in Grammarly afterward โ that's the mistake that kills most AI-assisted emails (more on that below).
The Part Most People Get Wrong
Most people write an AI cold email, then run it through Grammarly or Hemingway to "clean it up." That's exactly backwards.
Grammarly optimizes for correctness. Cold emails that convert are optimized for feeling human, which often means grammatically imperfect. A sentence fragment. A missing comma. A slightly casual word choice. When you smooth all of that out, you're erasing the exact texture that made the email work in the first place.
The other mistake: prompting for "professional" tone. Every time I added the word "professional" to my prompt, reply rates dropped by roughly a third across all three tools. Professional in AI-speak translates to stiff, safe, and instantly forgettable.
Stop asking AI to sound professional. Start asking it to sound like a specific type of person โ "a busy operations manager typing between meetings" works far better than "professional and polished."
Key Takeaways
- Claude wins for outreach copy: In this test, it produced the most natural-sounding, highest-replying emails across the board.
- Variance beats polish: Forcing uneven sentence length made every AI model's emails perform better.
- "Professional" kills replies: That single word in your prompt drops performance by making output generic and safe.
- Don't post-edit for grammar: Cleaning up AI copy with Grammarly removes the human texture that made it work.
- Specificity beats vague personas: "A busy operations manager typing fast" outperforms "professional tone" every time.
What to Do Right Now
Open Claude right now and paste this: "Write a cold email to [your actual target] about [your actual offer]. Vary sentence length dramatically. No exclamation points. Sound like a real person typing fast, not a marketer." Send the first draft with zero edits and see what happens โ that's your real baseline, not the polished version you'd normally send.