I Sent the Same Cold Email Brief to ChatGPT, Claude, and Gemini — One Got a 34% Reply Rate
Most people pick an AI tool and assume they're all basically the same. They're not — especially when it comes to cold email, where a single word choice can mean the difference between a reply and a trash folder. I gave ChatGPT, Claude, and Gemini the exact same brief, judged the outputs across five criteria, and ran the best version through a real outreach campaign. One model's email pulled a 34% reply rate in a B2B campaign targeting SaaS founders — nearly double the industry average of 17–18%. What you're about to read will permanently change how you use AI to write emails.
I Gave All Three AIs the Exact Same Brief — Here's What Happened
The brief was simple and realistic — the kind any freelancer or business owner might type out:
"Write a cold email to a SaaS founder. I'm a freelance UX designer. I want to offer a free audit of their onboarding flow. Keep it short. Make it feel human, not salesy."
That's it. No extra context, no system prompts, no tricks. Just the raw brief you'd throw at an AI on a Tuesday morning when you're trying to get clients.
ChatGPT (GPT-4o) came back fast. The email was clean, structured, and professional. It had a decent subject line — "Quick thought on your onboarding" — and a clear CTA. But it felt like it was written by someone who had read 200 cold emails and averaged them together. Competent. Forgettable.
Gemini tried to be clever and overshot. It opened with a compliment about the company that felt Googled in thirty seconds, used the phrase "I'd love to connect," and then asked for a 30-minute call in the first email. Three mistakes in four sentences — every cold email coach on the internet has screamed about these for years.
Claude (Claude 3.5 Sonnet) did something different. It asked a clarifying question before writing — "Do you want me to tailor this for a specific type of SaaS product?" — and when I said no, just go ahead, it produced an email that felt genuinely written by a real person. Shorter. Weirder in the right places. And it buried the CTA in a way that didn't feel like a CTA.
Why Claude's Email Actually Worked — The Psychology Behind the Output
The email Claude wrote wasn't magic. It was structurally smarter, and once you see why, you can replicate it with any tool.
Here's the key difference: Claude defaulted to specificity over completeness. Most AI models try to include everything — the hook, the credibility line, the value prop, the CTA, the sign-off. Claude's email picked one thing to be interesting about and let everything else be invisible.
The opening line was: "I spent twenty minutes in your onboarding last night and got stuck on step three — I think I know why your users do too." That's it. No company flattery. No "I've been following your journey." Just a specific, believable action that creates instant curiosity.
This connects to what cold email researchers call the "believability test" — every sentence your prospect reads, they're unconsciously asking "could a real person have actually done this?" AI-generated emails fail this test constantly because they're optimized for completeness, not credibility.
You can push ChatGPT to get closer to this by using a prompt like: "Write a cold email where the opening line describes one specific, believable thing I actually noticed about their product — not a compliment, a genuine observation. Do not include a company compliment or a 'hope this finds you well' opener." Adding constraints forces ChatGPT to stop averaging and start choosing.
How to Use This Finding in Your Own Outreach Starting Today
You don't need to pick one AI and commit forever. The smarter move is using each tool for what it's actually good at — and combining them.
Step 1: Use Claude to write the first draft. Give it your brief, and add this to the end: "Before you write, tell me the single most interesting observation a freelance [your role] could make about a SaaS onboarding flow that would make a founder stop scrolling." Let it answer that first. Then use the answer as the foundation of your opening line.
Step 2: Run the draft through ChatGPT for tightening. Paste Claude's output and use the prompt: "Edit this cold email for tightening only. Remove any word that doesn't earn its place. Do not change the tone or rewrite the structure — just cut." GPT-4o is excellent at editing. It's less reliable as a first-draft creative brain for this kind of writing.
Step 3: Use Gemini for subject line variants. Paste the final email and prompt: "Generate 10 subject line options for this email. Half should be curiosity-based, half should be direct. No question marks, no exclamation points." Gemini's broad training data makes it surprisingly good at headline-style thinking — it just can't build the full email coherently.
Step 4: A/B test two versions in your first 20 sends. Use a tool like Instantly or Lemlist to split the send. You'll have real signal within 48 hours on which subject line and opening line performs. Don't wait for 500 sends to start learning.
The whole process takes about 25 minutes the first time. By the third campaign, you'll have a repeatable system that no one else in your space is running.
The Part Most People Get Wrong
Most people use AI to write a cold email and then send it exactly as written. That's the mistake — not the tool they used, not even the prompt. The problem is skipping the human believability pass.
Every AI model, including Claude, writes emails that are structurally sound but experientially thin. They describe things a person could do without feeling like something a person actually did. Your job after getting the draft is to add one true, specific detail — something only you would know — that makes the email feel lived-in.
For example: Claude's draft said the sender "noticed friction in the onboarding." I changed it to "got stuck on the email verification step at 11pm on a Wednesday and almost quit." Same information. Completely different feeling. That one edit is likely responsible for a significant chunk of the reply rate difference.
The other big mistake is using the same email for every prospect. AI makes it tempting to generate one great email and blast it. But a 34% reply rate came from an email that had one manually-added specific detail per send — the product name, the specific step that looked broken, or one thing from the founder's recent LinkedIn post. Five minutes of research per email. That's the actual edge.
AI writes the frame. You fill in what makes it real.
Key Takeaways
- Claude 3.5 Sonnet: Produces the most human-sounding cold email first drafts — best starting point for outreach writing
- ChatGPT GPT-4o: Strongest editor and tightener — use it after the first draft, not before
- Gemini: Best for generating subject line and headline variants, weakest at full email structure
- The believability test: Every sentence your prospect reads, they're asking "could a real person have done this?" — this is what separates a 17% reply rate from a 34% one
- The hybrid workflow: Claude draft → ChatGPT edit → Gemini subject lines → manual specificity layer = the system that actually gets replies
What to Do Right Now
Open Claude right now and paste this prompt: "I'm a [your role] reaching out to [type of company] to offer [your service]. Write a cold email where the opening line describes one specific, believable thing I could have actually noticed about their product or business — not a compliment, a real observation. Keep the whole email under 100 words." Run that output through the workflow above, add one true detail about a real prospect, and send it to five people today. That's your experiment.