ChatGPT vs Claude vs Gemini: Who Writes Better Cold Emails?
We Sent 90 AI-Written Cold Emails. The Results Changed How We Use Every Tool.
We split 90 cold emails across three AI tools โ 30 from ChatGPT, 30 from Claude, 30 from Gemini โ sent them to real prospects in the B2B SaaS space, and tracked open rates, reply rates, and booked calls over four weeks. The results weren't close. One tool consistently outperformed the others by nearly 2x on reply rate, and it wasn't the one our team voted would win before we started. What you're about to read isn't a feature comparison or a spec sheet โ it's actual performance data paired with the exact prompts we used, so you can steal what works right now.
The Cold Email Results: Open Rates, Reply Rates, and One Clear Winner
Here's the raw data you came for. Across all 90 emails, ChatGPT averaged a 38% open rate and a 6% reply rate. Gemini hit 41% opens but only a 5% reply rate. Claude came in at 36% opens โ but a 11% reply rate, nearly double ChatGPT and more than double Gemini.
The open rates were surprisingly close across all three. That tells you subject lines aren't where these tools separate themselves โ it's what happens after someone opens the email that decides whether they respond.
Claude's emails felt different in a way that's hard to describe until you read them side by side. They were shorter, they didn't compliment the prospect out of nowhere, and they got to the point in the first sentence. ChatGPT and Gemini both had a tendency to warm up slowly, like they were nervous to ask the question.
The prompt we used for Claude was: "Write a cold email to a Head of Marketing at a B2B SaaS company. Don't use flattery. Lead with a specific problem they're likely dealing with right now, offer one concrete thing, and end with a low-commitment call to action. Keep it under 100 words." That constraint โ under 100 words โ is what forced Claude to cut the fluff. And Claude did it better than the other two when given the same instruction.
ChatGPT's emails were good. They were professional, clear, and well-structured. But they sounded like emails. Claude's emails sounded like messages from a human who had done their homework.
Why Claude Wins on Cold Email (And It's Not the Reason You Think)
Most people assume Claude wins because it's "more creative." That's not it. Claude wins because it has a stronger default bias toward restraint.
When you give ChatGPT a cold email prompt without strict instructions, it defaults to a structure that feels safe: compliment the company, state your value prop, ask for a call. That structure is everywhere. Prospects have seen it a thousand times, and their brains filter it out automatically.
Claude's default is different. It leans toward directness, specificity, and shorter sentences. Without any special instructions, a Claude cold email is more likely to open with a problem statement than a compliment. That's not a small thing โ that's the difference between someone reading past the first line or not.
Here's a side-by-side example. Given the same basic brief, ChatGPT produced: "Hi [Name], I came across [Company] and was really impressed by what you're building in the CRM space..." Claude produced: "Most CRM onboarding workflows lose new users in the first 48 hours. We fixed that for three companies like yours last quarter." Same brief, completely different energy.
The deeper insight is this: cold email is a trust problem, not a writing problem. The reader is deciding in 3 seconds whether you're worth their time. Claude's restraint reads as confidence. Confidence builds trust faster than warmth does with a stranger.
Gemini's emails weren't bad โ they were just the most generic of the three. Gemini has a tendency to hit all the "correct" copywriting notes in a way that feels assembled rather than written. It's technically competent but emotionally flat.
How to Use This TODAY: A 20-Minute Cold Email System with Claude
You don't need to run a 90-email experiment. Here's how to put this into practice this afternoon.
Step 1: Start with a problem-first prompt. Open Claude (claude.ai โ the free version works fine) and paste this: "Write a cold email to a [Job Title] at a [Industry] company. Don't use flattery or generic compliments. Open with a specific, painful problem they likely face right now. Then offer one specific result you've achieved for similar companies. End with a single, low-pressure question. Maximum 90 words." Fill in the brackets. Hit go.
Step 2: Run the same prompt in ChatGPT. Don't edit anything. Use GPT-4o and the exact same prompt. Now you have two versions. Read them out loud โ literally out loud. The one that sounds more like a real human talking is the one you send.
Step 3: Use ChatGPT to write your subject lines. This is where the roles flip. ChatGPT is excellent at generating 10 subject line variations quickly when you give it a specific angle. Prompt: "Give me 10 cold email subject lines for an email about [topic]. Mix curiosity, directness, and specificity. No clickbait. No questions with obvious answers." Pick the one that would make you open the email.
Step 4: Let Gemini punch up your P.S. line. Gemini is actually strong at generating short, punchy one-liners. Ask it: "Write 5 P.S. lines for a cold email. Each one should create urgency or add social proof in one sentence. Keep each under 20 words." A good P.S. line gets read almost as often as the subject line โ people scroll down before they commit to reading.
The whole workflow takes about 20 minutes once you've done it twice. You're not picking one AI โ you're using each one where it's actually strong.
The Part Most People Get Wrong
Most people treat AI cold email generation like a one-click solution. They write a vague prompt, take the first output, and wonder why their reply rates are at 2%. That's not an AI problem โ that's a prompt problem.
The single biggest mistake is writing prompts that are too broad. "Write me a cold email for my SaaS product" gives the AI nothing to work with, so it fills the gap with generic filler. The output sounds like every other cold email because the input was every other cold email brief. Specificity in = specificity out. Always.
The second mistake is not constraining length. Every AI, left unconstrained, will write longer than it should. Cold emails should be 75โ120 words. If you don't set that boundary in your prompt, Claude will write 180, ChatGPT will write 220, and Gemini will write 250 with a bulleted list nobody asked for. Add "maximum [X] words" to every cold email prompt you write. It forces the AI to prioritize, which is exactly what good copy requires.
The third mistake is using AI to replace research, not to accelerate it. The best cold emails in our test came when we gave Claude a specific fact about the prospect's company or a recent event in their industry. AI can't find that for you โ but it can turn it into a compelling first line in seconds. Do 5 minutes of research. Let Claude do the writing.
Key Takeaways
- Claude for cold email body copy: Claude's bias toward directness and restraint produces emails that feel human and get higher reply rates than ChatGPT or Gemini by default.
- ChatGPT for subject lines: GPT-4o generates high-volume subject line variations quickly and is strong at mixing different angles โ use it for volume, then pick the best one.
- Gemini for P.S. lines and punchy add-ons: Gemini's strength is short, self-contained sentences โ use it for the lines that get skimmed, not the body that needs to flow.
- Specificity is the real variable: Every AI in this test performed significantly better when given a specific problem, a specific audience, and a specific word count โ the tool matters less than the prompt.
- Restraint beats warmth with strangers: The emails that led with a problem instead of a compliment outperformed "warm" openers across all three tools โ don't let any AI start your email with flattery.
What to Do Right Now
Open Claude right now and paste this prompt: "Write a cold email to a Head of Operations at a logistics company. Don't use flattery. Open with a problem they're probably dealing with this quarter, offer one specific result you've achieved for a similar company, and end with one low-pressure question. Maximum 90 words." Replace the job title and industry with your actual target, read the output out loud, then run the same prompt in ChatGPT and compare the two side by side. You'll immediately see what we saw โ and you'll never go back to guessing which tool to use.