We Gave ChatGPT, Claude, and Gemini the Same Cold Email Brief β One Blew Us Away
Most AI cold email comparisons are useless. They pick a vague prompt, paste three generic outputs side by side, and call it a "test." We did it differently β same specific brief, same target audience, same goal, judged on the only metric that actually matters: would a real person reply to this? The results were genuinely surprising, and if you're using AI to write outreach right now, what we found could change which tool you reach for first.
The Exact Brief We Used β And Why It Was Designed to Be Hard
We didn't give the AIs an easy job. The brief was deliberately challenging because easy prompts produce identical results from every tool.
Here's the exact prompt we fed all three:
"Write a cold email to a Head of Marketing at a mid-sized SaaS company. The sender is a freelance video editor. The goal is to get a 15-minute discovery call. The email should be under 120 words, have a specific hook, and not use the phrase 'I hope this email finds you well.' Tone: confident, concise, human."
Word count limit. Banned phrases. A specific job title. A concrete ask. This is the kind of brief you'd actually write in real life β or should be writing β and it separates the tools that can handle nuance from the ones that just pattern-match to "professional email template #7."
We ran each prompt three times to account for variation, then picked the strongest output from each tool. Here's what we got.
How Each AI Handled the Hook β This Is Where They Split Apart
The hook is everything in cold email. You have roughly three seconds before someone hits delete.
ChatGPT (GPT-4o) opened with: "Your last three product launches probably had great features. The videos, though β that's where most SaaS brands quietly lose deals." Specific, a little provocative, directly relevant to a Head of Marketing's actual pain. It didn't just reference the industry β it named a specific moment (product launches) and a specific failure mode (bad video). That's a real hook.
Claude (Claude 3.5 Sonnet) went a different direction: "I watched your demo video on the pricing page. Thirty seconds in, I wanted to keep watching β but the pacing made me click away. I can fix that." This felt the most human of the three. It implied research, used a first-person observation, and made a specific, testable claim in under 25 words. The "I can fix that" ending is almost uncomfortably direct β which is exactly why it works.
Gemini (Gemini 1.5 Pro) opened with: "As a video editor specializing in SaaS content, I've helped brands like yours increase engagement through compelling visual storytelling." And that's where it fell apart. "Brands like yours." "Compelling visual storytelling." These are the phrases every Head of Marketing has seen a thousand times. The hook didn't hook β it blended in.
The pattern here is important: ChatGPT was confident and punchy, Claude was specific and human, Gemini defaulted to safe corporate language. That gap in the hook tells you almost everything about how each tool approaches persuasive writing.
The Deeper Difference No One Talks About β Specificity Under Constraint
Here's the insight most AI comparison articles completely miss: the real test isn't which tool writes well with no constraints β it's which tool writes well when you give it rules.
When you add word limits, banned phrases, and a specific CTA, you're essentially testing whether an AI understands the purpose of a cold email or just the format. Most tools know what a cold email looks like. Far fewer understand why each element exists.
Claude showed the clearest understanding of purpose. Its full email came in at 112 words and included a specific CTA β "Would Tuesday or Wednesday work for a 15-minute call this week?" β that gives the recipient two concrete options instead of a vague "let me know when you're free." That's a small detail with a measurable impact on reply rates. Claude didn't include it by accident β it understood that reducing friction is the job of the closing line.
ChatGPT's email was strong but slightly over-engineered. The hook was great, the body had one line too many, and the CTA was good but generic: "Happy to send over a quick video sample if you'd like β or we can just jump on a call." It gave options, but two different options in the same sentence creates a small decision paralysis that Claude avoided entirely.
Gemini's email stayed under the word count, but it did so by being vague rather than precise β every sentence was safe enough not to break anything, but not specific enough to create real interest. Vagueness is the enemy of cold email. You can have a short email that still has a razor-sharp idea in every line. Gemini didn't achieve that.
The takeaway: Claude wins when precision and human tone matter most. ChatGPT wins when you need a punchy, high-energy hook fast.
How to Run This Test Yourself β And Build a Workflow That Wins
You don't need to just read about who won. You can run this yourself in the next 20 minutes and immediately improve your outreach.
Step 1: Write a real brief. Don't write "write me a cold email for my freelance business." Write something specific: recipient's job title, your offer, desired outcome, word limit, tone, and at least one banned clichΓ©. The more constraints, the more useful the comparison.
Step 2: Run the same prompt in ChatGPT, Claude, and Gemini. Don't tweak it for each tool β the whole point is that you're testing the tools, not your prompting skills. Use the free tiers if you haven't upgraded yet, though GPT-4o and Claude 3.5 Sonnet are noticeably better than their free versions for this kind of task.
Step 3: Score each output on three criteria β Hook quality (would you keep reading?), Specificity (does it feel like it was written for this person?), and CTA clarity (is the ask obvious and low-friction?). Score each out of 10. You'll start to see patterns across your niche.
Step 4: Steal the best line from each. Here's a workflow most people skip β take the best hook from one, the best body from another, and the best CTA from the third. Paste them into one email, clean up the transitions, and you've got something better than any single tool produced on its own. This hybrid approach consistently outperforms single-tool outputs in our testing.
If you want to go further, use ChatGPT's Advanced Voice Mode or Claude's longer context window to iterate on the email in real time β paste your draft and ask "What's the weakest line in this email and why?" The answers are usually uncomfortably accurate.
The Part Most People Get Wrong
Most people treat cold email prompts like order forms. They type "write a cold email for [job] to [person]" and expect magic. That's wrong, and here's why: AI tools write to the level of detail you give them.
If your prompt is vague, the output will be generic β because the AI is filling in gaps with the most statistically average version of a cold email it has ever seen. That average is what everyone else is also sending.
The most common mistake we see is prompting without a persona. You need to tell the AI who the sender is as specifically as who the recipient is. "Write as a freelance video editor with 6 years of SaaS experience who has worked with Notion-style product teams" produces a fundamentally different email than "write as a freelance video editor." The specificity of the sender changes the confidence level, the vocabulary, and the examples the AI reaches for.
The second mistake is accepting the first output. Every tool we tested produced a noticeably better email on the second or third attempt β not because the prompt changed, but because AI outputs have natural variation, and you want to catch it in a high moment, not a low one. Run every cold email prompt at least twice before you judge the tool.
Key Takeaways
- Claude 3.5 Sonnet: Produces the most human-sounding cold emails β ideal when tone and specificity matter more than punch.
- ChatGPT (GPT-4o): Best for high-energy hooks and iterating fast β slightly over-explains, but the raw quality of the opening lines is hard to beat.
- Gemini 1.5 Pro: Defaults to safe, corporate language under constraint β strongest for research and drafting, weaker for persuasive outreach.
- The hybrid method: Combining the best element from each tool's output consistently outperforms any single AI working alone.
- Prompt specificity: The single biggest factor in output quality isn't which AI you use β it's how detailed your brief is before you hit enter.
What to Do Right Now
Open Claude.ai and paste this prompt exactly: "Write a cold email from [your role] to [specific job title] at a [type of company]. Goal: book a 15-minute call. Under 120 words. No clichΓ©s. End with a specific day and time option." Then run the same prompt in ChatGPT. Compare the hooks side by side β you'll immediately see the difference this article just described, and you'll have a real email you can send before the hour is up.