The reply that taught me the most about AI cold email was four words long: "We never announced that." A prospect, replying to a first line that congratulated her company on a Series B. The Series B did not exist. The model, asked to find a trigger for every account on a list of 400, had found one for every account whether or not one was there. About a third were fiction. Beautifully written fiction, in our brand voice, sent from a warmed domain, at the right time of day.
That is the whole problem with the question "what is the best AI for cold email." It assumes one tool does one job. Cold email is two jobs. Somebody has to write a true, short, specific email, and somebody has to get it into an inbox and follow up three times without burning the domain. The tools that are good at the first are useless at the second, and the tools that are good at the second write like a brochure. I run GTM systems at Docket and most of my outbound tooling runs through Claude, so read the model section with that in mind. The sender section has no dog in the fight.
The verdict, in the first hundred words
Two tools. For the writing, a frontier model, and my pick is Claude: it drafts shorter, cuts cleaner, and holds a template as a Skill so the hundredth email comes out in the same shape as the first. ChatGPT is a close second and a fine choice if your team already lives in it. Gemini is the right answer only when the email is a reply you are writing inside Gmail right now.
For the sending, a purpose-built sender: Instantly, Smartlead, or lemlist, pick by budget and taste. They do the plumbing no model does: warm the inboxes, rotate them, watch deliverability, run the three-step sequence, stop on reply. A brilliant email into a burned domain is a brilliant email in a spam folder.
AI SDR tools (11x, Artisan, AiSDR, and the rest of that category) bundle both jobs behind one login. They are a reasonable buy when the volume is bigger than your team and a poor one for the accounts you actually care about, because the personalization is the part they compress first.
TipIf a vendor says its AI does "end-to-end cold email," ask which job it is best at, writing or sending. The honest ones will tell you. The others are selling you a compromise on both.
Cold email is two jobs, and they hate each other
The writing job wants restraint, specificity, and truth. One real trigger, one line of relevance, one ask, under ninety words. It rewards a model that can read a careers page and a press release and notice the thing that matters, then say it in a sentence a human would send. It punishes volume, because volume is where the triggers start getting invented.
The sending job wants reliability, patience, and paranoia. Warm the inbox for two weeks before it sends anything real. Keep daily volume per inbox low. Rotate across domains. Watch the bounce rate like a smoke detector. Follow up on day three and day seven, stop the instant someone replies, and never send to the same domain twice in a week. None of this is intelligence. All of it is plumbing, and plumbing is what the purpose-built senders sell.
The two jobs pull in opposite directions, which is why the all-in-one tools end up mediocre at both. The sender's economics want more emails; the writer's craft wants fewer, better ones. Hold them in separate tools and you can tune each without wrecking the other.
The writers: Claude vs ChatGPT vs Gemini on the email itself
I judge writing tools on edit distance: how much a rep has to change before the draft is sendable. Across a few thousand drafts my experience is that Claude's is lowest. It respects a length rule, it does not reach for "I hope this finds you well," and when the Skill says one trigger and one ask, it produces one trigger and one ask instead of a bulleted tour of your features. The cold email templates skill on this site is the shape I use: the template lives in the Skill, the account data comes in, the email comes out already in our voice.
ChatGPT is close, and with a well-built custom GPT it is very close. Its default register runs a little longer and a little warmer than I want in a cold email, and it likes a bullet list, but those are instruction problems, not model problems. If your team already has ChatGPT installed and trusts it, build the GPT, put the length rule in the instructions, and you will get 90 percent of the way there.
Gemini earns its place in one situation: the email is a reply and you are already in Gmail. The side panel drafts it where you are, off the thread you are looking at, zero tabs away. For a net-new cold sequence generated off a list, it is fine and not where I would start, mostly because the list-to-email loop wants a Project or a Skill, and Claude's version of that is more mature.
- Claude: lowest edit distance, holds a template as a Skill, reads a full account file in one go. My pick for the writing job.
- ChatGPT: close second; longer and warmer by default, fixable with a custom GPT and a hard length rule.
- Gemini: best for replies written inside Gmail; fine for sequences, less mature for the list-to-email loop.
TipPut the length rule in the Skill, the Project, or the GPT instructions, never in the prompt. Reps send what comes out of the box, and the box has to be short by default.
The senders: Instantly, Smartlead, lemlist and the plumbing
This is the part of the stack no model replaces. Instantly and Smartlead are the volume-oriented senders: unlimited or near-unlimited inbox connections, built-in warmup, rotation across many mailboxes, deliverability dashboards, and sequencing with reply detection. lemlist sits closer to the personalization end, with more built-in enrichment and multichannel steps (LinkedIn touches alongside email) at a somewhat higher per-seat price. All three will happily accept AI-written copy; all three also have their own AI writing features, which I mostly turn off, because I want the words to come from the tool with the Skill and the context.
Pick by shape of your motion. Sending from a dozen inboxes at low daily volume each, with cold domains you bought last month, is a Smartlead or Instantly job. A smaller, higher-touch motion with LinkedIn steps and a rep who reads every reply leans lemlist. Prices for all three move often enough that I will not quote them here; check the pricing pages, and budget the extra mailboxes and domains as part of the sender cost, because that is where the real bill lives.
The sender's job is to make the writing job's output land. Treat it that way. Every deliverability setting is a promise to the inbox providers that you are a person, not a cannon, and the model upstream cannot keep that promise for you.
The AI SDR tools: buying the whole loop
11x, Artisan, AiSDR, and their cousins sell a different thing: a virtual rep that finds the accounts, writes the emails, sends them, and books the meeting, behind one login and one invoice. It is a genuine category and it solves a genuine problem, which is that most teams cannot staff the volume they want to run.
The trade is personalization for throughput. To send thousands of emails a week, the writing has to be templated hard, the triggers have to be shallow (a job posting, a funding round, a tech install), and the review step has to be gone. That is fine for a wide net over a big TAM where a 1 percent reply rate on a huge base is the plan. It is a poor fit for the 200 accounts you actually want, where a shallow trigger reads as spam to exactly the buyer you were trying to impress.
My rule: AI SDR tools for the long tail you would otherwise not touch, a model-plus-sender stack for the accounts a human would be embarrassed to lose. If you only have budget for one, decide which of those two lists your quarter depends on.
The stack I run, and why each piece has one job
Four tools, four jobs. Clay holds the list and does the enrichment: the firmographics, the tech stack, the careers page, the recent news, one row per account. Claude does the words: a Skill with our template reads the row and writes the email, one true trigger, one line of relevance, one ask. A sender delivers and follows up. HubSpot records everything, so the AE who picks up the reply can see what was sent and why. The full build is on this site as the personalized cold email use case.
The design rule that came out of the Series B incident: the model may only cite triggers that exist in the row. If Clay found no funding news, the email does not mention funding. If there is no trigger at all, the row gets flagged for a human instead of getting a fabricated one. That single constraint cut our invented-trigger rate to zero and, mildly annoyingly, improved reply rate, because a plain true email beats a clever false one every time.
Where volume is high enough that a human cannot review every email, the review moves to the row level: a Claude Code run scores each account's trigger quality before the email is written, and only rows above the bar get an email. Below the bar goes to a lighter sequence or to nobody. It is the same reasoning the inbound routing agent does for inbound, turned around to face outbound.
TipWrite the constraint into the Skill in these words: "Use ONLY triggers present in the account data. If none, output NO_TRIGGER." Then route NO_TRIGGER to a human. This is the highest-return sentence in my outbound stack.
What "best" means at your volume
The right stack changes with how many emails you actually send, so match the tools to the number rather than to the demo.
- Under 200 a month: a model (Claude in a Project or with the Skill) plus your normal inbox and a light sequencer. You do not need warmup infrastructure; you need better words. Read every reply yourself.
- 200 to 2,000 a month: model plus a real sender. Buy the extra domains and mailboxes, warm them, keep per-inbox volume low. Review at the row level, not the email level.
- Over 2,000 a month: this is where AI SDR tools earn a look for the long tail, while the model-plus-sender stack keeps working the named accounts. Two lanes, two standards.
- Any volume: the trigger must be true. The volume that turns fiction into a policy is the volume you should not be sending.
Pitfalls: how AI cold email goes wrong
Every one of these I have either done or cleaned up after. They are listed roughly in the order they show up.
- The invented trigger. A model asked for a "why now" on every row will find one on every row. Constrain it to the data and route the blanks to a human.
- Brilliant copy, burned domain. Skipping warmup or over-sending per inbox puts your best email in spam. The sender's settings are not optional.
- The six-paragraph cold email. Unedited AI drafts run long. The length rule lives in the Skill or GPT instructions, not in someone's memory.
- Sender AI on top of model AI. Two writing layers fighting each other produces the blandest possible email. Pick one writer and turn the other off.
- Personalization theater. "I saw you are hiring" across 400 companies is a template with a variable, and buyers can tell. If the trigger would fit any company, it is not a trigger.
- No record. Emails that never reach the CRM mean the AE who picks up the reply has no idea what was promised. Log every send.
Where I land
Claude for the words, a sender for the delivery, Clay for the data, HubSpot for the record, and a one-sentence constraint that the trigger has to be real. If you already live in ChatGPT, swap the writer and keep everything else. If you cannot staff the long tail, add an AI SDR tool for that lane and hold the named accounts to a higher standard.
The tool comparison is the smaller half of this. The bigger half is deciding that a plain true email beats a clever false one and building the loop so the model cannot argue. What is the trigger line your team sends most, and would it survive a prospect replying "we never announced that"?
How to set it up
Put the template in a Skill, with the truth constraint
Start from the B2B cold email templates skill. Add your ICP, your voice rules, the 90-word ceiling, and the line "Use ONLY triggers present in the account data. If none, output NO_TRIGGER." Save it as a Skill in Claude Code or paste it into a Project's instructions.
Build the account row in Clay
One row per account: company, size, industry, tech stack, the careers page summary, recent news with dates, and the contact's title. This is the only material the model is allowed to write from, so the quality of the row is the quality of the email.
TipAdd a column for the trigger you found and its source URL. If the column is blank, the model should not be inventing one, and you can verify the ones it does use in a click.
Generate and review at the row level
Run the Skill over the rows. Read the NO_TRIGGER count first; that is your list quality. Then spot-check ten emails against their source rows. A claim without a source is a bug in the prompt, not a one-off.
Load into the sender, with the plumbing done first
Warm the mailboxes for two weeks before the first real send. Keep per-inbox daily volume low, rotate across domains, set the three-step sequence with stop-on-reply. Upload the 371 emails as the first step's body, personalized per row. Turn the sender's own AI writer off.
Log to the CRM and hand the NO_TRIGGER rows to a human
Every send lands on the contact in HubSpot with the trigger used, so the AE reading the reply knows what was promised. The 41 flagged rows go to a rep with the row open; sometimes a human finds the trigger the data missed, and sometimes the right move is to leave the account alone this quarter.
Frequently asked questions
What is the best AI for cold email?
Two tools: a model for the writing and a sender for the delivery. Claude is my pick for the writing (shortest drafts, holds a template as a Skill, cites only real triggers when told to), with ChatGPT a close second. For sending, a purpose-built tool like Instantly, Smartlead, or lemlist handles warmup, rotation, deliverability, and sequencing, which no chat model does.
Is Claude or ChatGPT better for writing cold emails?
Claude by a small margin, measured by how much a rep has to edit before sending. It respects a length rule, avoids the greeting-card register, and holds your template in a Skill. ChatGPT with a well-instructed custom GPT gets close. On both, the length rule belongs in the instructions, because reps send what comes out.
Can ChatGPT or Claude send cold emails?
Claude's Gmail connector can send an email from a chat, and ChatGPT has Gmail plugins, but neither is a cold email sender. Cold outbound at any volume needs warmup, inbox rotation, deliverability monitoring, and sequencing with stop-on-reply. That is what Instantly, Smartlead, and lemlist are for. The model writes; the sender sends.
Are AI SDR tools like 11x or Artisan worth it?
For volume you cannot staff over a wide TAM, yes: they bundle finding, writing, sending, and booking behind one login. The trade is personalization for throughput, so triggers get shallow and review disappears. Use them for the long tail and hold a model-plus-sender stack to a higher standard for the accounts you actually want.
How do I stop AI from making up facts in cold emails?
Constrain the model to the account data. The sentence I use in the Skill is "Use ONLY triggers present in the account data. If none, output NO_TRIGGER," and NO_TRIGGER rows go to a human. Give the model the source URL for each trigger and spot-check ten emails against their rows every run. Invented triggers are a prompt bug, not a model quirk.
How long should an AI-written cold email be?
Under 90 words for a first touch: one true trigger, one line of relevance, one ask. Put the ceiling in the Skill, Project, or GPT instructions rather than the prompt, so the short version is the default. Unedited AI drafts run long by nature, and reps do not edit.
What is a good cold email stack with AI?
Four tools with one job each: Clay (or your data provider) builds the enriched account row, Claude writes the email through a Skill, a sender like Smartlead or Instantly delivers and follows up, and HubSpot records every send so the AE picking up the reply knows what was promised. The full build is the personalized cold email use case on this site.
Should I use Gemini for cold email?
For a reply you are writing inside Gmail right now, yes: the side panel drafts it where you are, off the thread you are reading. For a net-new sequence generated off a list, Claude or ChatGPT with a Project, Skill, or custom GPT is the more mature loop, and either way the sending still belongs to a dedicated sender.
Sources & further reading
- Claude models overview (platform.claude.com)
- Skills in Claude (Claude docs)
- Google Workspace connectors for Claude (Anthropic Help Center)
- Apps and plugins in ChatGPT (OpenAI Help Center)
- Instantly
- Smartlead
- lemlist
Claude ships fast. This page was last reviewed Sep 4, 2026; verify time-sensitive details against the official docs above before relying on them.