GTM Governance GTM GovernanceROI

How to measure the ROI of Claude in GTM (a model you can defend)

To measure AI ROI in a GTM org, keep three separate ledgers instead of one blended number: a time ledger (minutes per play before and after, timed on real instances), an output ledger (more of the work that touches pipeline, like accounts researched and same-day follow-ups), and a pipeline ledger (meetings, conversion, and cycle time for a pilot cohort against a holdout). Divide by an honest cost that includes seats, tokens, and the builder's time. Then halve the pipeline number for the attribution you cannot prove. Most AI ROI decks skip the third ledger and the honest denominator, which is why finance does not believe them.

Overview

The dashboard I am least proud of said 1,400 hours saved last quarter. It had a big green number, a trend line, and a per-rep breakdown, and the CFO looked at it for about four seconds before asking two questions: which hire did we not make, and where is the pipeline? I did not have either answer. The number was self-reported minutes per task multiplied by usage counts multiplied by a loaded hourly rate. It was arithmetic, not evidence.

That meeting is why this guide exists. Measuring the return on Claude in a GTM org is doable, and the model fits in a spreadsheet, but it only works if you separate the things you can measure from the things you are guessing, and price the whole thing against an honest cost. This is the model I run now, with a worked example a skeptical head of sales has actually sat through.

Most AI ROI decks are a vanity ledger

Most AI ROI decks are a vanity ledger

The standard AI ROI slide goes like this: ask users how long a task used to take, ask how long it takes now, multiply the difference by how many times they did it, multiply by a fully loaded hourly rate, and present the total as savings. Every input is a self-report, every multiplication inflates the error, and the output is a number nobody can spend. I call it the vanity ledger, and I built one, so this is not a lecture from a clean seat.

The tell is that the number is always huge and never changes a decision. 1,400 hours is 0.7 of a full-time person. Nobody stopped hiring. Nobody found a rep with 30% more calendar. The time went into more of the same, into slack, into slightly better work, all of which may be good and none of which is on the slide. Finance is right to nod and move on.

The fix is structural rather than cosmetic. Stop producing one blended number and produce three ledgers that each measure one kind of thing with the method that fits it. Some of what Claude does for a GTM team is measurable to the minute. Some is measurable to the deal. Some is a judgment call, and it is fine to label it one.

The three ledgers: time, output, pipeline

The three ledgers: time, output, pipeline

The time ledger asks: for a specific play, how many minutes did it take before and after, measured on real instances. The output ledger asks: are we doing more of the work that plausibly moves pipeline, accounts researched, follow-ups sent the same day, inbound leads routed inside an hour. The pipeline ledger asks the only question finance actually banks: did the reps running these plays book more meetings, convert better, or close faster than the reps who did not.

Keeping them apart is the whole trick. When you blend them, a stopwatch-grade time number and a testimonial-grade pipeline claim end up in the same cell, and the weakest input sets the credibility of the total. Presented as three ledgers, each with its own confidence label, the same data reads as a case instead of a pitch. The CFO can argue with one ledger without dismissing the others, which is exactly what you want.

A useful rule for which ledger a claim belongs in: if you measured it with a stopwatch, it is time; if you counted it in a system, it is output; if you compared a cohort against a holdout in the CRM, it is pipeline. If you asked someone how they felt about it, it is a quote, and quotes go in the appendix.

Three ledgers, kept apart, over one honest cost
01Time ledgerstopwatch on real instances, then discounted
minutes before / afterplays per monthrecapture rate
02Output ledgercounted in a system, cohort vs rest
accounts researchedsame-day follow-upsleads routed in an hour
03Pipeline ledgercohort vs frozen holdout, halved
meetings per repconversioncycle time
04Honest denominatorwhat it actually costs
seatstokensenablementbuilder's time
Time tells you which plays deserve plumbing, output tells you the time became work, and pipeline is the only ledger finance banks. Blend them and the weakest input sets the credibility of the total.
The time ledger: measured, then discounted

The time ledger: measured, then discounted

Time is the easiest ledger to get right and the easiest to fake. Get it right by timing ten real instances of the play before Claude and ten after, with a stopwatch, on the reps' actual accounts. For a pre-call brief at a team I ran this on, the before was 38 to 55 minutes and the after was 8 to 14, so I carry 45 and 10 and refuse to quote the best case. Then multiply by the play count from the usage analytics dashboard (skills usage and connector reads, exported to CSV), never by a survey of how often people think they did it.

Then discount it, out loud. Saved time is spent on something, and you cannot show me what. I use a recapture rate of 50%: half the saved minutes turn into more calls, more accounts, or better prep, and half evaporate into the day. The number is a judgment, so label it one, and let the head of sales argue it up or down. A discounted time number people believe beats an undiscounted one they smile at.

What the time ledger is for: it tells you which plays are worth the plumbing and whether a rep's day actually changed. What it is not for: converting into dollars and calling it savings. The moment you multiply minutes by a loaded salary, you are back in the vanity ledger, and the CFO's first question will be which hire you did not make.

💡

TipWrite the timing protocol down before the pilot: who times, how many instances, which accounts. Ten stopwatch numbers collected on purpose beat a hundred survey answers collected afterwards.

The output ledger: more of what touches pipeline

The output ledger: more of what touches pipeline

Recaptured time should show up as more of something, and the output ledger is where you check. Pick two or three counts that live in a system and plausibly lead to pipeline: accounts researched per rep per week, follow-ups sent within 24 hours of a call as a percentage, inbound leads routed inside an hour, personalized first lines per sequence. These are counted in the CRM or the sequencer, so nobody has to estimate anything.

The output ledger is also where you catch the play that saves time and changes nothing. If briefs went from 45 minutes to 10 and the number of accounts touched per week is flat, the time went somewhere else, and that is worth knowing before you claim it. In one rollout the recaptured prep time turned into same-day follow-ups going from about 40% of calls to about 75%, which is a real, countable, pipeline-adjacent change. That is the sentence I lead with now, not the hours.

Keep this ledger short. Three counts, tracked weekly, cohort versus rest of team. More than that and you are building a dashboard again, and the dashboard is where I got into trouble.

The pipeline ledger: the only one finance banks

The pipeline ledger: the only one finance banks

This is the ledger that answers the CFO's second question, and it needs a design, not a data pull. Take the pilot cohort from the rollout guide, five to eight reps running the plays, and compare them for a quarter against a holdout of similar reps who are not, on meetings booked per rep per month, stage-to-stage conversion, and cycle time. Match the cohorts on segment and tenure as best you can and say where you could not. Same territory mix, same quarter, same quota structure, or the comparison is decoration.

Be honest about what a cohort comparison can carry. With eight reps a quarter you will see a direction and a rough size, and you will not see statistical certainty, so do not dress it up as certainty. Report the delta, the sample size, and the two or three things that could also explain it (a new SDR, a pricing change, a territory reshuffle). A head of sales who has run comps before will trust a number with caveats over a clean one without.

Convert to dollars only here, and conservatively. One extra meeting per rep per month, at your real meeting-to-opportunity rate, your real win rate, and your real average contract value, gives you an expected closed-won value per meeting. It is a small number per meeting and a meaningful one across a team and a year, which is exactly how sales tooling is supposed to pay back.

💡

TipName the holdout before the quarter starts and freeze it. A holdout picked after you have seen the numbers is a highlight reel.

The honest denominator

The honest denominator

Every ROI I have seen understated the cost, and mine did too. The denominator has four lines. Seats, at the current per-seat price for the plan you run, for every seat provisioned, including the ones nobody used yet. Tokens, for anything running on the API or in Claude Code on a schedule: the overnight scoring, the enrichment, the digests, priced from the usage export and governed with the levers in the cost control guide. Enablement hours, the working sessions and office hours, at the loaded rate of the people in them. And the builder's time, which is mine, and which was the biggest line by a distance.

That last one is the line people delete because it is embarrassing. Four hours a week of a systems person building and maintaining skills, connectors, and Projects is real money, and leaving it off makes every ROI look better than it is. Put it in. If the case only works without your salary in it, the case does not work, and it is better to know that in a spreadsheet than in a budget review.

The honest denominator is also your best cost-control instrument, because it shows where spend actually sits. In my model the seats were a rounding error next to my time, which told me the highest-return work was packaging plays so they needed less of me, not shaving seats.

Where the money goes, and the levers
01Model choiceright size per task
the biggest lever
02Prompt cachingreuse the preamble
stop re-paying for context
03Effortreason deep only where it counts
not everywhere
04Batchfor patient work
about half price
Four levers control most of your Claude spend across the app, Code, and API. Pull the ones that fit the workload.
The spreadsheet model, with a worked example

The spreadsheet model, with a worked example

One tab, four blocks. Cost per month: seats, tokens, enablement, builder time. Time ledger: minutes before and after per play, plays per month from the analytics export, a recapture rate. Output ledger: two or three counts, cohort versus holdout. Pipeline ledger: meetings per rep per month for cohort and holdout, meeting-to-opportunity rate, win rate, ACV, and an attribution haircut. Every input has a cell you can point at and a source you can name.

Here is a twelve-rep team, with numbers I would defend on a call and you should replace with your own. Costs: twelve standard Team seats at $25 a seat a month billed monthly ($20 on an annual plan, per the current price list), about $300; API tokens for the routines, about $150; enablement, 20 hours in the quarter, about $600 a month spread; my time, four hours a week at a $100 loaded rate, about $1,700. Total, around $2,750 a month, or about $33,000 a year. Time ledger: the pre-call brief at 45 minutes before and 10 after, eight calls a rep a week, gives about 4.7 hours a rep a week, or about 56 hours a week across twelve reps. At a 50% recapture rate, 28 hours a week of real, redeployed time. Quoted, not banked.

Pipeline ledger: the cohort booked one more meeting per rep per month than the holdout. At a 25% meeting-to-opportunity rate, a 20% win rate, and a $40,000 ACV, each meeting is worth about $2,000 in expected closed-won. Twelve reps, twelve months, 144 meetings, about $288,000 expected. Now the haircut: halve it for the attribution you cannot prove, $144,000, against $33,000 of honest cost. That clears with room. If your version only clears at full credit, you do not have a case yet; you have another quarter of measurement to do.

  • Cost: 12 seats (~$300) + tokens (~$150) + enablement (~$600) + builder time (~$1,700) = ~$2,750 a month, ~$33k a year.
  • Time: 45 -> 10 minutes per brief, 8 calls a rep a week, 12 reps = ~56 hours a week; 50% recapture = ~28 hours a week redeployed. Quote it, do not bank it.
  • Output: same-day follow-ups from ~40% to ~75% of calls in the cohort; accounts researched per rep per week up from 6 to 11.
  • Pipeline: +1 meeting per rep per month x 25% to opportunity x 20% win x $40k ACV = ~$2k per meeting; 144 meetings = ~$288k expected.
  • Haircut: halve it. ~$144k against ~$33k. The case clears without the vanity math.
Where it goes wrong

Where it goes wrong

The failures cluster around the two things the vanity ledger lets you skip: measurement design and cost honesty. Surveying time instead of timing it. Multiplying minutes by salary and calling it savings. Skipping the holdout, so the pipeline number is a before-and-after on a team whose quarter changed for ten other reasons. Deleting the builder's time from the denominator. Presenting one blended number so a soft input sets the credibility of the whole thing.

Then the subtler one: measuring the tool instead of the plays. Claude across a GTM org is a dozen different jobs with different returns, and an org-level ROI hides the two plays paying for everything and the five that are hobbies. Run the model per play. It is more work and it is the version that tells you what to build next, which is the point of measuring anything.

  • Time from surveys instead of a stopwatch on real instances.
  • Minutes multiplied by a loaded salary and presented as savings.
  • No holdout, so the pipeline delta is a story about the quarter, not the plays.
  • The builder's time and enablement hours left out of the cost.
  • One blended number instead of three labeled ledgers.
  • ROI at the tool level, which hides which plays actually pay.
The GTM version

The GTM version

For a RevOps lead who has to defend the line item: three ledgers, an honest denominator, one tab, and a haircut you apply before anyone asks. Time tells you which plays deserve plumbing. Output tells you whether the recaptured time became work. Pipeline, cohort against holdout, is what finance banks, and only after you halve it. Run it per play, and the model doubles as your roadmap, because the plays that clear are the ones the rollout guide should expand next.

The 1,400-hours dashboard is gone. The replacement is one page with three small tables and a sentence at the top that says which plays cleared their cost and by how much, with the caveats attached. It is less impressive and it has never been argued with, which turns out to be the better outcome. Which of your Claude plays would clear its cost at half credit, and which one are you quietly carrying because the demo was good?

How to set it up

How to set it up

Pick the two or three plays you will measure

Measure plays, never the tool. Start with the plays already running in a pilot cohort, the pre-call brief and the post-call follow-up are typical, and write down for each one the time number, the output count, and the pipeline metric you expect it to move. If you cannot name a pipeline metric for a play, it lives in the time and output ledgers only, and you say so.

Time the before, on real instances

Before the cohort starts, time ten real instances of each play with a stopwatch on the reps' actual accounts. Record the range and carry the median. Repeat in week three of the pilot. This is the only time data you will ever need, and it cannot be reconstructed afterwards.

💡

TipAsk the frontline manager to do the timing rather than the reps. Self-timed numbers drift toward whatever the rep thinks you want.

Turn on usage analytics and export the CSV monthly

As a Team plan owner or an Enterprise owner or admin, open the usage analytics dashboard, confirm skills usage and connector activity are showing for the cohort, and export the spend CSV monthly for the token line of the denominator. Play counts come from skills and connector reads plus a log in the play's destination, not from anyone's memory.

Freeze a holdout and run the quarter

Name the cohort and a matched holdout before the quarter begins and do not change either. Track meetings booked per rep per month, stage conversion, and cycle time for both in the CRM. Note anything else that changed for either group during the quarter, because it goes on the slide next to the delta.

Fill the sheet and apply the haircut before you present

Four blocks on one tab: cost, time, output, pipeline. Have Claude Code do the joining so the numbers are reproducible next quarter: point it at the analytics export and your play log, and ask for the per-play table.

zsh
$# join the Claude usage export with the CRM cohort pull, per play
$claude -p "Read claude-usage-aug.csv and cohort-meetings-q3.csv. For each play (skill name), give plays per rep per week, token cost, and meetings per rep per month for cohort vs holdout. Apply a 50% recapture rate to time and a 50% haircut to pipeline. Output a markdown table."
| Play | Plays/rep/wk | Tokens/mo | Mtgs/rep/mo (cohort) | Mtgs/rep/mo (holdout) | Pipeline (haircut) |
| account-brief | 7.4 | $61 | 6.1 | 5.0 | ~$132k/yr |
| call-followup | 5.9 | $48 | (shared cohort) | | ~$12k/yr (output only) |
$
💡

TipPresent the haircut number as the headline and the full-credit number in a footnote, not the other way round. The room relaxes when you discount yourself before they do.

FAQ

Frequently asked questions

How do you measure the ROI of AI in a sales team?

Keep three ledgers apart: time (minutes per play before and after, timed on real instances), output (counts in the CRM like same-day follow-ups and accounts researched), and pipeline (meetings, conversion, and cycle time for a pilot cohort against a matched holdout). Divide by an honest cost that includes seats, tokens, enablement, and the builder's time, then halve the pipeline number for attribution you cannot prove.

Why is 'hours saved' a bad ROI metric?

Because saved time is spent on something you cannot see, and multiplying self-reported minutes by a loaded salary produces a number nobody can bank. Time a play with a stopwatch, discount it with a recapture rate, and quote it as evidence that the day changed, not as dollars.

What should the cost side include?

Seats for every seat provisioned, API and Claude Code tokens for anything running on a schedule, enablement hours at a loaded rate, and the time of whoever builds and maintains the skills, connectors, and Projects. The builder's time is usually the biggest line and the one people delete.

How do I attribute pipeline to Claude honestly?

Run a cohort against a frozen holdout for a quarter, matched on segment and tenure, and report the delta with the sample size and the other things that changed. Convert to dollars using your real meeting-to-opportunity rate, win rate, and ACV, then halve the result before presenting it.

What does a believable ROI number look like?

One that clears the honest cost at half credit. In the worked example, a twelve-rep team with about $33,000 a year of true cost and one extra meeting per rep per month lands around $144,000 in expected closed-won after the haircut. If your case only works at full credit, keep measuring.

Where do the usage numbers come from?

The Team and Enterprise usage analytics dashboard shows active members, skills usage, connector reads and writes, and spend by model, and exports a per-user, per-model CSV. Use it for play counts and the token line. It is available to Team owners and to Enterprise owners and admins.

Should I measure ROI for Claude as a whole or per play?

Per play. An org-level number hides the two plays paying for everything and the five that are hobbies. Per-play ROI also tells you what to build and expand next, which is the actual reason to measure.

How often should I report it?

The three ledgers weekly inside the rollout, on one page. The dollar case quarterly, once a cohort has run against a holdout for a full quarter. Reporting dollars monthly on eight reps is noise dressed as precision.

Sources

Sources & further reading

Claude ships fast. This page was last reviewed Sep 4, 2026; verify time-sensitive details against the official docs above before relying on them.

Get the AI-for-GTM playbook in your inbox

New Claude guides, use cases, and prompts every couple of weeks.

Subscribe →