The call coaching scorecard
Five dimensions, a written standard for each level, and the evidence rule that makes a score quotable. Rewrite the level definitions from your own won calls, save it as rubric/call-coaching.md, and the agent scores every call against it each Friday, citing the moment per dimension.
- 4: asked about the current process, what it costs them, and who else feels it, with at least one follow-up that went a level deeper
- 3: asked about the current process and its cost; no follow-up past the first answer
- 2: asked what tools or vendors they use, then moved to the demo
- 1: no discovery question, or the buyer had to volunteer their situation
- 4: rep 35-50% on discovery, 50-65% on demo
- 3: within 10 points of the band
- 2: rep above 65% on a discovery call, or below 30% on a demo
- 1: rep above 75% on any external call
- 4: a specific step, person, and date stated on the call, and in the CRM within 24 hours
- 3: a specific step and date on the call; not in the CRM
- 2: 'I'll send some times' or 'let's reconnect', no date
- 1: the call ended with no next step
- 4: named the objection, asked what sat behind it, answered the real one
- 3: answered the stated objection directly and checked whether it landed
- 2: answered with a concession (discount, extra scope) in the next sentence
- 1: talked past the objection or deferred it without a plan
- 4: metrics, economic buyer, decision process, and pain each surfaced with a quote
- 3: three of the four
- 2: pain only
- 1: none surfaced; the call was a feature tour
- every score cites a timestamped moment; a level with no supporting moment defaults to the level below
- check-ins score only talk ratio and next-step clarity
- one practice point per rep per week, written as a behavior on the next call
- calibrate level definitions quarterly against closed deals; tighten any level that 80% of calls land on
Next-step clarity at x2 on every meeting type is the weight I would defend hardest. It is the dimension most teams are lowest on, and the only one that a stage definition and a required field can fix for the whole team at once.
The stack
- Per week
- one run over 100-150 transcripts, several dollars on a mid-tier model; cache the rubric and roster
- Call recorder + CRM + Slack
- free over MCP; no scorecard add-on to buy
- Setup
- a day: the rubric with the manager against ten won calls, the brief, two manual weeks
- The saving
- the 114 calls a week that were coached by nobody, and the manager's listening hour aimed at the calls that need it
The problem
A first-line manager with eight reps sits on top of roughly 120 recorded calls a week. They will listen to three, usually the ones a rep asked them to hear, usually the ones that went well. The coaching that follows is real but random: the rep who talked for 70% of a discovery call on Tuesday gets no note, because nobody heard Tuesday. The call recorder captured all of it. Almost none of it turns into a change in how anyone sells.
The standard fix is a call review form in the CRM. It fails the same way every time. Managers fill it in for two weeks, then a quarter-end arrives, and the form becomes a field with 94% blanks. The scoring is inconsistent across managers even when it gets done, so a 3 from one manager means something different from a 3 from another, and reps learn quickly that the score depends on who reviewed them.
The call recorder's built-in scorecards are better than nothing and worse than they look. They measure what is easy to measure, talk ratio and the number of questions, and they score against a generic notion of a good call, which has little to do with whether your reps ask the three discovery questions that predict a close in your market. And the output is a dashboard, which is to say it waits for someone to open it.
A Claude agent does what a manager with unlimited time would do: listen to every call, score it against the rubric your team wrote, and quote the moment that earned the score. Discovery depth, with the questions that were and were not asked. Talk ratio, from the transcript. Whether the call ended with a next step that has a date, and whether that step showed up in the CRM afterwards. How the pricing objection was handled, in the rep's own words. The rep gets one note a week that names the pattern across their calls and one thing to practice; the manager gets the team view and the two calls worth listening to in full. The scores are the index. The quotes are the coaching.
How it works
- 01 Fire Friday 3pm Claude Code (scheduled)the week's calls per rep
- 02 Read the calls Gong / Fireflies (MCP)transcript, speakers, timestamps
- 03 Read the deals HubSpot (MCP)stage, next step, activities
- 04 Score Clauderubric, quote per dimension
- 05 Find the pattern Claudeone dimension, one practice point
- 06 Note + team view Claudeone page per rep, one for the manager
- 07 Route Slack (MCP)rep gets a DM, manager gets the view
- A scheduled Claude Code job runs every Friday and pulls the week's calls per rep from Gong or Fireflies over MCP, with transcripts, attendees, and talk time
- For each call it reads the linked deal in the CRM: stage, amount, the meeting type, and whether a next step was logged within a day of the call
- It scores the call on your rubric's five dimensions, 1 to 4 each, and records the quote or the absence that earned each score
- Per rep, it looks across the week's calls for the pattern (the dimension that is consistently low, the objection that keeps landing badly) and writes a one-page note: scores, evidence, the pattern, one thing to practice, and the one call worth re-listening to
- It writes the manager a team view: every rep's dimension averages, movement from last week, the team's weakest dimension, and the two calls worth hearing in full
- Each Friday it re-scores, so the notes show whether last week's practice point moved, and the team view shows the trend by dimension
See it run
The playbook
Write the rubric your best calls would actually pass
Start from your own wins. Pull the ten discovery calls from the last two quarters that turned into closed-won and read them with the manager. What did those calls have in common that the losses did not? The answer is usually specific to your market: a question about the current process that surfaces the cost of doing nothing, a named economic buyer by minute twenty, a next step with a date in the calendar before the call ended. Those are your dimensions, and the winning calls give you the definition of a 4.
The rubric below is the one I start from: discovery depth, talk ratio, next-step clarity, objection handling, and MEDDICC coverage, each on a 1-4 scale with a written definition of every level and a rule for what counts as evidence. 'Good discovery' is a feeling; 'asked about the current process, the cost of it, and who else feels it, with a follow-up on at least one' is a thing the agent can find in a transcript and quote.
Decide what the rubric applies to. A discovery call and a pricing call should score differently on discovery depth, so tag the meeting type from the calendar or the CRM and let the rubric weight dimensions per type. A rubric that scores a 15-minute check-in the same as a 45-minute first call produces notes reps stop reading by week three.
TipDefine the 4 before you define the 1. A rubric written from the top down describes what your best call actually sounds like, which is the only version reps will believe.
Connect the call recorder and the CRM, read-only
Connect Gong or Fireflies over MCP and confirm the agent can list the week's calls by rep and read a transcript with speakers and timestamps. Talk ratio comes straight from the speaker segments, so you get it without any recorder feature turned on. The MCP directory covers what each recorder's server exposes.
Connect HubSpot or Salesforce so each call can be tied to its deal and stage, and so the agent can check the one thing the transcript cannot tell you: whether the next step the rep promised on the call was logged in the CRM afterwards. Next-step clarity that lives only in the recording is a 2, not a 4, and the agent should say why.
Keep the agent out of the write path entirely. It reads calls and deals and writes documents and Slack messages. It does not log activities, update fields, or comment in the call recorder. Coaching notes that show up as CRM entries turn into performance records, and the moment reps believe that, the rubric becomes something to game instead of something to learn from.
- Reads: this week's calls per rep: date, attendees, meeting type, speaker segments, transcript
- Reads: the linked deal: stage, amount, meeting type, next-step field and its date, activities logged within 24 hours of the call
- Reads: the rubric file and last week's notes, so the agent can report movement on the practice point
- Writes: one note per rep, one team view, Slack DMs and one channel post. Never a CRM field or a recorder comment
Write the brief so every score arrives with its quote
The brief is the rubric turned into a run, plus the rule that makes the note worth reading: every score cites the moment in the transcript that earned it, with a timestamp, and a score with no quotable evidence defaults to the lower level. A rep who reads 'discovery depth 2/4' will argue. A rep who reads 'discovery depth 2/4: you asked what tools they use (04:12) and moved to the demo; no question about the current process or what it costs them' will nod, because they remember the moment.
Make it look for the pattern before it writes the note. Five calls scored individually is a spreadsheet; the same dimension low on four of five calls is a coaching conversation. The brief asks the agent to find the one dimension that is consistently below the rep's others, name the one thing to practice next week in a sentence, and pick the single call worth re-listening to, with the timestamp to start from. One thing. A note with six action items is a note nobody acts on.
Tell it to be specific about what good looked like, too. When a rep handled the pricing objection well on Wednesday, the note should quote that, so the practice point on Friday is 'do the Wednesday thing on every call' and the rep has their own words as the model. Praise with a timestamp is coaching; praise without one is noise.
Every Friday at 3:00pm {{TIMEZONE}}, run the call coaching review for the week.
Read from {{CALL_RECORDER}}: every recorded external call this week for each rep in {{TEAM_ROSTER}}. For each: date, meeting type, attendees, speaker segments with timestamps, transcript.
Read from {{CRM}}: the deal linked to each call (stage, amount), the next-step field and its date, and any activity logged within 24 hours after the call.
Score each call against rubric/call-coaching.md. Rules:
- Apply the weights for the call's meeting type (discovery, demo, pricing, check-in).
- Each dimension scores 1-4. Cite the transcript moment (timestamp and a short quote) that earned the score. If no moment supports a level, score the lower level and say what was missing.
- Talk ratio is computed from speaker segments; report the rep's share.
- Next-step clarity is a 4 only if the step was stated with a date on the call AND appears in the CRM within 24 hours.
For each rep, write notes/{{DATE}}/{{REP}}.md, one page:
1. Header: rep, calls this week, average per dimension, change from last week.
2. Scores per call in a table, with the earning quote per dimension.
3. The pattern: the one dimension consistently below the rep's others, with the evidence across calls.
4. One thing to practice next week, in one sentence, written as a behavior on the next call.
5. The one call to re-listen to, with the timestamp to start from, and one moment this week that was a 4, quoted.
Write team/{{DATE}}.md for the manager: dimension averages per rep, week-over-week movement, the team's weakest dimension with three quotes that show it, and the two calls worth hearing in full (one strong, one to coach).
Re-read last week's notes and report, per rep, whether the practice point moved.
Post each rep's note to them as a Slack DM. Post the team view to {{MANAGER_CHANNEL}}. Write nothing to the CRM or the call recorder.
If a call has no transcript or is internal, skip it and list it at the end. Do not score a call you did not read.
TipCap the note at one practice point. The rep who gets one behavior to change next week changes it; the rep who gets six changes none and stops opening the DM.
Run it Friday, DM the rep, brief the manager
Schedule the job for Friday afternoon as a scheduled Claude Code job, so the rep reads their note before the weekend and walks into Monday with one thing to try. The manager's team view lands in the leadership channel at the same time, with the two calls worth hearing in full, so their listening time goes where the agent found something.
Run the first two weeks by hand with the manager and read every note against their gut. Where the agent scored a call low that the manager thought was strong, look at the quote: usually the rubric is missing a way your team legitimately does the thing, and the fix is a line in the rubric. Where the manager names a rep they worry about who scored well, find the dimension the rubric does not have. Both are edits to a file, and the second one is the more valuable, because the recorder's generic scorecard would have missed it too.
Tell the reps what the agent reads and what it does not touch, before the first note arrives. The note is a DM from a rubric the team wrote, it quotes their own calls, and it goes nowhere near their CRM record or their review. Reps who understand that read the note. Reps who suspect it feeds a performance file learn to talk less on calls so their ratio looks good, which is the opposite of the point.
Check whether the practice point moved
The re-score is where the play earns its keep. A rep whose discovery depth averaged 2.1 with 'practice: ask about the current process and what it costs' should, a week later, have that question on the transcript at minute six of most calls, and the dimension should move. The agent reads last week's note, looks for the behavior, and says whether it showed up. If it did not, the manager knows the coaching conversation is overdue while it is one week overdue, instead of one quarter.
The team view's trend by dimension is the manager's view of the whole team and the head of sales's view of the manager. A team that is weak on next-step clarity across every rep has a process problem, and the fix is a stage definition and a required field, which is a systems fix, not eight coaching conversations. The team view makes that visible in a way eight individual notes never would.
Send the outcomes back too. When a deal closes, won or lost, the run notes the average scores of its calls beside the outcome, so by the end of the quarter you can see which dimensions actually predict a close in your market. That is the calibration data for the next step, and the answer is rarely the one the rubric started with.
TipRead the team view for the dimension the whole team is low on. That is a process fix for one person to make, and it is worth more than eight versions of the same coaching conversation.
Calibrate the rubric against closed deals every quarter
Every quarter, lay the average call scores of each closed deal beside its outcome. A good rubric shows daylight: the won deals' calls scored higher on the dimensions that matter and the lost deals' calls did not. Where a dimension shows no difference between won and lost, question its weight. Where the won deals share something the rubric never scored, add it. In one team it was whether the rep named a competitor before the buyer did; nobody would have put that in a rubric on day one.
Calibrate the scale as well as the weights. If 80% of calls score a 3 on objection handling, the definition of 3 is too wide and the note is telling nobody anything. Tighten the level definitions until the scores spread, then re-run last quarter's calls against the new rubric so reps see the change explained rather than experienced.
Then share the calibration with the team. A rubric that has been tuned against two quarters of closed deals, and that the reps watched get right about the deals they had a bad feeling about, is a credibility asset. It turns the Friday DM from a machine's opinion into the team's own definition of a good call, applied to every call for once.
Inside the prompt
The scoring prompt is short, but every line is there for a reason. Here is what each one is doing and why.
- Weights per meeting type"Apply the weights for the call's meeting type"
- A discovery call and a check-in are different jobs. One rubric applied flat produces notes reps stop reading.
- Quote or the lower level"If no moment supports a level, score the lower level and say what was missing"
- The rule that makes the score arguable on the weight and not on the fact. It also stops the agent from being generous.
- Next step needs the CRM"a 4 only if ... appears in the CRM within 24 hours"
- The transcript cannot tell you whether the promise survived the call. The CRM can, and the agent checks.
- One pattern, one practice point"the one dimension consistently below ... one thing to practice next week, in one sentence"
- Five scored calls is a spreadsheet. One pattern with one behavior to change is coaching.
- Quote the 4"one moment this week that was a 4, quoted"
- The rep's own good call is the model for next week. Praise with a timestamp teaches; praise without one is noise.
- Read only"Write nothing to the CRM or the call recorder"
- Notes are documents and DMs. The day they become fields, the rubric becomes something to game.
What you get
One rep's Friday note as it lands in their Slack. Every score carries the moment that earned it, the pattern is named across the week, and there is exactly one thing to practice.
CALL COACHING · week of Sep 8 · Maya Chen · 14 calls scored
AVERAGES discovery 2.4 (+0.3) · talk ratio 3.1 (=) · next step 2.1 (-0.2) · objections 3.3 (+0.1) · MEDDICC 2.6 (=)
THE PATTERN
Next-step clarity is your low dimension for the third week. On 9 of 14 calls the call ended with 'I'll send over some times' or 'let's reconnect next week' and no date on the call. On 6 of those the next step never reached the CRM. Contrast Tue 10:00 (Northwind, discovery): 'Can we hold Thursday at 2 for the technical walkthrough with Sam?' (38:40). The deal moved to Evaluation the same day.
ONE THING TO PRACTICE
Before you say goodbye, name the day and the person: 'Thursday at 2 with Sam' instead of 'I'll send times'. Try it on every call this week, including the check-ins.
THE CALL TO RE-LISTEN TO
Wed 14:00, Harlow Logistics (pricing). Start at 22:10. The buyer said 'that's more than we budgeted' and you moved to a discount in the next sentence. On Thu 11:00 (Bexley) you handled the same line with 'what did you budget against?' (19:05) and it turned into a scoping conversation. Do the Thursday thing.
A 4 THIS WEEK
Discovery depth, Mon 15:00 (Corvid): 'Walk me through what happens today when a lead comes in at 6pm' (06:12), then 'and who notices when that breaks?' (08:40). That is the sequence.
Skipped: 2 internal calls, 1 call without a transcript.
- definition4: names the objection, asks what sits behind it, answers the real one
- From the rubric file. Written from the calls that closed, so the rep recognizes the standard as their own team's, not a vendor's.
- evidence'that's more than we budgeted' (22:10) -> discount offered (22:31)
- Timestamp and quote. 'Handled the pricing objection weakly' would be an adjective; this is a moment the rep remembers.
- what was missingno question about the budget, the alternative, or the cost of waiting
- The agent names the absence, so the note explains the 2 instead of asserting it.
- the contrastThu 11:00, Bexley: 'what did you budget against?' (19:05), scored 4
- The same rep did it well the next day. The practice point becomes 'do the Thursday thing', in their own words.
- weightpricing call: objection handling x 2
- Meeting type from the calendar. On a pricing call this dimension carries double weight; on a check-in it barely counts.
- written wherenotes/2026-09-11/maya-chen.md, DM'd; nowhere in the CRM
- Coaching notes are documents. The moment they become records, reps optimize for the score instead of the call.
Pitfalls to avoid
Scores without quotesA 2/4 with no timestamp is an opinion the rep will argue with. Every score cites the moment or defaults to the lower level.
Six action itemsA rep given six things to change next week changes none. One practice point, written as a behavior on the next call, and the one call to re-listen to.
Feeding the CRMThe moment coaching notes become CRM records, reps believe they feed performance reviews and start optimizing for the score. The agent writes documents and DMs, never fields.
One rubric for every meeting typeA check-in scored on discovery depth like a first call produces nonsense notes by week three. Tag meeting types and weight the rubric per type.
Measuring what is easyTalk ratio and question counts are what recorders measure because they can. Your rubric should score the three things your won deals' calls actually had in common.
Never calibratingLay each quarter's closed deals beside their call scores. A dimension that shows no daylight between won and lost is costing you reps' attention every Friday.
Questions people ask
- What is AI sales coaching?
- Using a model to read every recorded sales call, score it against a rubric your team wrote, and turn the scores into a weekly coaching note per rep with quoted evidence and one thing to practice. The version on this page runs as a Claude agent over your call recorder and CRM, and delivers the note to the rep and a team view to the manager. The manager still coaches; the agent does the listening no manager has time for.
- How is this different from Gong's or Fireflies' scorecards?
- Three ways. The rubric is yours, written from your won deals, instead of a generic definition of a good call. Every score cites the transcript moment that earned it, with a timestamp, so the rep can check it. And the output is a note in the rep's Slack every Friday with one practice point, instead of a dashboard that waits to be opened. The recorders' own scoring is a fine starting point for talk ratio; this play uses that data and adds the judgment.
- Will reps accept being scored by an agent?
- They accept it when three things are true: the rubric was written by their team from their own best calls, every score quotes the moment so it can be argued with on the facts, and nothing the agent writes goes into the CRM or a review file. Tell them all three before the first note. Reps who believe the notes feed a performance record will talk less to fix their ratio, which defeats the point.
- What should be in the rubric?
- Whatever your won deals' calls had in common that your lost deals' calls did not. The five dimensions on this page, discovery depth, talk ratio, next-step clarity, objection handling, and MEDDICC coverage, are a starting point most B2B teams recognize. Weight them by meeting type, define every level with a written standard, and calibrate quarterly against closed deals. The dimension you add after a quarter of data is usually the one that matters most.
- Does it need Gong, or does Fireflies work?
- Either. It needs transcripts with speakers and timestamps, which both provide over MCP, and a way to tie a call to a deal, which comes from the CRM. Talk ratio is computed from the speaker segments, so you do not need a paid scorecard feature on either recorder.
- How many calls can it score a week?
- A team of eight reps produces 100-150 external calls a week, and one scheduled run handles that in well under an hour for a few dollars of tokens. Larger teams run per manager. The constraint is never the volume; it is keeping the rubric tight enough that the notes stay short.
- How long does it take to build?
- A day. A morning with the manager writing the rubric against ten won calls, an afternoon connecting the recorder and the CRM and writing the brief, then two Fridays run by hand and read against the manager's gut before it posts on its own.
Related plays
- Post-Call Follow-Up + CRM Note →The per-call twin: the follow-up and the CRM note the moment the call ends.
- Win-Loss Analysis on Evidence →The quarterly view: which of these dimensions actually predicted a close.
- Discovery Call Prep Skill →The prep that raises discovery depth before the call starts.
- Discovery call questions by role →The question sets a 4 on discovery depth tends to draw from.
- Claude + Gong integration guide →Reading transcripts, speakers, and timestamps over MCP.
- Running Claude agents safely →Why the coaching agent writes documents and never touches a CRM field.
- Objection handling scripts by objection →Sixteen objections, each answered in five beats you can say out loud, with a Claude prompt to tailor the script to your product.