---
name: scoring-sales-calls
description: >-
  Scores a recorded or transcribed sales call against a weighted behavioral
  rubric using Gong Labs conversation benchmarks, then writes coach-ready
  feedback naming exactly one focus behavior for the next call. Use when the
  user says "score this call", "review this call recording", "call coaching",
  "coaching scorecard", "call rubric", "grade this discovery call", "here is
  the transcript of my demo", "talk to listen ratio", "how did this rep do",
  "objection handling feedback", or pastes a call transcript, Gong or Chorus
  link, or a call summary and asks what the rep should improve. Use this
  skill whenever the task involves evaluating seller behavior on a specific
  recorded conversation, even if the user does not say "score" or "rubric".
  Do NOT use for designing a new hire ramp or certification program (see
  building-onboarding-ramp-plans), for aggregating closed-deal buyer
  interviews (see running-win-loss-analysis), for forecast or pipeline
  reviews, or for writing call summaries and CRM notes.
metadata:
  version: "1.0"
---

# Scoring sales calls

Score one call against a weighted behavioral rubric and return a coaching brief with a single focus behavior. Deal strategy, forecast judgment, CRM note writing, and rep performance management across a quarter are out of scope.

## What you need

A transcript or a detailed call summary, the call type (discovery, demo, negotiation), the rep's tenure, and the deal stage. If talk-time percentages are not supplied and you have a transcript, compute seller talk share from speaker word counts and label it estimated, because a rubric line scored on an unstated guess is not coachable.

## Benchmarks the rubric is anchored to

Score against these observed values, not intuition. All are correlational findings from vendor-analyzed call corpora, so treat a miss as a flag for a conversation, not proof of failure.

| Behavior | Benchmark | Source |
|---|---|---|
| Seller talk time, closed-won deals | 57% | [Gong, 326,000 calls](https://www.gong.io/blog/talk-to-listen-conversion-ratio) |
| Seller talk time, lost deals | 62% | [Gong](https://www.gong.io/blog/talk-to-listen-conversion-ratio) |
| Aspirational "golden ratio" (2016 study) | 43% talking / 57% listening | [Gong](https://www.gong.io/blog/talk-to-listen-conversion-ratio) |
| Talk share associated with lower conversion and win rates | above 65% | [Gong](https://www.gong.io/blog/talk-to-listen-conversion-ratio) |
| Questions asked by sellers who won | 15–16 per call | [Gong](https://www.gong.io/blog/talk-to-listen-conversion-ratio) |
| Questions asked by sellers who lost | ~20 per call | [Gong](https://www.gong.io/blog/talk-to-listen-conversion-ratio) |
| Targeted questions on a discovery call | 11–14 | [Gong, 519,000 calls](https://www.gong.io/blog/nailing-your-sales-discovery-calls) |
| Buyer contacts, successful vs unsuccessful deals | 2x as many | [Gong, 1.8M opportunities](https://www.gong.io/blog/the-best-sales-insights-of-2025) |
| Win-rate lift from multithreading on deals over $50K | +130% | [Gong](https://www.gong.io/blog/the-best-sales-insights-of-2025) |
| Win-rate lift when an enterprise rep brings a sales engineer to the technical demo | up to +30% | [Gong](https://www.gong.io/blog/the-best-sales-insights-of-2025) |

Do not conflate 43:57 with 57%. Cite 43:57 as the aspirational target from the 2016 study and 57% as the observed won-deal average in the 2025 refresh ([Gong](https://www.gong.io/blog/talk-to-listen-conversion-ratio)).

## Rubric weights

Use these dimension weights unless the user supplies their own ([MuchBetter.ai 25-point rubric](https://muchbetter.ai/blog/sales-call-coaching-scorecard-a-25-point-rubric-for-managers)):

| Dimension | Weight | Observable criteria |
|---|---|---|
| Objection handling | 30% | Pauses after the objection to let the buyer speak; reframes against stated urgency; does not answer a concern the buyer did not raise |
| Value communication | 25% | Ties capability to the buyer's stated "why"; quantifies impact in the buyer's units |
| Closing | 25% | Quality of next step; builds value in the next step; positions it as natural progression rather than a sales tactic; confirms resolution before ending |
| Opening and qualification | 15% | Credibility built in the first 30 seconds; identifies all relevant buyers; validates decision criteria |
| Process adherence | 5% | Methodology fields captured; agenda set and confirmed |

Score each dimension on a 1–5 band and define the band before assigning it. Each component needs three anchors — exceeding, meeting, below (same source). Worked anchor for talk-to-listen: *exceeding* = asks thoughtful questions that uncover buying motivations; *meeting* = asks qualifying questions and listens actively; *below* = dominates the conversation without seeking customer input (same source).

## Workflow

Copy this checklist into your reply and tick items as you go:

```
- [ ] 1. Classify the call and pull the measurable counts
- [ ] 2. Score all five dimensions with evidence quotes
- [ ] 3. Compute the weighted score
- [ ] 4. Run the consistency vs reactivity read
- [ ] 5. Select exactly one focus behavior
- [ ] 6. Validate against the checks, fix, re-validate
- [ ] 7. Emit the coaching brief
```

**1. Pull the counts.** Extract seller talk share, total seller questions, question distribution across call thirds, number of buyer participants, and longest uninterrupted seller monologue. Front-loaded questions are a distinct failure from too few questions: top performers distribute questions evenly and average reps work a checklist at the top of the call ([Gong](https://www.gong.io/blog/nailing-your-sales-discovery-calls)). Report counts before scores, because the score is arguable and the counts are not.

**2. Score each dimension.** Every score needs a verbatim quote or a timestamped moment as evidence. Judgment step: you decide the band. Decision criterion — if you cannot cite the moment that made it a 2 rather than a 4, the score is not defensible and you should widen to a band ("2–3, insufficient evidence"). Never mark a dimension the call had no opportunity to exercise; mark it N/A and redistribute its weight proportionally, because scoring an absent negotiation as 1 punishes the rep for the call type.

**3. Compute the weighted score.** Weighted score = Σ(dimension score × weight), reported on the 1–5 scale to one decimal, with the N/A redistribution stated.

**4. Read consistency versus reactivity.** High performers hold roughly the same talk ratio whether they win or lose; low performers swing 10 points, from 54% on won deals to 64% on lost ones ([Gong](https://www.gong.io/blog/talk-to-listen-conversion-ratio)). If prior calls from the same rep are available, compare talk share across them and say whether the pattern is consistency or reactivity, because a rep who only over-talks under pressure needs a different intervention than one who always over-talks.

**5. Select one focus behavior.** Non-negotiable: one, not three. Pick it by expected impact — the lowest-scoring dimension weighted by its rubric weight, unless a lower-weighted item is a hard blocker (no next step booked, no economic buyer identified). Name the replacement behavior, not the deficiency. Write it as a rehearsable move the rep can run on the next call, drawing on the four active-listening mechanics: the two-second pause before responding; paraphrase and confirm ("So what I'm hearing is that efficiency is your top concern. Does that sound right?"); open-ended phrasing ("How does your team currently handle automation?" rather than "Do you use automation?"); and "That's interesting … tell me more" ([Gong](https://www.gong.io/blog/talk-to-listen-conversion-ratio)).

**6. Validate, fix, re-validate.** Run every check. If one fails, fix it and re-run the whole list; only emit the brief once all pass, because a coaching brief with three focus areas produces zero behavior change.

```
- [ ] Every dimension score has a quote or timestamp attached
- [ ] Counts (talk share, question count, distribution) are stated before scores
- [ ] Dimensions with no opportunity are N/A, not 1, with weight redistributed
- [ ] Exactly one focus behavior, phrased as a move to run, not a flaw to fix
- [ ] Every benchmark comparison names the number and links the source
- [ ] No competency judgment about the rep as a person, only about behaviors on this call
```

**7. Emit the brief.** Use the template below.

## Output format

Use this exact structure. Managers paste it into 1:1 docs and compare week over week, so keep the section order and the single-focus rule; wording inside sections is yours.

```markdown
# Call score — <rep>, <account>, <call type>, <date>

## Measured
Seller talk share: X% (won-deal benchmark 57%) | Questions: N (discovery range 11-14)
Distribution: front-loaded / even / back-loaded | Buyer participants: N
Longest seller monologue: M:SS

## Scores
| Dimension | Weight | Score (1-5) | Evidence |
| Objection handling | 30% | | "<quote>" |
| Value communication | 25% | | |
| Closing | 25% | | |
| Opening and qualification | 15% | | |
| Process adherence | 5% | | |
**Weighted: X.X / 5**

## What worked
<2-3 behaviors to keep, each with the moment it happened>

## Focus behavior for the next call
**<One behavior.>** Trigger: <when to run it>. Script: "<exact words>".
How we will know it worked: <observable signal on the next call>

## Coaching question to open the 1:1
"<one question, e.g. 'Did they create appropriate urgency without damaging trust?'>"
```

## Gotchas

- More questions is not better. Sellers who lost asked ~20 questions per call while winners asked 15–16 ([Gong](https://www.gong.io/blog/talk-to-listen-conversion-ratio)), so a rep at 22 questions has an interrogation problem, not a discovery strength. Score question *targeting and distribution*, and treat a high count as a flag.
- Talk share alone under-diagnoses. High performers generate more buyer-seller interactions at the same talk ratio, and lost deals feature long seller monologues ([Gong](https://www.gong.io/blog/talk-to-listen-conversion-ratio)). A rep at 55% talk share who delivered one seven-minute monologue scores worse on interactivity than a rep at 60% with rapid exchange, so always report the longest monologue next to the ratio.
- These benchmarks are observational correlations from one vendor's customer base, not controlled experiments ([Gong](https://www.gong.io/blog/the-best-sales-insights-of-2025)). Present a miss as "outside the pattern associated with won deals", never as a causal claim, because reps disengage from coaching that overstates its evidence.
- Single-threaded calls are the highest-leverage miss the rubric will not catch on its own. Multithreading lifts win rate by 130% on deals over $50K and successful deals carry 2x the buyer contacts ([Gong](https://www.gong.io/blog/the-best-sales-insights-of-2025)), so when a call has one buyer participant, flag it in the brief regardless of the score.
- For enterprise technical demos, note whether a sales engineer was present: bringing one in for the demo and technical questions is associated with up to +30% win rate ([Gong](https://www.gong.io/blog/the-best-sales-insights-of-2025)). This is a staffing fix, not a rep skill gap, and misfiling it as a coaching item wastes the 1:1.
- The published 25-point rubric gives dimension weights but does not decompose per-item point values ([MuchBetter.ai](https://muchbetter.ai/blog/sales-call-coaching-scorecard-a-25-point-rubric-for-managers)). Do not invent a 25-point breakdown; score 1–5 per dimension and apply the weights.
- When rolling scores up across a team, pilot the rubric with one team for 4–6 weeks before wider rollout, report composites monthly, and adjust goals quarterly (same source). Useful review filters: objection handling below 6/10, talk-listen ratio below 40%, and calls from reps in their first 90 days (same source).
