Make conversation data do work (L3)
Call recordings are the only genuinely un-copyable data a revenue team owns, and in most companies they sit in a library and rot. Extracting a fixed handful of qualification fields from tools like Gong, Chorus or Avoma with the supporting quote attached gets the forecast populated before the pipeline review starts, letting the meeting be about decisions instead of reading status aloud. Strongest at 200 to 5,000 employees, B2B with a real multi-call sales cycle and a qualification framework already in use.
The steps
- 01
Pick six fields, exactly six
Tool: Salesforce
Choose the six fields that genuinely change a forecast conversation. A workable default: quantified metric, economic buyer, decision criteria, decision process, identified pain, competitive alternative. For each, write what a good answer looks like and what a bad one looks like; reps and models both need the bad example. Six is the cap because every field added is another thing a rep has to confirm, and confirm rate is what makes or breaks this. • Owner: Sales leadership plus RevOps • Tool options: your existing qualification framework, MEDDPICC, SPICED, BANT, whatever is already agreed • Pitfall or what breaks: extracting 30 fields because the model can, so reps stop confirming and the data goes stale. • Definition of done: six fields exist in Salesforce, HubSpot, Dynamics, Pipedrive or Zoho with written definitions plus one good and one bad example each.
- 02
Grade 40 historical calls by hand to build ground truth
Tool: Gong
Use 20 calls from won deals and 20 from lost. Managers fill all six fields manually from the recordings, and note the timestamp where they found each answer. This is simultaneously your accuracy benchmark and your example set. Without it you have no way to know whether the model is right, only whether it sounds right. • Owner: two front-line managers • Tool options: the Gong, Chorus or Avoma library plus a spreadsheet • Pitfall or what breaks: skipping this step, leaving no way to verify whether the model's extraction is correct. • Definition of done: 40 graded calls exist with all six fields populated by a human and a timestamp on each.
- 03
Write the extraction prompt and measure it against ground truth
Tool: Claude
The prompt takes a transcript and returns, per field: the value, a confidence score, and the verbatim quote it drew from. The quote requirement is non-negotiable; it makes every extraction auditable in two seconds instead of requiring someone to rewatch a call. Add one hard rule: return "not discussed" rather than inferring; an empty economic-buyer field is a coaching signal, a guessed one is a corrupted forecast. Run it against all 40 graded calls and measure agreement field by field, not in aggregate, because aggregate accuracy hides the one field that is always wrong. • Owner: Enablement plus a GTM engineer • Tool options: Claude or GPT via API, or Gong's native AI if it will write to your fields • Pitfall or what breaks: letting the model infer instead of quoting, so empty fields get filled with plausible fiction and the coaching signal disappears. • Definition of done: agreement hits 85%+ on each individual field, or the failing field is cut until the prompt improves.
- 04
Write to CRM as a draft, never straight to truth
Tool: n8n
Extraction lands in a staging field or a task, not the live field. The rep sees the proposed value plus the supporting quote and confirms or corrects with one click. Build the correction path to be as fast as the confirm path, because corrections are your best training data. Reps should be the source of excellence in relationships, not of data entry; they still own the judgment call, which is why the click stays. • Owner: GTM engineer • Tool options: the native integration, or Zapier, Make or n8n into your CRM, with the notification landing in Slack or Teams • Pitfall or what breaks: writing directly into the live CRM field, removing the human confirm step. • Definition of done: 50 calls have flowed through, one-click confirm and one-click correct both work, and confirm rate is being tracked per rep.
- 05
Change the pipeline review agenda, in writing
Tool: Slack
Fields are populated before the meeting starts, so no minute goes to reading status aloud. Rewrite the agenda in the invite itself: the three deals with the lowest field confidence, the three with an empty economic buyer, and what happens this week on each. Then hold the line for four weeks, because the old habit returns instantly if you do not. • Owner: VP Sales • Tool options: the recurring meeting invite • Pitfall or what breaks: letting the meeting slide back into reading status aloud within the first few weeks. • Definition of done: four consecutive reviews have run on the new agenda with zero status recitation.
- 06
Feed one coaching topic per rep per month
Tool: Looker
Monthly, per rep, show which of the six fields they most often fail to surface on calls. That is the coaching topic: one field, one month. Frame it as a topic, not a scorecard, and never put it in a comp conversation. The moment it becomes a metric reps are graded on, they start telling the model what it wants to hear. • Owner: Enablement plus managers • Tool options: extraction output aggregated by rep, charted in Looker, Tableau or Power BI • Pitfall or what breaks: turning this into a rep compliance metric, which inverts data quality. • Definition of done: every rep has had one coaching conversation tied to one named field.
- 07
Instrument the result
Track percentage of open opportunities with all six fields populated and confirmed, then forecast accuracy at 30 days out, quarter over quarter. • Where it breaks: extracting 30 fields because the model can, causing reps to stop confirming and data to go stale; the model infers instead of quoting, filling empty fields with plausible fiction; or it becomes a rep compliance metric and data quality inverts. • Visual guidance: a linear flow: call, transcript, six extracted fields with a quote chip under each, one-click confirm, forecast. Put the confirm step in a different color and label it "the human gate."
Tools in this playbook
Next playbooks
Unfamiliar terms are defined in the AI and Revenue Dictionary. Related frameworks live in the framework library.
