Make conversation data do work (L3)

Call recordings are the only genuinely un-copyable data a revenue team owns, and in most companies they sit in a library and rot. Extracting a fixed handful of qualification fields from tools like Gong, Chorus or Avoma with the supporting quote attached gets the forecast populated before the pipeline review starts, letting the meeting be about decisions instead of reading status aloud. Strongest at 200 to 5,000 employees, B2B with a real multi-call sales cycle and a qualification framework already in use.

WORKFLOW1Pick six fields, exactly …ixSalesforce2Grade 40 historical calls…by hand to build ground truthGong3Write the extraction prom…t and measure it against grounClaude4Write to CRM as a draft, …ever straight to truthn8n5Change the pipeline revie… agenda, in writingSlack6Feed one coaching topic p…r rep per monthLooker7Instrument the resultManual
7 steps, in order, with the tool that owns each one.
Adoption ladderSix levels from Starter to Rebuilt. This item sits at level 3.L1 StarterOne tool, no workflow changeL2 AssistedAI drafts, humans approveL3 IntegratedWired into CRM and SlackL4 OrchestratedMulti-step, owned, measuredL5 AutonomousAgent runs, human auditsL6 RebuiltThe process itself changes
This playbook belongs at L3 Integrated. Running it above your level is how pilots stall.
Measures of success% opportunities with six fields confirmed; Forecast accuracy at 30 days; Field-level agreement rate; Confirm rate per repPROVE IT WORKED% opportunities with six fieldsconfirmedForecast accuracy at 30 daysField-level agreement rateConfirm rate per rep

The steps

  1. 01

    Pick six fields, exactly six

    Tool: Salesforce

    Choose the six fields that genuinely change a forecast conversation. A workable default: quantified metric, economic buyer, decision criteria, decision process, identified pain, competitive alternative. For each, write what a good answer looks like and what a bad one looks like; reps and models both need the bad example. Six is the cap because every field added is another thing a rep has to confirm, and confirm rate is what makes or breaks this. • Owner: Sales leadership plus RevOps • Tool options: your existing qualification framework, MEDDPICC, SPICED, BANT, whatever is already agreed • Pitfall or what breaks: extracting 30 fields because the model can, so reps stop confirming and the data goes stale. • Definition of done: six fields exist in Salesforce, HubSpot, Dynamics, Pipedrive or Zoho with written definitions plus one good and one bad example each.

  2. 02

    Grade 40 historical calls by hand to build ground truth

    Tool: Gong

    Use 20 calls from won deals and 20 from lost. Managers fill all six fields manually from the recordings, and note the timestamp where they found each answer. This is simultaneously your accuracy benchmark and your example set. Without it you have no way to know whether the model is right, only whether it sounds right. • Owner: two front-line managers • Tool options: the Gong, Chorus or Avoma library plus a spreadsheet • Pitfall or what breaks: skipping this step, leaving no way to verify whether the model's extraction is correct. • Definition of done: 40 graded calls exist with all six fields populated by a human and a timestamp on each.

  3. 03

    Write the extraction prompt and measure it against ground truth

    Tool: Claude

    The prompt takes a transcript and returns, per field: the value, a confidence score, and the verbatim quote it drew from. The quote requirement is non-negotiable; it makes every extraction auditable in two seconds instead of requiring someone to rewatch a call. Add one hard rule: return "not discussed" rather than inferring; an empty economic-buyer field is a coaching signal, a guessed one is a corrupted forecast. Run it against all 40 graded calls and measure agreement field by field, not in aggregate, because aggregate accuracy hides the one field that is always wrong. • Owner: Enablement plus a GTM engineer • Tool options: Claude or GPT via API, or Gong's native AI if it will write to your fields • Pitfall or what breaks: letting the model infer instead of quoting, so empty fields get filled with plausible fiction and the coaching signal disappears. • Definition of done: agreement hits 85%+ on each individual field, or the failing field is cut until the prompt improves.

  4. 04

    Write to CRM as a draft, never straight to truth

    Tool: n8n

    Extraction lands in a staging field or a task, not the live field. The rep sees the proposed value plus the supporting quote and confirms or corrects with one click. Build the correction path to be as fast as the confirm path, because corrections are your best training data. Reps should be the source of excellence in relationships, not of data entry; they still own the judgment call, which is why the click stays. • Owner: GTM engineer • Tool options: the native integration, or Zapier, Make or n8n into your CRM, with the notification landing in Slack or Teams • Pitfall or what breaks: writing directly into the live CRM field, removing the human confirm step. • Definition of done: 50 calls have flowed through, one-click confirm and one-click correct both work, and confirm rate is being tracked per rep.

  5. 05

    Change the pipeline review agenda, in writing

    Tool: Slack

    Fields are populated before the meeting starts, so no minute goes to reading status aloud. Rewrite the agenda in the invite itself: the three deals with the lowest field confidence, the three with an empty economic buyer, and what happens this week on each. Then hold the line for four weeks, because the old habit returns instantly if you do not. • Owner: VP Sales • Tool options: the recurring meeting invite • Pitfall or what breaks: letting the meeting slide back into reading status aloud within the first few weeks. • Definition of done: four consecutive reviews have run on the new agenda with zero status recitation.

  6. 06

    Feed one coaching topic per rep per month

    Tool: Looker

    Monthly, per rep, show which of the six fields they most often fail to surface on calls. That is the coaching topic: one field, one month. Frame it as a topic, not a scorecard, and never put it in a comp conversation. The moment it becomes a metric reps are graded on, they start telling the model what it wants to hear. • Owner: Enablement plus managers • Tool options: extraction output aggregated by rep, charted in Looker, Tableau or Power BI • Pitfall or what breaks: turning this into a rep compliance metric, which inverts data quality. • Definition of done: every rep has had one coaching conversation tied to one named field.

  7. 07

    Instrument the result

    Track percentage of open opportunities with all six fields populated and confirmed, then forecast accuracy at 30 days out, quarter over quarter. • Where it breaks: extracting 30 fields because the model can, causing reps to stop confirming and data to go stale; the model infers instead of quoting, filling empty fields with plausible fiction; or it becomes a rep compliance metric and data quality inverts. • Visual guidance: a linear flow: call, transcript, six extracted fields with a quote chip under each, one-click confirm, forecast. Put the confirm step in a different color and label it "the human gate."

Tools in this playbook

Next playbooks

Unfamiliar terms are defined in the AI and Revenue Dictionary. Related frameworks live in the framework library.

Share this playbook

Posting to Instagram or TikTok? Copy the link, it carries the title, summary and share image.

Arrives weekly by email. Free. Unsubscribe anytime. By subscribing you agree to our Privacy policy and Terms. We never sell or share the list.