Framework · Adoption model

The L1 to L6 AI Maturity Ladder

The AI Maturity Ladder is a six-level model The Revenue AI Report uses to place an AI deployment: L1 Starter is one tool with no workflow change, L2 Assisted is AI drafting with human approval, L3 Integrated wires AI into the systems of record, L4 Orchestrated runs multi-step work with a named owner and kill criteria, L5 Autonomous lets an agent act with humans auditing exceptions, and L6 Rebuilt changes the process itself. Each level has its own benefit, its own failure mode, and its own exit condition. Skipping a level is the most common cause of a reversal.

What it is

Most AI adoption arguments are stuck because two people are describing different rungs with the same word. One means a rep using a chat window. The other means an agent emailing customers unsupervised. The ladder gives both a number, so the conversation becomes a placement instead of an opinion.

Score one workflow at a time. The rung tells you the benefit you should expect, the failure you are exposed to, and the condition you have to satisfy before climbing.

The six AI playbook adoption stagesL1 Starter adds one tool. L2 Assisted uses AI drafts with human approval. L3 Integrated connects AI to core systems. L4 Orchestrated assigns ownership and measurement. L5 Autonomous lets an agent run with human audit. L6 Rebuilt changes the process itself.AI ADOPTION MODELFrom useful tool to rebuilt workflowL1StarterOne tool, noworkflow changeSTART HEREL2AssistedAI drafts, humansapproveNEXT RUNGL3IntegratedWired into CRM andSlackNEXT RUNGL4OrchestratedMulti-step, owned,measuredNEXT RUNGL5AutonomousAgent runs, humanauditsNEXT RUNGL6RebuiltThe process itselfchangesNEW OPERATING MODEL
The stage describes how work changes, not how advanced the tool is.
Leverage versus risk across the six AI maturity levelsTwo lines plotted from L1 to L6. Leverage rises steadily from low at L1 Starter to highest at L6 Rebuilt. Risk rises sharply from L3 Integrated through L5 Autonomous, where it peaks, then falls slightly at L6.THE CLIMBWhere the risk outruns the leverageCONTRACT AND CONTROL ZONE0255075100L1StarterL2AssistedL3IntegratedL4OrchestratedL5AutonomousL6RebuiltLeverageRisk
Leverage climbs the whole way. Risk climbs faster between L3 and L5, then flattens once the operating model is redesigned and owned. That crossing zone is where contracts, spend caps, and kill switches earn their keep.

The six levels

Each level carries the tech normally in play, what you gain, what it costs you, and the gate you have to clear to move up. Tools are tagged as examples of the rung, not as recommendations.

L1

Starter

One tool, one person, no workflow change.

The question at this rung: Is anyone getting value from this at all?

A rep pastes a call transcript into a chat window and gets a follow-up email back. Nothing in the CRM changes. Nobody else sees the output. The tool is often bought on a personal card or used on a free tier.

Tech normally in play

  • ChatGPT
  • Claude
  • Microsoft Copilot
  • Gemini
  • Perplexity
  • Otter

Upside

  • Time to first value is hours, not quarters. No integration, no security review, no procurement cycle.
  • You find out fast which tasks AI is actually good at inside your business, in your language.
  • Cost is near zero, so a failed experiment costs a week, not a budget line.

Downside

  • Nothing compounds. The prompt lives in one person's history and leaves when they leave.
  • Shadow AI risk. Customer data moves into consumer tools with no retention terms and no audit trail.
  • Leaders read individual enthusiasm as adoption and report it upward before anything is measured.

Gate to the next rung

  • You can name the three tasks where output quality was consistently good.
  • An approved tool list exists, with one sanctioned account per person instead of personal logins.
  • One prompt has been written down somewhere another person can find it.
What to measure
Number of people who used a sanctioned tool on real work in the last 14 days.
Who owns it
Enablement, with RevOps holding the tool list.

Teams who have been here

  • Shadow AI survey data

    Open dataset behind the shadow AI reporting. Download it and compare against your own tool list.

See the L1 playbooks →

L2

Assisted

AI drafts, a human approves, the workflow stays the same.

The question at this rung: Is the draft good enough that the human edit is faster than writing it?

Sequences, call summaries, account research, and first-draft content come out of AI. A person reviews before anything reaches a customer. The process on the org chart is unchanged. Volume goes up first.

Tech normally in play

  • Gong
  • Clari
  • Outreach
  • Salesloft
  • HubSpot AI
  • Jasper
  • Notion AI
  • Grammarly

Upside

  • Real hours come back, especially in research, summarization, and first drafts.
  • Quality floor rises for newer reps, because the blank page is gone.
  • You can measure edit distance and finally argue about output quality with evidence.

Downside

  • Saved time gets absorbed rather than reinvested. Hours saved is not revenue until you decide where the hours go.
  • Output volume rises without unique reach rising. More content, same pipeline.
  • Review becomes the bottleneck. If a manager approves everything, the manager is now the rate limit.

Gate to the next rung

  • A written standard exists for what good output is, so review is a measurement instead of an opinion.
  • The reclaimed hours have a named destination, agreed before the next rollout.
  • At least one workflow is stable enough that the same prompt produces usable output for a different person.
What to measure
Edit rate on AI drafts and the reinvestment destination for hours saved.
Who owns it
Frontline managers, with enablement owning the standard.

Teams who have been here

See the L2 playbooks →

L3

Integrated

AI is wired into the systems of record, not a separate window.

The question at this rung: Is the data underneath good enough for the output to be trusted?

Summaries write back to the CRM object. Enrichment fires on record creation. Slack gets the signal. Nobody copies and pastes anymore, and the failure mode moves from prompt quality to data quality.

Tech normally in play

  • Salesforce
  • HubSpot
  • Snowflake
  • Slack
  • Zapier
  • Workato
  • Clay
  • Census

Upside

  • Output reaches the place the work happens, so adoption stops depending on willpower.
  • One record, one version. Fewer contradictory answers between seats.
  • Usage becomes measurable at the object level, which makes ROI arguments possible.

Downside

  • Bad CRM data becomes confident bad output at scale. The pilot ran on a clean sandbox, production is not clean.
  • Integration debt. Every new field and every new tool adds a break point somebody has to own.
  • Cost moves from per seat to per event, and nobody has forecast the event volume.

Gate to the next rung

  • A data owner exists per object, and field-level completeness is measured, not assumed.
  • You can trace one AI-written field back to its source and its timestamp.
  • Spend is forecast against event volume, not seat count.
What to measure
Field completeness and accuracy on the objects AI reads, plus cost per event.
Who owns it
RevOps and GTM engineering.

Teams who have been here

See the L3 playbooks →

L4

Orchestrated

Multi-step work runs end to end, with a named owner and a measured result.

The question at this rung: Who owns the outcome, and what kills it?

A sequence of steps runs across tools without a human relay in the middle: research, draft, route, log, alert. A person owns the outcome number. Kill criteria are written before launch, not after the incident.

Tech normally in play

  • Agentforce
  • Copilot Studio
  • n8n
  • LangGraph
  • Zapier Agents
  • Gong Forecast
  • Clay workflows

Upside

  • Cycle time drops on the whole workflow, not one task inside it.
  • The work is repeatable across people, so it survives turnover.
  • You can finally answer what the spend bought, because a single owner holds a single number.

Downside

  • Failures compound silently. One bad step three nodes back shows up as a wrong customer email.
  • Ownership gets ambiguous fast. Marketing built it, RevOps runs it, sales gets blamed for it.
  • Pilot-to-production gap. The version that worked with five accounts breaks at five thousand.

Gate to the next rung

  • Written kill criteria with a threshold and a date, agreed before launch.
  • Observability on every step, so a failure is attributable to a node rather than a mood.
  • A rollback path that a single person can execute inside an hour.
What to measure
Cycle time on the full workflow and the exception rate per hundred runs.
Who owns it
One named operator per workflow, at director level or above.

Teams who have been here

See the L4 playbooks →

L5

Autonomous

The agent acts on live records. The human audits the exceptions.

The question at this rung: How do we stop it, and who audits the meter?

An agent resolves cases, updates records, or contacts customers without a human in the loop on every action. Humans review a sample and the exceptions. Contracts move to consumption meters the vendor defines.

Tech normally in play

  • Agentforce
  • Sierra
  • Decagon
  • Intercom Fin
  • 11x
  • Artisan
  • Claude agents

Upside

  • Coverage becomes constant. Nights, weekends, and the long tail get worked.
  • Unit cost per resolved interaction can drop hard, if the exception rate stays low.
  • Response latency becomes a competitive surface rather than a staffing problem.

Downside

  • The public failure. A wrong answer at this rung is a customer-facing event, sometimes a legal one.
  • Spend runs ahead of budget. Consumption meters bill in arrears and the vendor defines the billable unit.
  • No kill switch by default. Most teams discover this during the incident, not during procurement.

Gate to the next rung

  • An audit right in the contract, plus a parallel count you run yourself.
  • A hard spend cap and a tested kill switch you control, not a dashboard alert.
  • An escalation path with a human owner and a service level for exceptions.
What to measure
Exception rate, cost per resolved outcome, and time to stop.
Who owns it
Exec sponsor, with revenue finance on the meter.

Teams who have been here

See the L5 playbooks →

L6

Rebuilt

The process itself changes. The old workflow is not automated, it is gone.

The question at this rung: Would we design this motion this way if we were starting today?

Stage-based handoffs give way to a loop. Roles are redrawn around judgment and exceptions. Comp, quota, and headcount plans are rewritten because the unit of work changed, not because a tool was bought.

Tech normally in play

  • Custom agent platforms
  • Snowflake or Databricks
  • Internal MCP servers
  • Salesforce or HubSpot as data layer

Upside

  • Structural margin change rather than task savings. The cost curve bends, not the timesheet.
  • Defensible advantage. A competitor can buy your tools and still not have your operating model.
  • Clarity of ownership, because the new model was designed rather than inherited.

Downside

  • The public reversal risk is highest here. Announced headcount replacement is the most reversed decision in the ledger.
  • Change load lands on the same people who still have a number to hit this quarter.
  • Comp and quota design lags the new work, so the best operators are paid for the old job.

Gate to the next rung

  • This is the top rung. The test is whether the new model survives a full planning cycle, including comp.
  • Reversals are documented and published internally, so the next redesign starts from evidence.
What to measure
Cost per outcome, net revenue retention, and whether the model survived a planning cycle.
Who owns it
CEO and the exec team, not a function.

Teams who have been here

See the L6 playbooks →

The gate between each AI maturity levelA staircase of six steps rising from L1 Starter to L6 Rebuilt. Between each step is a gate naming the condition required to climb: a written tool list, a written quality standard, a data owner, kill criteria, and audit rights with a kill switch.GATES, NOT MILESTONESL1StarterOne tool, one person, n…Sanctioned tool listL2AssistedAI drafts, a human appr…Written quality standardL3IntegratedAI is wired into the sy…Named data ownerL4OrchestratedMulti-step work runs en…Kill criteria in writingL5AutonomousThe agent acts on live …Audit right and kill switchL6RebuiltThe process itself chan…
Each rung sits on a gate. The gate is the condition that has to be true before the next level is safe. Every reversal on record is a gate somebody stepped over.
Playbook count by maturity levelBar chart of how many playbooks in the library sit at each maturity level, from L1 to L6. The largest counts are at L3 Integrated and L4 Orchestrated.WHERE THE WORK ACTUALLY SITS22L1Starter22L2Assisted45L3Integrated34L4Orchestrated22L5Autonomous21L6Rebuilt
Distribution of the playbooks in the library by rung. The bulge sits at L3 and L4, which is where most revenue teams are actually working, not at L5.

Four rules for using it

Level is per workflow, not per company
A company is not at L4. One workflow is at L4 while eleven others sit at L1. Score the workflow, then look at the spread. A wide spread is normal. A single high number reported as a company grade is a vanity metric.
You cannot skip a rung and keep the receipts
L5 without the L3 data work produces confident wrong output on live customers. L4 without the L2 quality standard means nobody can say whether a run was good. Every reversal in the ledger is a skipped rung that surfaced later.
Risk rises faster than leverage between L3 and L5
The middle of the ladder is where cost moves from seats to meters, failures become customer-facing, and ownership gets ambiguous. This is the stretch that needs contract terms and kill switches, not enthusiasm.
The top rung is an operating model decision
L6 is not a bigger tool purchase. It is a redesign of who does what, measured on cost per outcome and whether the model survives a planning cycle including comp.

The mistakes it prevents

Reporting the highest rung anyone reached
One team ran an agent, so the board hears the company is autonomous. Report the median workflow, not the maximum.
Buying L5 tooling for an L2 problem
If the constraint is that nobody agrees what good output looks like, an agent does not fix it. It scales the disagreement.
Treating the ladder as a schedule
There is no obligation to reach L6. Plenty of workflows should stop at L2 or L3 permanently, because the risk at the next rung is not worth the margin.
Climbing without a stop condition
Every rung above L3 needs a written kill criterion with a threshold and a date. Without one, the rollback becomes a public event.

Terms used here are defined in the glossary and explained in plain language in the AI and Revenue Dictionary.

Common questions

What is the L1 to L6 AI maturity ladder?
The AI Maturity Ladder is a six-level model The Revenue AI Report uses to place an AI deployment: L1 Starter is one tool with no workflow change, L2 Assisted is AI drafting with human approval, L3 Integrated wires AI into the systems of record, L4 Orchestrated runs multi-step work with a named owner and kill criteria, L5 Autonomous lets an agent act with humans auditing exceptions, and L6 Rebuilt changes the process itself. Each level has its own benefit, its own failure mode, and its own exit condition. Skipping a level is the most common cause of a reversal.
How do I find what level we are at?
Pick one workflow, not the company. Ask four questions: does the output leave a single person's screen, does a human approve before it reaches a customer, does it write back to a system of record, and can it run end to end without a human relay. The last yes you can give honestly is your level.
Can we skip levels?
You can move fast, but you cannot skip the conditions. L5 depends on the data hygiene built at L3 and the quality standard set at L2. Skipped conditions surface as customer-facing failures, which is what the Reversal Ledger records.
Is L6 the goal for every workflow?
No. Many workflows should stay at L2 or L3 because the added risk at the next rung outweighs the margin. The ladder measures where you are and what it costs to move, not where you are obliged to end up.
What changes most between L3 and L5?
Cost structure and blast radius. Pricing shifts from seats to consumption meters the vendor defines, and failures move from an internal annoyance to a customer-facing event. That stretch needs audit rights, a spend cap, and a kill switch you control.
How does the ladder relate to the playbooks on this site?
Every playbook in the library is tagged L1 through L6, so you can filter to the rung you are on and see the steps, owners, pitfalls, and KPIs for that stage.

Every playbook in the library is tagged L1 through L6, with steps, owners, pitfalls, and KPIs for that rung.

Go to the playbooks →