Autonomous engineering team with Devin + Claude Code (L5)

A handful of senior engineers + a fleet of autonomous agents (Devin for project-level work, Claude Code for repo-level chores, Codex for issue triage). Humans set strategy and architecture, write the hard 20% of code, review everything. Output per human engineer 3\,5x.

WORKFLOW1Re-org around agent orche…trationManual2Tier the agent fleetDevin3Continuous evals on agent…outputGitHub4Comp + career-ladder rewr…teManual
4 steps, in order, with the tool that owns each one.
Adoption ladderSix levels from Starter to Rebuilt. This item sits at level 5.L1 StarterOne tool, no workflow changeL2 AssistedAI drafts, humans approveL3 IntegratedWired into CRM and SlackL4 OrchestratedMulti-step, owned, measuredL5 AutonomousAgent runs, human auditsL6 RebuiltThe process itself changes
This playbook belongs at L5 Autonomous. Running it above your level is how pilots stall.
Measures of successfeatures shipped per human engineer; agent cost per merged PR; staff-eng time spent on architecture vs implementationPROVE IT WORKEDfeatures shipped per human engineeragent cost per merged PRstaff-eng time spent on architecture vsimplementation

The steps

  1. 01

    Re-org around agent orchestration

    Tool: Manual

    Restructure eng: pods of 2 senior engineers + 1 staff eng now own what used to take 6 engineers. The staff role is "agent wrangler", decomposes work into agent-shaped tasks, defines acceptance criteria, reviews and rejects. Owner: VP Eng + CEO. This is a real org change, read "AI-native operating model (L6)" before doing this. Pitfall: keeping headcount flat and just "adding agents", you'll get burnout from review overload. DoD: re-org communicated; new pod structure and role definitions live.

  2. 02

    Tier the agent fleet

    Tool: Devin

    Three tiers: • Tier 1 (Codex): issue triage, small bug fixes, dependency bumps. Cheap, high-volume. • Tier 2 (Claude Code): multi-file refactors, test-coverage backfill, migrations. Mid-cost. • Tier 3 (Devin): multi-day project work, "build the X service to spec", with checkpoints. Expensive, requires senior oversight. Document which tier owns which work shape. Owner: staff eng per pod. DoD: a routing rubric, "this kind of work goes to this agent", exists and is followed.

  3. 03

    Continuous evals on agent output

    Tool: GitHub

    Stand up an internal eval harness: every merged agent PR is scored on (a) defects in next 30 days, (b) revert rate, (c) follow-on cleanup tickets. Aggregate per agent tier. If Tier 3 (Devin) has a defect rate >2x human PRs, narrow what you assign to it. Owner: platform eng. Pitfall: trusting marketing benchmarks. Trust your own eval. DoD: monthly agent-quality scorecard published internally.

  4. 04

    Comp + career-ladder rewrite

    Tool: Manual

    Promote engineers based on "systems shipped", "agents managed", "customer outcomes", NOT lines of code or PRs authored. Otherwise you'll punish your best agent-wranglers. Rewrite the career ladder before performance season, not after. Owner: VP Eng + People. DoD: new ladder live; first promotions under the new model done before AI ROI is reported externally.

Tools in this playbook

Next playbooks

Unfamiliar terms are defined in the AI and Revenue Dictionary. Related frameworks live in the framework library.

Share this playbook

Posting to Instagram or TikTok? Copy the link, it carries the title, summary and share image.

Arrives weekly by email. Free. Unsubscribe anytime. By subscribing you agree to our Privacy policy and Terms. We never sell or share the list.