Sponsors ask “how much did AI improve us?” and often want a clean multiplier: 2×, 5×, sometimes slogans in the thousands of percent. The louder the number, the more it sounds like finished delivery evidence.

Reality is messier. Developer productivity is multi-dimensional—SPACE already warns that satisfaction, performance, activity, collaboration, and efficiency/flow can move in different directions; collapsing them into one throughput slogan is usually a worse decision aid, not a better one. Worse, perception can diverge from measurement: in a METR RCT of early-2025 AI tools, experienced open-source developers forecasted speedup yet were about 19% slower when AI was allowed.

What organizations still lack is rarely another louder N×. It is a management-usable assessment discipline: one that answers both “are capabilities being institutionalized?” and “what throughput / cycle-time lift should we expect?”, and that labels how every number was obtained. What follows is the reading we use in consulting practice—maturity and measurement together; a multiplier claim is not a substitute for that pairing.

1. Three failures of the bare multiplier

1. Build-local speedup sold as end-to-end ROI

A large gain confined to the build (code-generation) stage is not the same as an equal gain across the lifecycle. Plan, review/test, and deploy can stay near human speed—Anthropic’s AI-native SDLC playbook likewise argues that when build accelerates, surrounding stages can become the bottleneck. If management narrates “coding felt N× faster” as “delivery is N× faster,” the denominator has already changed.

2. Missing comparator and fixed conditions

Any N× needs an explicit comparator and fixed measurement conditions. Without them the number is not interpretable. AI4SE also changes how work is decomposed, verified, and shipped—“same work, faster” breaks down when the unit of delivery itself shifts. Our post on the three anchors for efficiency comparison compresses the discipline into three: comparability (hold Product Backlog, team, and iteration window), separability (freeze story points; After must not re-point), and honesty (high-consensus promissory expectation, not disguised platform ROI). Without those nails, percentages are slide decoration.

3. Missing signal-type labels

Maturity scores, workshop expert estimates, and platform-measured throughput/cycle time are three different evidence strengths. Mixing them into one “ROI” table without labeling each cell invites the most dangerous narrative swap: treating consensus expectation as realized return, or treating domain-mean maturity as business outcome. Published experiments typically report moderate, conditional gains (for example, roughly +26% completed tasks across pooled field experiments, and about 21% shorter time with a wide CI on a complex enterprise task)—already a different species from folklore organizational N×. Without labels, folklore wins the slide.

2. Dual-track dashboard: maturity and measurement together

Sponsors are really asking two faces:

TrackQuestionHonest read
Capability maturityAre efficiency-related capabilities being institutionalized?Modest domain-mean lifts (+0.3 to +1 on a 1–5 encoding) in a short pilot are often an expected finding, not automatic failure
Delivery-result signalsUnder fixed conditions, what throughput / cycle-time lift should we expect?Treat workshop numbers as promissory expectations first; validate or refute them later with platform telemetry under the same definitions

Maturity alone under-answers “how much throughput / CT?”; metrics alone ignore institutionalization. Publishing both on separate scales blocks collapsing them into one unvalidated multiplier.

On the maturity side, we use a process-capability portrait (for example 6 domains × 18 capabilities, L1–L5) for before/after contrast—see maturity model design. It tracks organizational capability change, not product quality or business outcome, and not platform-measured ROI.

On the metrics side, a practical starting pair (recipes and co-design flow in co-designing efficiency metrics) is two different constructs:

  • Throughput (volume): Σ\Sigma story points of stories completed in the fixed iteration window (monotone-equivalent to points/person-day when team and window are fixed).
  • Cycle time (speed): elapsed duration of a story between two workflow endpoints—e.g. CT1 (DEV accept → latest related commit) and CT2 (same start → QA testing completed). Primary-report percentiles (e.g. P75); compute paired Δ\Deltas first, then cross-group median + IQR—not means of means, and not the maximum pair as the organizational headline.

Critical discipline: throughput improve Δ\Delta and CT speed-improve Δ\Delta must not be narrated as one interchangeable “≈2× everywhere.” One is in-window completed volume; the other is duration speedup. Keep them separate in prose.

Joint reading is three lines:

  1. Capability lift ≠ ROI
  2. Estimate ≠ platform-measured
  3. Both tracks appear labeled on the same decision artifact

3. Signal types: every published cell carries a label

Every number you publish should carry a signal type:

signal_typeMeaning
process_maturityFacilitated process-capability maturity assessment (domain/capability levels)
metric_estimateThroughput / CT from SME or pair judgment—promissory expectation
metric_platformThe same Throughput / CT definitions computed from the requirements platform (and Git timestamps when needed) in real iterations

Honest assessment here means something concrete: every published cell carries a label; Throughput/CT medians stay promissory on an estimate→measured promotion path until matching metric_platform values validate or refute them. An empty platform cell is itself information—do not paper it over with a louder N×.

A first-order credibility trap: Before can also be an estimate. To contrast methods on the same Product Backlog, team, and timebox, workshops sometimes elicit both traditional delivery and the new method (e.g. Spec-Driven under a high-consensus After prompt) as metric_estimate. Then Before→After is an estimate-to-estimate contrast—not measured-baseline→forecast, and not platform→platform. The gap can reflect genuine method expectation and shared optimism (an “imagination gap”)—which is why metric_platform validation remains mandatory before ROI language.

4. Sponsor decision tree (anti-misquote)

When the pattern is “modest maturity movement beside large workshop median Δ\Deltas”—common in short enablement pilots—read in this order instead of forcing one success story:

  1. Fund next-phase method sedimentation and instrumentation, do not declare ROI.
  2. Schedule metric_platform validation against the same definitions before any funding slide quotes “about +100% Σ\SigmaSP” or “about 2× CT” as achieved productivity.
  3. Primary-claim median + IQR; forbid retelling one pair’s maximum as the organizational headline.
  4. Ask who owns the dual dashboard (engineering leadership + delivery leads co-sign labels) and what incentive pressure would punish honest estimate≠measured divergence.
  5. Refuse to narrate a build-stage or workshop N× as end-to-end lifecycle ROI without an explicit comparator and fixed conditions.
  6. Quality remains a precondition: efficiency asks sit under non-degrading quality gates; omitting a parallel quality track here is not a claim that quality is unimportant—it means do not let throughput slogans replace the client’s existing quality governance.

Executive narrative may round to “about +100% Σ\SigmaSP under workshop conditions (promissory)”; precision stays in tables and CSVs.

5. Anti-patterns and minimal actions

The scheme fails if:

  • Published rows drop signal_type
  • After re-points work (separability is broken)
  • Workshop maxima replace median + IQR as primary claims
  • Estimate rows are promoted to ROI without platform-measured confrontation

Enablement fails if:

  • Later metric_platform Δ\Deltas stay near zero under the same definitions while “high method consensus” was claimed
  • Maturity domains tied to the method pack stay flat across reassessment cycles with no asset-adoption evidence
  • Efficiency narratives advance while quality red-line indicators are knowingly breached or dropped from governance

Either failure is reportable news—not a reason to invent a multiplier.

Minimal actions (one slide is enough to start):

  1. On one page: maturity domain means (process_maturity) + workshop median/IQR (metric_estimate).
  2. Name who owns the estimate→measured promotion path and when platform validation is scheduled.
  3. Write the bans: do not collapse the two tracks into one N×; do not headline extreme pairs.

Close

Literature point estimates already contradict folklore organizational multipliers; perception and measurement can even move opposite directions. What transfers is not another industry-wide “2×,” but assessment discipline: maturity beside labeled measurement; every cell tagged by how it was obtained; empty platform cells kept visible until filled or falsified.

Multipliers make good slogans. They rarely make good stage-gate evidence.