[research@ai4se] : ~ $
cd ../
[measurement] | | 18 min

# Designing the AI4SE Maturity Model: Differential Diagnosis, Not One-Size-Fits-All

Every team is using AI — that doesn't mean every team needs the same improvement plan. The AI4SE maturity model uses an evidence-based profile across 6 domains × 18 capabilities to identify weak spots, then prescribes a dosage-matched capability-building plan by maturity type — the point of assessment isn't scoring, it's differential diagnosis.

[maturity-model][measurement][ai4se-framework]

This article is based on the design draft and assessment sheet for AI4SE team maturity model v0.3.5, aimed at engineering managers, effectiveness leads, and consultants who need to diagnose current-state maturity, produce evidence-based scores, and drive continuous improvement within an organization. The companion assessment sheet is oriented toward capability profiles and risk-item output, and does not compute a team total score.

The most dangerous habit in a clinic isn’t “not knowing how to prescribe” — it’s “prescribing the same drug at the same dose for every case that looks like a cold.”

AI4SE transformations have the same failure mode: once every team has a coding agent installed, everyone gets the same workflow template, the same training course, the same “autonomy” target. The usual result: the tooling layer looks busy, but process and human-agent collaboration contracts haven’t kept up — or the basics aren’t even in place yet, and the team is already chasing an L4/L5 systematized narrative.

The key point of maturity assessment isn’t slapping a level label on a team — it’s that teams at different maturity levels need genuinely different improvement plans. Even if every team eventually needs to add “evidence-driven review,” the set of actions and the investment dose differ between going from L2 to L3 versus L3 to L4. The model exists to find capability gaps, then prescribe an AI4SE capability-building plan through differential diagnosis, case by case.

This article explains why the model is designed this way, how leveling works, and how to read assessment output as a “prescription by type.” The complete L1–L5 behavioral anchors for each capability item are governed by the assessment sheet; this article retains only the design logic and the reading method.


I. Why a Team-Level Maturity Model Is Needed

Buying tools, running training, and announcing “full AI embrace” are all easy to turn into mere activity. The hard part is answering three questions:

  1. What is stably happening right now? Not what’s written in the plan, and not what was demoed.
  2. Where exactly is the weak spot? Unclear effectiveness goals? Unchanged process? No established methods? Unbounded tooling? No collaboration contract? No organizational soil?
  3. What “medicine” should be prescribed next, and at what “dose”? Keep stacking tools, or first fix the gates and the chain of accountability?

DORA, SPACE, and DevEx answer “what happened in delivery and experience.” The layered technology model answers “what structure should we use to understand engineering capability.” The maturity model fills the third gap: the team’s current profile across AI4SE capability domains, and the improvement priorities that follow from it.

It deliberately avoids four things:

What It AvoidsWhy
A single team total scoreAn average masks the critical weak spot; “total score 3.2” doesn’t tell you what to fix next week
Team rankingsAssessment serves improvement, not competition
Fixed-timeline maturity commitmentsThe model supports re-assessment, but doesn’t bake in “must reach L4 in three months”
Binding to a specific vendor/frameworkIt assesses capability, not a procurement checklist

II. Model Positioning: Diagnostic Profile, Prescriptive Output

The model assesses a team’s current capability profile with respect to AI4SE, identifies gaps, and provides an evidence base for continuous improvement. It is not tied to a single project context, and does not judge teams by a single total score.

Should output:

  • Capability-domain profile: maturity distribution across each of the six domains
  • Capability-item heatmap: L1–L5 status for the 18 capability items
  • Evidence strength and scoring confidence: whether the score is reliable
  • Bottleneck items and evidence gaps: why a higher level can’t be assigned
  • Continuous-improvement recommendations: next-step direction based on the gaps

Should not output: a total score, a ranking, a timeline-based leveling commitment, or a mandate for a specific tool.

In one line:

Assessment is differential diagnosis; the prescription is a capability-building plan matched to the weak spots and evidence gaps.


III. Design Principles

PrincipleDescription
Current-state scoringOnly score practices that are already stably happening — not plans, visions, or one-off demos
Evidence firstScores are based on observable facts, artifacts, process records, system data, and cross-validated interviews
Team consistencyOne person achieving something doesn’t mean the team can; individual practice typically caps at L2
Score to the lower boundIf only some conditions of a higher level are met, score at the highest level whose conditions are all stably satisfied
Capability, not toolingAnchors describe what capability the team has, not which vendor or framework is hardcoded
Re-assessableThe same model can be reused in later cycles, but doesn’t bake in a fixed time target
No total scoreOutput is profile, bottlenecks, and improvement recommendations — not a ranking score

These principles all serve one purpose: making the “prescription” rest on verifiable facts, not slogans or gut feel.


IV. Capability Structure: 6 Domains × 18 Items

The model uses 6 capability domains and 18 capability items, aligned with the AI4SE layered technology model: the Effectiveness foundation, the three layers of Process / Methods / Tools, the Harmony cross-cutting layer, and organizational/cultural support.

DomainCapability DomainFocusCapability Items
D1Effectiveness FoundationQuality + Efficiency + ValueValue goals and outcome hypotheses; quality and trust baseline; efficiency and flow measurement
D2Process LayerAuditable closed loop for human-agent collaborationAI4SE workflow design; small-batch task decomposition; evidence-driven review/ship gates
D3Engineering Methods & Work DisciplineReusable, trainable, verifiable methodsSpec formalization; context organization; verification and evidence discipline
D4Tools LayerAgent runtime, context, and automated verificationAgent runtime and permissions; AI-accessible engineering context; automated verification and toolchain integration
D5Harmony (Human/Agent Collaboration)Collaboration contracts cutting across all three layersCollaboration responsibilities and independent verification; HITL/HOTL/HOOL risk routing; accountability and audit trail
D6Organization & CultureSustained adoption and team learningPractice champions and training; community and knowledge assets; adoption confidence and behavior change

The structure itself hints at “differential typing”: if D4 is clearly ahead of D2/D5, the typical prescription isn’t buying an even stronger agent — it’s filling in workflow, gates, responsibilities, and risk routing. Otherwise, the stronger the tooling, the faster it outruns governance.

Mapping to Reference Frameworks

This ModelCorresponding Source
D1 Effectiveness FoundationAI4SE Effectiveness Focus; DORA Value/Flow perspective
D2 Process LayerPressman Process; Research → Plan → Execute → Review → Ship
D3 Engineering MethodsPressman Methods; SDD, Agentic Engineering, Harness, and other stable capability abstractions
D4 Tools LayerPressman Tools; agent runtime, permissions, context, automated verification
D5 HarmonyHarmony cross-cutting layer; HITL / HOTL / HOOL supervision spectrum
D6 Organization & CultureAdoption, training, communities of practice, and other organizational soil

The measurement frameworks (DORA / SPACE / DevEx) tell you “how well outcomes and experience are going”; the maturity model tells you “where the capability structure is missing, and what to fill in next.”


V. Level Semantics: L1–L5

The primary labels use L1–L5. If alignment with Microsoft’s CMM-style naming is needed, they can be mapped to L100–L500.

LevelNameGeneral Leveling Semantics
L1InitialPractices are scattered, individual, non-repeatable, lacking team standards and evidence
L2TooledLocalized practices or preliminary rules exist, but rely on individual experience; weak team consistency and governance
L3EngineeredClear standards, role responsibilities, required artifacts, and review mechanisms are established and repeatable
L4SystematizedStandards are embedded in process and toolchain, key data is traceable, and a continuous-improvement loop exists
L5AutonomousWithin clear governance boundaries, low-risk paths run stably and automatically, exceptions escalate, decision evidence is traceable, and continuous optimization is driven by operational data

Note: L2’s “Tooled” is shorthand meaning AI4SE has moved from scattered ideas into localized practice or preliminary rules — it does not mean only tool procurement is assessed.

For “prescribing,” the gap between levels matters more than the absolute score:

  • Most capability items stuck at L1–L2: prioritize building team standards, templates, execution records, and evidence discipline, rather than chasing an autonomy narrative.
  • Most already at L3, some pushing toward L4: embed standards into the task system, review gates, and toolchain, so that missing evidence actually blocks delivery.
  • Tools high, process/Harmony low: fix the collaboration contract and gates first, before expanding agent permissions.

The same “drug” (e.g., evidence-driven gates) needs a different dose: L2→L3 is usually “require a change description / test / risk note before review”; L3→L4 is “missing evidence blocks or delays delivery, and must be embedded in the toolchain.”


VI. Evidence-Based Leveling: Making the Prescription Verifiable

1. The Five-Step Leveling Method

Final level = the highest level whose necessary conditions are met, constrained by evidence strength, team coverage, and cap rules.

  1. Confirm current stable practice: only look at behavior that is already happening stably.
  2. Compare against behavioral anchors: choose the L1–L5 description closest to the current facts.
  3. Check necessary conditions: if conditions are incomplete, drop to the highest level whose conditions are fully met.
  4. Apply cap rules: cap based on evidence strength, coverage, execution records, and the “boundary of ordinary engineering practice.”
  5. Produce the final interpreted level: record the evidence summary, AI4SE incremental evidence, evidence gaps, confidence, and improvement recommendations.

The v0.3.5 assessment sheet uses expert consultant scoring as the primary maturity input; what ultimately enters the dashboard is the final interpreted level, capped by evidence and coverage rules — not a subjective gut-feel score.

2. Evidence Strength and Coverage Caps

Evidence StrengthMeaningCap Tendency
LowPersonal description, single case, unverifiable screenshotUsually caps at L2
MediumSome documentation/interviews/limited samples, but insufficient continuity or coverageUsually caps at L3
HighA complete work cycle or 4+ weeks, multiple samples, multiple sources, linkable to work artifactsOnly this qualifies for interpreting L4/L5
CoverageMeaningCap Tendency
IndividualA few members can do itUsually caps at L2
SubgroupA single role or localized scenarioUsually caps at L3
TeamMain delivery roles and key steps use it stablyBaseline threshold for L3+
Cross-teamMultiple teams share governance and knowledge mechanismsSignal of organizational-level diffusion

3. AI4SE Incremental Evidence

Ordinary PR, CI, Git, and test records can serve as baseline evidence, but cannot by themselves prove AI4SE maturity. The assessor must be able to explain:

  • Which step AI was involved in;
  • How the team adjusted process, verification, permissions, responsibility, or knowledge assets as a result;
  • Whether these adjustments can be independently verified.

Without incremental evidence, it’s easy to mistake “we’ve always had CI” for “we’ve already systematized AI4SE” — and the prescription will be wrong too.

4. Boundaries Between Adjacent Capability Items

To avoid the same piece of evidence being reinterpreted repeatedly, assessment should distinguish:

Capability ItemPrimarily AssessesDoes Not Substitute For
D2.3 Evidence-driven gatesWhether evidence is required before review/delivery, and whether missing evidence blocks deliveryNot a substitute for D3.3 verification discipline, nor D5.3 accountability chain
D3.3 Verification and evidence disciplineHow AI output is verified and what evidence is retainedNot automatically high-scored just because ordinary CI exists
D4.3 Automated verification and toolchainWhether AI-assisted changes are identifiable and enter a unified verification pipelineNot a substitute for gate decisions or human accountability
D5.1 Collaboration responsibility and independent verificationHuman/agent responsibilities, handoff points, whether verification is independent of the generating party”Everyone eyeballed it” ≠ automatic high score
D5.3 Accountability and audit trailWhether who approved, who verified, who merged, who is accountable is traceableOrdinary code review records ≠ AI4SE accountability chain

Clear boundaries make weak spots clear; clear weak spots make the prescription avoid becoming “one drug for every disease.”


VII. From Profile to Prescription: Same Disease, Different Treatment

The product of a maturity assessment isn’t the radar chart itself — it’s the differentiated, executable improvement plan. Below are three common profile types to illustrate “differential typing” — all of them might be verbally described as “the team is using AI,” yet the prescriptions are completely different.

Type A: Tools Ahead, Process and Harmony Weak

Profile characteristics: D4 is relatively highest; D2 and D5 clearly lag; L1/L2 items cluster around workflow, risk routing, and accountability records.

Wrong prescription: keep buying stronger agents, expand auto-code-change permissions, use tool-usage rate as a success metric.

Matched prescription:

  1. First define a team-level AI4SE workflow and small-batch task standard (D2.1 / D2.2);
  2. Establish evidence-driven review/ship requirements (D2.3), aligned with verification discipline (D3.3);
  3. Clarify responsibilities, independent verification, and HITL/HOTL/HOOL risk routing (D5);
  4. Tool-permission expansion must follow the collaboration contract, not precede it.

Type B: Stuck at L1–L2 Across Many Domains

Profile characteristics: widespread L1/L2 across the board; occasional individual experts; knowledge scattered across chats and personal notes.

Wrong prescription: benchmark directly against “autonomous agent teams,” adopt complex multi-agent orchestration.

Matched prescription:

  1. Pick a few high-value scenarios and build repeatable spec/context/verification templates (D3);
  2. Identify practice champions and provide role-based coaching rather than one-off training (D6.1);
  3. Turn success and failure patterns into team knowledge assets (D6.2);
  4. Target repeatable engineering (L3) first, before talking about systematization.

Type C: Most Items Near L3, Some Pushing Toward L4

Profile characteristics: standards and execution records already exist; what’s missing is toolchain embedding, data closed-loops, and decision influence.

Wrong prescription: run another round of “awareness-raising” training, or use satisfaction surveys as a substitute for engineering change.

Matched prescription:

  1. Embed evidence requirements, permissions, and verification results into the task system / PR / release gates (D2 / D4);
  2. Let metrics influence prioritization and governance decisions, not just retrospective reporting (D1);
  3. Bring the accountability chain and risk routing into auditable records (D5.2 / D5.3);
  4. Use re-assessment to verify “have low-maturity items been cleared,” not just whether domain averages rose.

”Same Drug, Different Dose” Example

Take D2.3 Evidence-driven review/ship gates as an example:

Current Rough PositionNext Dose of ImprovementSuccess Signal
L1–L2Define a required evidence checklist before review, and leave execution records on real tasksShift from “self-reported” to “a stable evidence format exists”
L2–L3Evidence requirements become team consensus; gaps are flagged and filledConsistent execution visible across multiple real delivery samples
L3–L4Evidence requirements are embedded in gates; missing evidence blocks or delays deliveryGate-failure records are queryable and affect real delivery
L4–L5Gate strictness adjusts dynamically by risk, change type, and historical qualityOperational data proves the dynamic policy is effective

If an assessment gives a single “D2.3 = L2” score but prescribes “add a dynamic gating engine,” that’s a classic dosage error.


VIII. How to Use the Assessment Sheet (v0.3.5)

The formal assessment sheet collects the 18 capability items, behavioral anchors, necessary conditions, suggested evidence, cap rules, and evidence fields on the same scoring page. The recommended pace for the consultant version:

  1. Fill in expert consultant scores based on interviews, evidence, and delivery records; when facts are unknown, mark “unclear/no evidence” rather than defaulting to L1.
  2. Read the L1–L5 anchors, necessary conditions, and cap rules in each capability item’s notes.
  3. Add an evidence summary, evidence type, AI4SE incremental evidence, sources/links, and sample period.
  4. Fill in the evidence score, evidence strength, team coverage, confidence, evidence gaps, and improvement recommendations.
  5. Review rule-conflict warnings and the final level: when the consultant’s score exceeds the “suggested maximum level,” it is explained by the evidence-capped result.
  6. Use the dashboard to view the capability-domain profile; use the risk register to aggregate low-maturity, low-evidence, low-coverage, and incremental-evidence-gap items.

The dashboard defaults to the final interpreted level, does not use the average score as a team total, and does not rank. Domain averages are only used to view the overall shape; decisions should still look at the lowest-scoring items, the count of L1/L2 items, risk items, and evidence gaps.


IX. Reading Assessment Output: A Mock Walkthrough

The charts and data below are reproduced from a visualization mock (ai4se-maturity-visualization-mock-v0.1): the radar chart shows overall shape, the bar chart shows domain differences, the improvement bars show the effect of actions taken, and the heatmap shows individual capability items. All data is simulated and does not represent a real team or a maturity commitment — it’s used here to demonstrate “read the chart → diagnose the type → prescribe.”

1. Start With the Dashboard’s Key Signals

Mock dashboard key signals overview

Four things you can read at a glance: whether L1/L2 items form a cluster, whether any L4/L5 already exists, which domains changed the most, and which typical profile type currently applies. In this example, the first assessment is flagged as Type A: tools ahead, process/Harmony weak — that’s already a prescription direction, not just a score summary.

2. First Assessment: A Typical Structural Imbalance

Six-domain maturity profile: before vs. after improvement

On the radar chart, the gray “first assessment” contour sits close to L2 overall, with only D4 Tools Layer clearly extended outward; the teal “after improvement” contour is rounder and further out. When reading the first-pass contour, the key point isn’t “what’s the average” — it’s whether the shape is imbalanced.

Capability-domain before/after comparison bar chart
DomainCapability DomainDomain Profile (mean)Interpretation Clue
D1Effectiveness Foundation2.00Some local goals/gut sense, not yet engineered
D2Process Layer1.67Workflow and gates are weak
D3Engineering Methods2.00Spec/context/verification still individual-driven
D4Tools Layer2.67Relatively ahead
D5Harmony1.67Risk routing and collaboration contract insufficient
D6Organization & Culture2.00Training and knowledge assets not yet mechanized

The global signal is even more stark: 16 of the 18 items fall at L1/L2, and L4/L5 is zero. This isn’t a “buy more tools” signal — it’s Type A: tools relatively ahead, process and Harmony weak.

Drilling further into the capability-item heatmap — domain averages can mask individual weak spots, but the heatmap doesn’t:

18-item capability heatmap: before vs. after

At the item level, the weak spots are very specific (highlighted in red in the chart):

  • D2.1 AI4SE workflow design: L1 (still at individual-usage stage)
  • D5.2 HITL/HOTL/HOOL risk routing: L1 (no risk tiering in place)
  • D4.2 / D4.3: already at L3 (context and verification pipeline relatively ahead)

The first prescription drawn from this should target D2 and D5, not keep piling onto D4: turn individual usage into a team workflow, upgrade “everyone eyeballs it” into tiered risk routing with an auditable accountability chain, and simultaneously fill in spec/context/verification templates (D3) and value hypotheses (D1).

3. Post-Improvement Assessment: Look at Cleared Weak Spots, Not Just a Rising Average

MetricFirst AssessmentAfter Improvement
L1/L2 capability items160
L4/L5 capability items03
Six-domain profile~L2 range, D4 slightly aheadAll at or near L3, D3/D4 locally reaching L4
Improvement magnitude by capability domain

The domains with the largest gains are D2 / D3 / D5 (roughly +1.33). This matches the prescription direction: the weak domains were prioritized for improvement. The right-hand column of the heatmap shows items that locally reached systematization:

  • D3.3 Verification and evidence discipline → L4 (evidence now feeds into review and delivery decisions)
  • D4.2 / D4.3 → L4 (context is controlled and maintainable; AI-assisted changes stably enter the verification pipeline)

When reading the charts, it helps to consistently ask three questions:

  1. What does the shape look like? (Type A/B/C or some combination) — start with the radar chart and dashboard
  2. Which items are dragging things down? (the lowest items and the L1/L2 list) — then check the heatmap’s red zones
  3. Can the evidence hold up? (would low evidence, insufficient coverage, or incremental-evidence gaps make a “high score” unexplainable)

A rising domain average with L1/L2 items still numerous means the weak spots haven’t actually been cleared; more L4/L5 items with risk routing still empty means the autonomy narrative is running ahead of governance. A good re-assessment proves the prescription worked — not that the score looks good.


X. Wrap-Up: Assessment in Service of Actionable Capability Building

The design choices behind AI4SE Maturity Model v0.3.5 can be summed up in four points:

  1. Use 6×18 to align with the layered engineering structure, avoiding assessing only tools or only culture.
  2. Cap using evidence and coverage, avoiding treating plans, demos, and individual heroics as team capability.
  3. Output profile and gaps, not a total score, avoiding letting rankings substitute for diagnosis.
  4. Read the assessment as a typed prescription: different maturity, different weak spots, genuinely different improvement plans and doses.

Back to the opening analogy: patients may all have “a cold,” but a good doctor still asks about history, constitution, and complications before deciding on a different dose or a different treatment. Every team is “using AI” — professional AI4SE improvement should work the same way: diagnose the type first, then build capability case by case.

Related reading: