[research@ai4se] : ~ $

ls -la ~/notes/measurement/

MEASUREMENT

7 articles

cd ../
01

Three Anchors for Efficiency Comparison: Comparability, Separability, Honesty

After AI enters software engineering, efficiency comparisons often become apples-to-oranges. This article proposes three anchors—comparability (fixed backlog / team / window), separability (freeze the complexity baseline), and honesty (expectation, not fake measured ROI)—plus reading rules for throughput, cycle time (speed framing), and paired aggregation, with common anti-patterns.

| 10 min | [dev-productivity] [measurement] [cycle-time] [throughput] [ai4se]
02

Co-Design AI4SE Efficiency Metrics with Your Team: From Fuzzy “% Gains” to Comparable Measurement

Stop stressing each other with vague “300% faster” claims. Align end-to-end roles, lock DORA and lean vocabulary, co-design metrics in two paired groups, then compare before/after on the same team, iteration, and demand scope—using percentiles—plus recipes for throughput, multi-segment cycle time, quality, and token cost.

| 12 min | [dev-productivity] [measurement] [dora] [ai4se] [kpi]
03

Designing the AI4SE Maturity Model: Differential Diagnosis, Not One-Size-Fits-All

Every team is using AI — that doesn't mean every team needs the same improvement plan. The AI4SE maturity model uses an evidence-based profile across 6 domains × 18 capabilities to identify weak spots, then prescribes a dosage-matched capability-building plan by maturity type — the point of assessment isn't scoring, it's differential diagnosis.

| 18 min | [maturity-model] [measurement] [ai4se-framework]
04

Productivity Data in Anthropic's 2026 Report: What to Measure, and What Not To

Distilling observable signals from Anthropic's 2026 trends report: 60% AI penetration vs. 0–20% full delegation, 27% net-new work, and the measurement implications behind cases like TELUS and CRED.

| 7 min | [dev-productivity] [measurement] [anthropic]
05

From DORA to DevEx: A Panorama of AI4SE Measurement Frameworks

DORA measures delivery, SPACE measures productivity, DevEx measures experience — AI4SE needs to integrate all three frameworks, because agents simultaneously affect delivery speed, how people work, and developer experience.

| 9 min | [dev-productivity] [measurement]
06

DORA 2025: AI Is an Amplifier, Not a Silver Bullet

Key findings from Google's DORA 2025 report: 90% of software professionals are using AI, 59% report improved code quality — but AI amplifies both good and bad engineering practices.

| 10 min | [dora-metrics] [measurement]
07

Developer Productivity in the AI4SE Era: Why Traditional Metrics Are Failing

Measuring AI4SE-era developers by lines of code, story points, or PR counts? These metrics are not just useless — they actively encourage the wrong behavior. We need a new framework.

| 11 min | [dev-productivity] [measurement]