ls -la ~/notes/measurement/
MEASUREMENT
7 articles
cd ../Three Anchors for Efficiency Comparison: Comparability, Separability, Honesty
After AI enters software engineering, efficiency comparisons often become apples-to-oranges. This article proposes three anchors—comparability (fixed backlog / team / window), separability (freeze the complexity baseline), and honesty (expectation, not fake measured ROI)—plus reading rules for throughput, cycle time (speed framing), and paired aggregation, with common anti-patterns.
Co-Design AI4SE Efficiency Metrics with Your Team: From Fuzzy “% Gains” to Comparable Measurement
Stop stressing each other with vague “300% faster” claims. Align end-to-end roles, lock DORA and lean vocabulary, co-design metrics in two paired groups, then compare before/after on the same team, iteration, and demand scope—using percentiles—plus recipes for throughput, multi-segment cycle time, quality, and token cost.
Designing the AI4SE Maturity Model: Differential Diagnosis, Not One-Size-Fits-All
Every team is using AI — that doesn't mean every team needs the same improvement plan. The AI4SE maturity model uses an evidence-based profile across 6 domains × 18 capabilities to identify weak spots, then prescribes a dosage-matched capability-building plan by maturity type — the point of assessment isn't scoring, it's differential diagnosis.
Productivity Data in Anthropic's 2026 Report: What to Measure, and What Not To
Distilling observable signals from Anthropic's 2026 trends report: 60% AI penetration vs. 0–20% full delegation, 27% net-new work, and the measurement implications behind cases like TELUS and CRED.
From DORA to DevEx: A Panorama of AI4SE Measurement Frameworks
DORA measures delivery, SPACE measures productivity, DevEx measures experience — AI4SE needs to integrate all three frameworks, because agents simultaneously affect delivery speed, how people work, and developer experience.
DORA 2025: AI Is an Amplifier, Not a Silver Bullet
Key findings from Google's DORA 2025 report: 90% of software professionals are using AI, 59% report improved code quality — but AI amplifies both good and bad engineering practices.
Developer Productivity in the AI4SE Era: Why Traditional Metrics Are Failing
Measuring AI4SE-era developers by lines of code, story points, or PR counts? These metrics are not just useless — they actively encourage the wrong behavior. We need a new framework.