Teams bolt on slash commands, skills, and rules for a Coding Agent—then swap models or sessions and watch behavior drift. The usual failure is not “one more prompt missing.” Those assets are not hung on a stable process spine, and they lack clear entry points, autonomy levels, and stop conditions.
A pile of plugins is not a User Harness. Harness Engineering describes the control system; this article is about landing that control system as an extensible tool framework—ai4se-harness.
Thesis in one line: Entry = four work-item Commands × three Profiles; execution = three Gate Agents hung on P0–P5. Method and gates stay fixed; asset packs and human–AI depth vary.
Core spine: P0–P5 and three quality gates
The Core of ai4se-harness does not invent a new methodology. It inherits the stage flow from the SDD end-to-end practice guide. User-state assets hang on this spine:
| Stage | Name | Key outputs | Related gate |
|---|---|---|---|
| P0 | Baseline & init | Specs library, Rules carrier, knowledge dirs in place | — |
| P1 | Intent & proposal | proposal.md (scope / non-scope / assumptions) | — |
| P2 | Spec definition & align | Delta Spec, acceptance scenarios | G1 Align |
| P3 | Design & task breakdown | design, tasks, acceptance design | — |
| P4 | Implement & verify | implementation, tests, verification report | G2 Verify |
| P5 | Codify & evolve | Living Spec merge, archive, asset improvements | G3 Merge & Archive |
The three gates intercept different failure modes at different rework costs:
| Gate | Asks | Intercepts |
|---|---|---|
| G1 Align | What does “correct” look like? | Directional errors (edit text; cheapest) |
| G2 Verify | Was it built to the Spec? | Execution errors (edit code; medium cost) |
| G3 Merge & Archive | Was learning retained; is the source of truth consistent? | Consistency and missing codification (hurts the long term) |
Profiles may change confirmation frequency; they must not relax a gate’s DoV (Definition of Verified). For Spec as source of truth, see Spec as Source of Truth.
Zero-layer architecture: User asset packs on Core flows
The zero layer splits the system into two bands: User above holds extensible asset packs; Core below provides SDD flow capability and supports the matching assets upward.
How to read it:
- Layering: User owns the customization surface; Core owns the flow engine and defaults.
- One-to-one: Each Core flow maps to a User Assets pack (default flow ↔ standard pack; mode-specific flows ↔ lean / agile / hybrid packs).
- Extensible: Assets and workflows can grow sideways without rewriting the engine.
The on-site project knowledge layer covers “what Guides know”; this article covers how whole asset packs bind to the stage flow.
User state: five assets and call / governance boundaries
Align with the usual Coding Agent stack (Claude Code / Cursor / Codex and peers):
| Asset | Role | Industry counterpart |
|---|---|---|
| Command | Thin entry: normalize a work item into a flowing Change Request | slash command / command entry |
| Agent | Goal-driven loop: plan → pick Skill/tools → act → observe | Agent / sub-agent loop |
| Skill | Reusable capability pack (prompts, steps, context, tool usage) | Skills / prompt+tool bundles |
| MCP / CLI / API | Concrete execution surface | MCP servers, shell, HTTP |
| Rules | Cross-cutting governance: decisions, boundaries, tool permissions | CLAUDE.md / .cursor/rules / AGENTS.md |
Constraints:
- Main path is Command → Agent; Command→Skill direct calls fit only light cases.
- Agent is the orchestrator: it may call Skills or MCP / CLI / API directly.
- Skill is not a mandatory middle tier; execution still lands on tools.
- Rules are not on the call chain; they cross-cut the execution stack—Guides / Permissions in Harness terms.
The goal is not one frozen call chain, but three assets that can stand alone and compose on demand.
Three Flow layers: thin entry, thick execution, platform skeleton
There are three independently configurable workflow layers, with different thickness and ownership:
| Layer | Flow | Thickness | Role |
|---|---|---|---|
| Command | Command Flow | Thin | Entry orchestration: trigger Agent / Skill, pick Profile |
| Agent | Agent Flow | Thick | Goal-driven: plan, call Skill/tools, loop until the Gate |
| Core | AI4SE Workflow | Platform | P0 ~ Ps default / mode-specific flows |
Command chooses which work-item door; Agent chooses how to drive toward a Gate; Core defines stage skeleton and gate bars. Each layer’s workflow is independently configurable—complementary to Loop Engineering: Loop asks what is worth cycling; Harness asks how a single run stays reliable.
Decoupling analogy: Command ≈ Controller
A familiar layering metaphor (kept short):
| ai4se-harness | Analogy | Spring |
|---|---|---|
| Command | → | Controller |
| Agent | → | Service |
| Skill | → | Method |
Design intent:
- Skill is an atomic capability—run alone or orchestrated by an Agent.
- Agent is goal-driven—run alone or re-orchestrated.
- Command may touch a Skill, a single Agent, or an Agents+Skills workflow.
Glue them into one hard-coded pipeline and you are back to maintaining a giant prompt.
Command design: four work items
Commands cover a few stable, memorable entries for common work items. Their job is not to “write code” directly, but to normalize input into a Change Request that can enter P0–P5 and help choose an execution Profile.
| Work item | Command | Typical input | Goal |
|---|---|---|---|
| Requirement | /ai4se:requirement | One-liner, link, product doc, Issue | Proposal / Delta Spec start → clarify and G1 |
| Defect | /ai4se:defect | Repro steps, logs, screenshots, bug link | Defect change: scope impact, then verify-driven fix |
| Hotfix | /ai4se:hotfix | Incident, alert, rollback constraints | Converged flow: risk control, verified fix, archive |
| Tech debt | /ai4se:tech-debt | Module path, refactor goal, quality metrics | Improvement change: benefit, bounds, regression bar |
Examples:
/ai4se:requirement "Support login with phone number"
/ai4se:requirement https://example.com/product/req-123
/ai4se:requirement --profile strict "Support login with phone number"
/ai4se:defect --profile yolo ISSUE-456
/ai4se:hotfix --profile strict "Payment callback timeouts in prod; limited-scope fix"
/ai4se:tech-debt "Refactor order-module state machine; cut branch complexity"
Users may pass only Command + content. The AI recommends a Profile from complexity, blast radius, risk, context sufficiency, and verification availability—and lists alternatives. Recommendation is not mandatory; humans accept or override.
Three Profiles: autonomy, not gate relaxation
Profiles resemble driver-assist levels. They tune AI autonomy and human checkpoints, and map onto the HITL / HOTL spectrum:
| Profile | Analogy | Autonomy | Fit | Human role |
|---|---|---|---|---|
strict | L1 assist | Low | High risk, unclear bounds, prod hotfix | Confirm key decisions; no silent overreach |
copilot | L2 co-drive | Medium | Routine requirements, defects, debt | AI drives the main path; ask at gates / risks |
yolo | L3 conditional auto | High | Low risk, clear bounds, strong verification | Advance under Rules; stop at gates or anomalies |
Recommended or overridden, the change still must meet the gate’s pass criteria. Faster progress ≠ skipping G1/G2/G3.
Gate Agents: goal-driven, not step-complete
Agents are Goal Driven: understand context → plan → call Skill/tools → observe → iterate until the Gate bar is met. Done means gate pass, not “checklist finished.”
Three core Gate Agents map to the end-to-end flow:
| Agent | Span | Target gate | Goal | Typical outputs |
|---|---|---|---|---|
AlignGateAgent | P1 → P2 (G1) | G1 Align | Fuzzy intent → complete, precise, verifiable, non-conflicting Delta Spec | proposal, Delta Spec, acceptance scenarios, G1 report |
VerifyGateAgent | G1 → P4 (G2) | G2 Verify | Spec → design/tasks/implementation verification, faithful to Spec | design, tasks, tests & verify report, G2 verdict |
ArchiveGateAgent | G2 → G3 | G3 Merge & Archive | Merge Delta, archive history, codify reusable assets | updated specs/, archive, learnings & asset suggestions |
Boundary rules:
- Different spans—do not let one Agent both author Spec and self-adjudicate G2.
- Different goals—G1 blocks direction; G2 blocks execution; G3 blocks consistency/codification gaps.
- Different inputs—intent/proposal vs aligned Spec/code vs Delta/verify report/global Specs.
- Different stop conditions—each Agent’s Done is its Gate’s DoV.
- Profiles stack—the same Gate Agent runs under all three Profiles; only autonomy changes.
This differs from the intuition of “six Agents cut by P0–P5”: fine cuts create expensive cross-stage handoffs. Aggregating by Gate goals fits “drive to a gate,” not “report steps.”
On-disk shape: extensible and configurable
Commands
Commands/
├── requirement/
│ ├── workflow.yml # entry orchestration for /ai4se:requirement
│ ├── command.md
│ ├── profiles.yml # strict / copilot / yolo
│ └── res.yml # asset references
├── defect/
├── hotfix/
└── tech-debt/
Agents
Agents/
├── AlignGateAgent/
│ ├── workflow.yml # P1 → P2 (G1) goal-driven loop
│ ├── agent.md
│ ├── goal.yml # gate goal, stop conditions, DoV
│ └── res.yml # Skills / Rules / Tools refs
├── VerifyGateAgent/
└── ArchiveGateAgent/
Skills
Skills/
├── SkillA/
│ ├── References/
│ └── skill.md
└── ...
Conventions: Commands / Agents are extensible; workflows are configurable; Skills stay atomic with low mutual dependence—entangled Skill graphs push debug cost back to giant prompts.
Mini walkthrough: one requirement to G3
Take /ai4se:requirement "Support login with phone number":
- Command Flow: normalize into a Change Request; AI recommends
copilot(routine, clarifiable); user accepts or overrides with--profile strict. - AlignGateAgent (→ G1): clarify scope/non-scope; produce proposal and Delta Spec; no G1 pass → no implementation.
- P3 + VerifyGateAgent (→ G2): design and tasks, then implement and verify; weak local/pipeline evidence stops at G2—not “code written = done.”
- ArchiveGateAgent (→ G3): merge Living Spec, archive the change, suggest Skills/Rules upgrades—write learning back into the User Harness.
The Core stage skeleton never changes. What changes is confirmation density under the Profile, and which Skills / Rules live in this team’s Assets Pack.
Anti-patterns
- Skills without entries and gates—lots of capability, no door, no stop.
- Profiles as gate bypass—
yolobecomes an inspection-free lane. - One mega-Agent for P1–P5—mixed duties, hard to audit, hard to swap per Gate.
- Tightly coupled Skills—atomicity lost; reuse and test cost explode.
- Core mixed into User—changing a business command requires touching the flow engine.
- User Harness grows without governance—conflicting rules/skills; outer-loop Review must cover the Harness itself.
Action checklist
- Lock Core first: does the team accept P0–P5 and G1/G2/G3 pass criteria?
- Ship only the four classic Commands; add entries only after proven frequency.
- Configure
profiles.ymlper Command—default recommendation + explicit override. - Land three Gate Agents; spell Done / DoV in
goal.yml. - Split Skills atomically; reference via
res.ymlinstead of pasting into Agents. - At every G3 retro: which Rules/Skills to upgrade, retire, or create—the Steering Loop changes the Harness, not only the code.
References
- Birgitta Böckeler, Harness engineering for coding agent users, Martin Fowler, 2026
- On this site: Harness Engineering, AI Harness as Function Composition, Project Knowledge Layer, Spec as Source of Truth, Loop Engineering, HITL / HOTL