Teams bolt on slash commands, skills, and rules for a Coding Agent—then swap models or sessions and watch behavior drift. The usual failure is not “one more prompt missing.” Those assets are not hung on a stable process spine, and they lack clear entry points, autonomy levels, and stop conditions.

A pile of plugins is not a User Harness. Harness Engineering describes the control system; this article is about landing that control system as an extensible tool framework—ai4se-harness.

Thesis in one line: Entry = four work-item Commands × three Profiles; execution = three Gate Agents hung on P0–P5. Method and gates stay fixed; asset packs and human–AI depth vary.

Core spine: P0–P5 and three quality gates

The Core of ai4se-harness does not invent a new methodology. It inherits the stage flow from the SDD end-to-end practice guide. User-state assets hang on this spine:

StageNameKey outputsRelated gate
P0Baseline & initSpecs library, Rules carrier, knowledge dirs in place—
P1Intent & proposalproposal.md (scope / non-scope / assumptions)—
P2Spec definition & alignDelta Spec, acceptance scenariosG1 Align
P3Design & task breakdowndesign, tasks, acceptance design—
P4Implement & verifyimplementation, tests, verification reportG2 Verify
P5Codify & evolveLiving Spec merge, archive, asset improvementsG3 Merge & Archive

The three gates intercept different failure modes at different rework costs:

GateAsksIntercepts
G1 AlignWhat does “correct” look like?Directional errors (edit text; cheapest)
G2 VerifyWas it built to the Spec?Execution errors (edit code; medium cost)
G3 Merge & ArchiveWas learning retained; is the source of truth consistent?Consistency and missing codification (hurts the long term)

Profiles may change confirmation frequency; they must not relax a gate’s DoV (Definition of Verified). For Spec as source of truth, see Spec as Source of Truth.

Zero-layer architecture: User asset packs on Core flows

The zero layer splits the system into two bands: User above holds extensible asset packs; Core below provides SDD flow capability and supports the matching assets upward.

Zero-layer architecture: User Assets and Core Workflow

How to read it:

  1. Layering: User owns the customization surface; Core owns the flow engine and defaults.
  2. One-to-one: Each Core flow maps to a User Assets pack (default flow ↔ standard pack; mode-specific flows ↔ lean / agile / hybrid packs).
  3. Extensible: Assets and workflows can grow sideways without rewriting the engine.

The on-site project knowledge layer covers “what Guides know”; this article covers how whole asset packs bind to the stage flow.

User state: five assets and call / governance boundaries

Align with the usual Coding Agent stack (Claude Code / Cursor / Codex and peers):

User-state structure: Command, Agent, Skill, Tools, and Rules

AssetRoleIndustry counterpart
CommandThin entry: normalize a work item into a flowing Change Requestslash command / command entry
AgentGoal-driven loop: plan → pick Skill/tools → act → observeAgent / sub-agent loop
SkillReusable capability pack (prompts, steps, context, tool usage)Skills / prompt+tool bundles
MCP / CLI / APIConcrete execution surfaceMCP servers, shell, HTTP
RulesCross-cutting governance: decisions, boundaries, tool permissionsCLAUDE.md / .cursor/rules / AGENTS.md

Constraints:

  1. Main path is Command → Agent; Command→Skill direct calls fit only light cases.
  2. Agent is the orchestrator: it may call Skills or MCP / CLI / API directly.
  3. Skill is not a mandatory middle tier; execution still lands on tools.
  4. Rules are not on the call chain; they cross-cut the execution stack—Guides / Permissions in Harness terms.

The goal is not one frozen call chain, but three assets that can stand alone and compose on demand.

Three Flow layers: thin entry, thick execution, platform skeleton

There are three independently configurable workflow layers, with different thickness and ownership:

Three Flow layers: Command / Agent / Core

LayerFlowThicknessRole
CommandCommand FlowThinEntry orchestration: trigger Agent / Skill, pick Profile
AgentAgent FlowThickGoal-driven: plan, call Skill/tools, loop until the Gate
CoreAI4SE WorkflowPlatformP0 ~ Ps default / mode-specific flows

Command chooses which work-item door; Agent chooses how to drive toward a Gate; Core defines stage skeleton and gate bars. Each layer’s workflow is independently configurable—complementary to Loop Engineering: Loop asks what is worth cycling; Harness asks how a single run stays reliable.

Decoupling analogy: Command ≈ Controller

A familiar layering metaphor (kept short):

Spring MVC analogy: Command / Agent / Skill

ai4se-harnessAnalogySpring
Command→Controller
Agent→Service
Skill→Method

Design intent:

  1. Skill is an atomic capability—run alone or orchestrated by an Agent.
  2. Agent is goal-driven—run alone or re-orchestrated.
  3. Command may touch a Skill, a single Agent, or an Agents+Skills workflow.

Glue them into one hard-coded pipeline and you are back to maintaining a giant prompt.

Command design: four work items

Commands cover a few stable, memorable entries for common work items. Their job is not to “write code” directly, but to normalize input into a Change Request that can enter P0–P5 and help choose an execution Profile.

Work itemCommandTypical inputGoal
Requirement/ai4se:requirementOne-liner, link, product doc, IssueProposal / Delta Spec start → clarify and G1
Defect/ai4se:defectRepro steps, logs, screenshots, bug linkDefect change: scope impact, then verify-driven fix
Hotfix/ai4se:hotfixIncident, alert, rollback constraintsConverged flow: risk control, verified fix, archive
Tech debt/ai4se:tech-debtModule path, refactor goal, quality metricsImprovement change: benefit, bounds, regression bar

Examples:

/ai4se:requirement "Support login with phone number"
/ai4se:requirement https://example.com/product/req-123
/ai4se:requirement --profile strict "Support login with phone number"
/ai4se:defect --profile yolo ISSUE-456
/ai4se:hotfix --profile strict "Payment callback timeouts in prod; limited-scope fix"
/ai4se:tech-debt "Refactor order-module state machine; cut branch complexity"

Users may pass only Command + content. The AI recommends a Profile from complexity, blast radius, risk, context sufficiency, and verification availability—and lists alternatives. Recommendation is not mandatory; humans accept or override.

Three Profiles: autonomy, not gate relaxation

Profiles resemble driver-assist levels. They tune AI autonomy and human checkpoints, and map onto the HITL / HOTL spectrum:

ProfileAnalogyAutonomyFitHuman role
strictL1 assistLowHigh risk, unclear bounds, prod hotfixConfirm key decisions; no silent overreach
copilotL2 co-driveMediumRoutine requirements, defects, debtAI drives the main path; ask at gates / risks
yoloL3 conditional autoHighLow risk, clear bounds, strong verificationAdvance under Rules; stop at gates or anomalies

Recommended or overridden, the change still must meet the gate’s pass criteria. Faster progress ≠ skipping G1/G2/G3.

Gate Agents: goal-driven, not step-complete

Agents are Goal Driven: understand context → plan → call Skill/tools → observe → iterate until the Gate bar is met. Done means gate pass, not “checklist finished.”

Three core Gate Agents map to the end-to-end flow:

AgentSpanTarget gateGoalTypical outputs
AlignGateAgentP1 → P2 (G1)G1 AlignFuzzy intent → complete, precise, verifiable, non-conflicting Delta Specproposal, Delta Spec, acceptance scenarios, G1 report
VerifyGateAgentG1 → P4 (G2)G2 VerifySpec → design/tasks/implementation verification, faithful to Specdesign, tasks, tests & verify report, G2 verdict
ArchiveGateAgentG2 → G3G3 Merge & ArchiveMerge Delta, archive history, codify reusable assetsupdated specs/, archive, learnings & asset suggestions

Boundary rules:

  1. Different spans—do not let one Agent both author Spec and self-adjudicate G2.
  2. Different goals—G1 blocks direction; G2 blocks execution; G3 blocks consistency/codification gaps.
  3. Different inputs—intent/proposal vs aligned Spec/code vs Delta/verify report/global Specs.
  4. Different stop conditions—each Agent’s Done is its Gate’s DoV.
  5. Profiles stack—the same Gate Agent runs under all three Profiles; only autonomy changes.

This differs from the intuition of “six Agents cut by P0–P5”: fine cuts create expensive cross-stage handoffs. Aggregating by Gate goals fits “drive to a gate,” not “report steps.”

On-disk shape: extensible and configurable

Commands

Commands/
├── requirement/
│   ├── workflow.yml    # entry orchestration for /ai4se:requirement
│   ├── command.md
│   ├── profiles.yml    # strict / copilot / yolo
│   └── res.yml         # asset references
├── defect/
├── hotfix/
└── tech-debt/

Agents

Agents/
├── AlignGateAgent/
│   ├── workflow.yml    # P1 → P2 (G1) goal-driven loop
│   ├── agent.md
│   ├── goal.yml        # gate goal, stop conditions, DoV
│   └── res.yml         # Skills / Rules / Tools refs
├── VerifyGateAgent/
└── ArchiveGateAgent/

Skills

Skills/
├── SkillA/
│   ├── References/
│   └── skill.md
└── ...

Conventions: Commands / Agents are extensible; workflows are configurable; Skills stay atomic with low mutual dependence—entangled Skill graphs push debug cost back to giant prompts.

Mini walkthrough: one requirement to G3

Take /ai4se:requirement "Support login with phone number":

  1. Command Flow: normalize into a Change Request; AI recommends copilot (routine, clarifiable); user accepts or overrides with --profile strict.
  2. AlignGateAgent (→ G1): clarify scope/non-scope; produce proposal and Delta Spec; no G1 pass → no implementation.
  3. P3 + VerifyGateAgent (→ G2): design and tasks, then implement and verify; weak local/pipeline evidence stops at G2—not “code written = done.”
  4. ArchiveGateAgent (→ G3): merge Living Spec, archive the change, suggest Skills/Rules upgrades—write learning back into the User Harness.

The Core stage skeleton never changes. What changes is confirmation density under the Profile, and which Skills / Rules live in this team’s Assets Pack.

Anti-patterns

  • Skills without entries and gates—lots of capability, no door, no stop.
  • Profiles as gate bypass—yolo becomes an inspection-free lane.
  • One mega-Agent for P1–P5—mixed duties, hard to audit, hard to swap per Gate.
  • Tightly coupled Skills—atomicity lost; reuse and test cost explode.
  • Core mixed into User—changing a business command requires touching the flow engine.
  • User Harness grows without governance—conflicting rules/skills; outer-loop Review must cover the Harness itself.

Action checklist

  1. Lock Core first: does the team accept P0–P5 and G1/G2/G3 pass criteria?
  2. Ship only the four classic Commands; add entries only after proven frequency.
  3. Configure profiles.yml per Command—default recommendation + explicit override.
  4. Land three Gate Agents; spell Done / DoV in goal.yml.
  5. Split Skills atomically; reference via res.yml instead of pasting into Agents.
  6. At every G3 retro: which Rules/Skills to upgrade, retire, or create—the Steering Loop changes the Harness, not only the code.

References