Skip to content

Test-First Evidence

Test-first development (TDD) establishes a behavioral claim before implementation, then shows that the implementation satisfies that same claim. In codeArbiter it is a gated procedure, not a promise to add tests near the end of a feature.

The important observation is the contrast. The test should distinguish the incorrect behavior from the required behavior. A test that only repeats an implementation detail, asserts on its own mock, or fails because its module cannot load has not yet established that contrast.

On the full feature lane, the specification and plan exist first. executing-plans delegates a scope to subagent-driven-development, which selects a fresh author for a task. That author follows TDD inside the task. Afterward the task engine performs specification review and fresh verification, then the combined scope receives its quality review and acceptance treatment.

The first visible chapter below shows the task loop. Use the other chapters to follow its caller back to the definition and forward through commit and PR preparation. Roles are rows, not phases.

Follow the work, not a catalog

A feature request becomes a reviewed pull request

Attended full-lane example, choosing Open a PR. Read left to right within each chapter, then continue downward. These are source-contract relationships, not a captured run or installed-host certification.

Rows identify roles. Gold commands enter or define a procedure; blue skills coordinate work; green agents perform bounded roles. Arrows cross rows in execution order. A dashed arrow reuses a procedure rather than calling its entry again. Dashed nodes are path-selected reviewers, not mandatory roles for every change.

Chapter 1 / 4

Define the change

What behavior is being agreed to?

Link to this chapter
  1. command
    /ca:feature

    Resolve existing work before classifying a new request. This map follows new full-lane work; small and resume routes are separate.

    Result: A selected lane and exact artifact route.

    full lane → New full-lane work enters brainstorming.

    Source for step 1
  2. skill
    brainstorming

    Turn the request into concrete acceptance criteria. Surface unresolved questions and obtain the required specification approval before planning.

    Result: The reviewed specification, not implementation code.

    approved → Specification approval precedes implementation planning.

    Source for step 2
  3. skill
    writing-plans

    Map the specification into tasks, paths, dependencies and verification. Preserve the selected format and its approval binding.

    Result: An executable plan only when its required bindings are current.

    ready plan → The attended coordinator receives the approved, bound plan.

    Source for step 3
  4. skill
    executing-plans

    Coordinate attended work. On HTML, use authored checkpoint membership rather than inventing smaller acceptance scopes.

    Result: One complete scope delegated to the task engine.

    Source for step 4

Before continuing: A specification, a linked plan, and the current approvals required by the selected artifact format.

Continuations, returns and repeat paths
  • Delegate the complete scope. The attended coordinator calls the task engine with the authored scope, then waits for its return. Next: subagent-driven-development: select. Owning contract

Chapter 2 / 4

Build and verify each task

Who writes, and what proves this task?

Link to this chapter
  1. skill
    subagent-driven-development: select

    Select eligible work inside the caller’s scope and retain its requirement, target paths and verification command. Typed work uses engine eligibility and a current context ticket.

    Result: One eligible task, dispatched with bounded context.

    dispatch → One fresh, scope-selected author receives the task.

    Source for step 5
  2. agent
    backend-author, frontend-author or infra-author

    Select the author from the project’s scope mapping. The author runs the test-first procedure inside this task; authoring is not a completed stage before TDD.

    Result: One fresh author context, not a persistent team of simultaneous writers.

    uses tdd → The author follows the test-first procedure, not a separate later test service.

    Source for step 6
  3. skill
    tdd: six gated phases

    Derive obligations, observe a meaningful failing test, implement the minimum change, then verify coverage of the obligations and project quality requirements.

    Result: Test-first work returned to its caller, not permission to commit.

    Inspect all six TDD phases
    1. Obligation scan: identify each verifiable claim.
    2. Red: the new test fails for the intended behavior.
    3. Green: satisfy the same assertion with the minimum change.
    4. Obligation verify: every claim has meaningful passing evidence.
    5. Coverage: apply the declared threshold or cited no-tooling treatment.
    6. Lint: pass the declared lint and applicable type checks.

    return → The caller performs specification review and fresh verification after authoring.

    Source for step 7
  4. skill
    subagent-driven-development: review and verify

    Check the change against its task obligation and run the plan’s verification afresh. An author’s report does not substitute for this evidence.

    Result: The task’s review and verification results; typed REVIEW remains provisional.

    Source for step 8

Before continuing: Per-task specification review and fresh verification. This is not yet whole-scope acceptance.

Continuations, returns and repeat paths
  • More tasks in this scope. Build the remaining eligible tasks. Combined quality review occurs after task specification review and fresh verification, despite its earlier phase number. Next: subagent-driven-development: select. Owning contract
  • Scope ready for quality review. Proceed only once every task has the required per-task evidence. Next: Scope-selected quality reviewers. Owning contract
  • Missing obligation returns to red. Phase 4 returns to Phase 2 for a meaningful failing test, then reruns implementation. Tests are not weakened to clear a gate. Next: tdd: six gated phases. Owning contract

Chapter 3 / 4

Accept the scope, then checkpoint

What returns to the human coordinator?

Link to this chapter
  1. agent · if applicable
    Scope-selected quality reviewers

    Review the combined diff once per scope after task review and verification. Dispatch security, auth/crypto, dependency or migration reviewers only where applicable; ordinary work retains its TDD/coverage quality bar.

    Result: Findings against the combined scope, not a blanket reviewer roster.

    findings → Applicable reviewer findings are classified before the scope can be accepted.

    Source for step 9
  2. agent · if applicable
    finding-triage

    Classify the dispatched findings. Security CRITICAL stops the loop; HIGH returns the affected work for correction. The active caller owns lower-severity disposition.

    Result: A severity-aware result with out-of-scope work kept separate.

    no blocker → Return quality results to the task engine; current evidence still governs acceptance.

    Source for step 10
  3. skill
    subagent-driven-development: accept

    Accept only after task proof and combined quality review. The HTML path records current whole-scope acceptance through the installed authority and engine, not an edited checkbox.

    Result: A completed scope returned to executing-plans.

    to caller → A scoped invocation returns to executing-plans, not directly to commit-gate.

    Source for step 11
  4. skill
    executing-plans: human checkpoint

    Report what landed, what is next and what remains open. Wait for acknowledgement. This attended stop is different from a periodic checkpoint report and from an HTML acceptance scope.

    Result: Acknowledgement, then another scope or the commit gate.

    Source for step 12

Before continuing: Accepted scope evidence and a human acknowledgement before the next attended batch.

Continuations, returns and repeat paths
  • HIGH finding requires correction. Correct the affected task, repeat its proof and rerun the combined quality review. Real authority/security blocks remain stops. Next: backend-author, frontend-author or infra-author. Owning contract
  • Next acknowledged scope. The coordinator delegates the next scope only after human acknowledgement. A changed typed partition uses the owning amendment path. Next: subagent-driven-development: select. Owning contract
  • Final scope acknowledged. Current acceptance of the entire selected plan is checked again before commit handoff. Next: commit-gate. Owning contract

Chapter 4 / 4

Commit and open the PR

Which action is actually authorized?

Link to this chapter
  1. skill
    commit-gate

    The coordinator hands off only after all required scopes and acknowledgements. The commit gate checks permission, selected work and fresh evidence; /ca:commit is also a standalone entry, not an extra command this path must type.

    Result: A selectively staged commit after the applicable gates.

    committed → The finishing procedure requires a cleared commit gate on current work.

    Source for step 13
  2. skill
    finishing-a-development-branch

    Present the supported terminal choices with actual state. This illustration follows the user choosing Open a PR; merging via PR and confirmed discard are separate outcomes.

    Result: An explicit terminal choice, with no direct write to main.

    reuse → Execute the PR contract in place; do not re-invoke its command.

    Source for step 14
  3. command
    /ca:pr procedure, reused by finishing

    Finishing executes the PR procedure here. The command is not re-invoked: routing back into it would create a loop. A standalone /ca:pr remains an entry to the same contract.

    Result: The required PR preparation and review, not a new authorization.

    review → PR creation includes its own path-selected review and verdict.

    Source for step 15
  4. agent
    PR reviewers, finding-triage and verdict-aggregator

    Apply the PR path matrix and read-only review funnel. Resolve blocking findings before opening. The final handoff names the PR URL, current head and actually observed checks.

    Result: An open pull request. Exact-head merge readiness and merge permission remain separate.

    Source for step 16

Before continuing: The chosen Open a PR outcome, with current source and evidence. The PR is not an automatic merge or release.

Continuations, returns and repeat paths

    An open PR, its source identity and observed check state. Merging and publishing remain separate decisions.

    Actual endpoint: An open PR, its source identity and observed check state. Merging and publishing remain separate decisions.

    Command labels use Claude Code spelling for the illustrated route. Codex and Pi use their own entries; shared names do not establish equal execution or authority support. Check the installed-host boundary.

    Source identity and diagram limits

    Reviewed source: 29f84168ab7592a5d7bb0550cb1f25cf29608a9d. The numbered stages explain one ordinary success path. The return notes retain the caller’s conditions and alternate paths. A conditional role participates only when its owning contract selects it. Existing Markdown, small-lane and sprint paths retain their separate contracts.

    Open the complete three-row diagram. The reading view and image are generated from the same editorial map. They neither run tools nor record approvals.

    The small feature lane enters TDD after its confirmed inline criteria. A confirmed fix has a regression-first entry. Refactoring has a behavior-preservation contract and must not be treated as permission to change old assertions. Those paths are not all identical to this full-feature illustration. See gated lanes for the distinction.

    An obligation is a verifiable claim with an ID, a source and a state. A specification criterion is one source; contract invariants and applicable security controls can add others. “Test the export” is too vague. “Preserve the input order in the output rows” is a claim a test can challenge.

    The procedure records the obligations before implementation. Spec-derived obligations already covered by approval should not trigger a redundant full reapproval. Beyond-spec obligations need the owning lane’s decision treatment. Genuinely missing facts remain explicit rather than being filled with an invented command or expected result.

    A plan references the specification’s stable criterion IDs. It does not create another independent criterion list. The artifact model explains why that ownership matters.

    Each new test maps to an obligation and fails for the right reason. Existing tests should remain green at this point; an unrelated regression is a separate problem to surface.

    For example, the saved-search fixture requires non-empty input to produce CSV text. Its baseline returns None. The recorded red output reports two failing assertions for that missing output and one passing empty-input test. This is not a demonstrated quoting-only defect: the baseline does not yet produce non-empty CSV at all.

    A missing import would establish only that the test environment is incomplete. A test that never executes the target would establish even less. Diagnose those problems before describing the result as the behavioral red phase.

    Implement the minimum change that satisfies the behavioral obligation. Preserve the assertions between red and green; fixture or setup corrections must not weaken the expected behavior. The project’s affected existing tests must still pass.

    In the same fixture, the completed serializer passes all three tests with the same test-file digest. The result is useful, bounded local evidence. It does not prove a codeArbiter host session ran, a human approved the example, or another application is correct.

    A green run against a weakened assertion would not establish the intended correction. “We changed the expected output to match the implementation” is a changed requirement, not a successful proof of the original one.

    Review the map from claims to passing evidence. A test file can exist and still miss the seam that matters. A MISSING obligation returns to red, then implementation, rather than disappearing from the list.

    For security-relevant or contract-critical logic, the TDD contract dispatches coverage-auditor to challenge whether tests exercise the claimed behavior. This is a conditional role, not a universal extra author or a substitute for the later scope review.

    A useful objection names the consequence: an untested export failure could deliver malformed records to a caller. A bare percentage or a generic “needs tests” comment does not identify that risk as clearly.

    Phases 5 and 6: coverage and quality tools

    Section titled “Phases 5 and 6: coverage and quality tools”

    Coverage measures exercised lines and branches under the declared tooling and maturity rules. It does not decide that an assertion is meaningful. Both required measures must satisfy the applicable contract; a high line percentage cannot hide a failing branch measure.

    Where the surface has no coverage tooling, the documented no-tooling treatment requires a citation to that actual project configuration. “The command was not found” is not equivalent to “this surface has no coverage command.” For platform-forked behavior, state what environment and any required union were measured rather than treating one host’s result as universal.

    Finally, run the project-declared lint and applicable type checks. Their success is a separate quality observation. They do not replace the obligation proof or grant approval for unrelated work.

    The task engine compares the implementation with the specification, runs the task verification afresh, and applies the required scope quality review. Source phase numbers should not be mistaken for a naive numeric execution list: the combined quality review explicitly waits until the tasks’ specification review and fresh verification have completed.

    Typed work then needs the actual current authority records and whole-scope acceptance. The attended coordinator receives the scope result and waits for its human checkpoint. Commit and PR have their own evidence and permission requirements. A successful author report skips none of them.

    Use Investigate and fix to work from a symptom into a regression, and Review and ship for the final handoff. The exact six-phase contract remains in tdd; the map above is an explanation, not its replacement.

    Pick one current criterion. Name the observation that should fail without the required behavior, show why the actual failure is that observation rather than test setup, and compare the assertion before and after the correction. Then name the review, acceptance and delivery evidence that still has to follow. That complete answer is stronger than “the suite is green.”