Skip to content

Evaluations: a run cannot stop at the hand over, so measuring a change of the interview pays for the whole loop #76

Description

@PierreMardon

Status

Closed on 2026-10-05: done by #90, and verified on one run. corpus --stop-at-hand-over plays planning alone. What it saves is the execution, which is a third of a run and not most of it: 0.90 USD of 2.83 on average, over the 17 runs of the three campaigns. Before that, campaigns 2 and 3 cost 2.6 and 3.4 USD per run on average, the whole loop each time, to measure changes that all sit in planning.

What was seen

corpus takes --case, --runs, --jobs and --developer-model, and always plays a run to its end: planning, approval, slices, gates, review. A change that touches only planning, the interview or the blueprint, is measured at the price of the execution too.

Orders of magnitude from campaign 1: a full run costs 1.53 to 4.43 USD at list price, 2.61 on average. In the trial run, the planning session of 01-overdue-list cost 0.78 USD of the run's 1.80.

Sources

Next

Worth adding when a campaign targets planning alone: an option of corpus that stops a run once the blueprint is handed over and the developer has read it, with the measures and the judgements that need no execution (the interview, the blueprint). The first campaign after the fixes needs full runs anyway, since it also reads the hand back at conformity.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions