Skip to content

Evaluations: one run per case, so the spread between runs is unknown and the next report cannot say what moved #74

Description

@PierreMardon

Status

Closed: measured. The noise of the judge is known on every criterion. Since campaign 5 every case has two runs or three on one version, for planning and for the execution: 02-reservations, 04-suspension, 05-fine-cap and 06-borrow-limit twice in campaign 5, 01-overdue-list and 03-overdue-reminders three times in campaign 2. The measures and their sources are in the comments below, campaign by campaign.

What was seen

Campaign 1 is 6 runs over 6 cases. The report compares two campaigns by their spreads: "a measure moved when its values lie wholly outside those of the campaign before". With one run per case the spread is a point, so any difference in the next campaign reads as a move.

What one case already shows of the spread: 01-overdue-list was run twice on the same version, in the trial and in the campaign. It drew one diagram, then none; went through two reviews and one fix, then one review and none.

Sources

Why it matters

Every other issue of this campaign that says "not confirmed" says so partly for this reason.

Next

Three runs each of 01-overdue-list and 03-overdue-reminders on the chain once the first fixes are merged: the two cases where the interview asked nothing, so the same runs measure the fix of the interview. The spread of the four other cases stays unknown until a campaign pays for it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions