Skip to content

The 2026-09-10 scorecard batch, and a picture of the panel - #10

Merged
kewmine merged 8 commits into
mainfrom
dev
Sep 12, 2026
Merged

kewmine merged 8 commits into
mainfrom
dev

Conversation

@kewmine

@kewmine kewmine commented Sep 12, 2026

Copy link
Copy Markdown
Member

What this changes, and why

The batch that has been sitting on dev since 2026-09-10, plus a screenshot.

Most of it is the scorecard learning to describe a run it previously only
scored. A card that says 43 of 48 and nothing else hides three different ways
of being wrong: a target the clock cut short reports its misses exactly as a
target that finished and failed does, a site that never got a baseline is
counted as a site found clean, and a pass that grew from minutes to 69 minutes
does so without a single number moving. Each of those now has a field that
cannot be read past.

  • truth: score what a run costs, not only what it finds adds seconds and
    requests per target and in total, and names the slowest target instead of
    averaging it away.
  • lab: a run the clock cut short is not a run that found nothing adds
    truncated_targets and recall_is_a_lower_bound to the totals.
  • lab: the scorecard carries how many sites were never tested separates a
    skipped site from a clean one.
  • lab: the scorecard carries where the time went, per class splits the time by
    detection class. File inclusion was half of one run and command injection was
    two percent, which is the reverse of what everyone assumed.
  • lab: keep the 2026-09-10 scorecard puts the run the current backlog came
    from in the repo, because the next card is only meaningful next to it.

Two fixes:

  • lab: run the control plane from the working tree, not the baked copy mounts
    the source read only over /opt/limed. Editing the scorer previously changed
    nothing until somebody rebuilt, and the symptom is that a correct edit looks
    wrong.
  • lab: call a fixture a fixture on the command line too stops the CLI
    reporting androgoat as stopped, which made a healthy lab look like it had a
    target down and made the CLI disagree with the panel and doctor.

And docs: show the panel at the top of the README. The README described every
page of the panel in a table and never showed one. The image is the Targets
page with the light fleet up, so it doubles as a list of what is in the lab.
It is referenced by its raw URL on main, so it renders here once this merges,
and anywhere the README is read outside GitHub.

Checks

  • ./lime audit passes (audit: PASS, 0 warnings)
  • ./lime doctor passes (doctor: PASS, no host-port clashes)

If this changes the panel or the daemon

  • npm run build succeeds in control/ui
  • Checked in a browser, not only in tests: the panel at
    http://127.0.0.1:7000 is what the screenshot in the README is of, served
    by the daemon running from the mounted working tree

A scorecard that records 43 of 48 and says nothing about the pass taking 69
minutes is measuring half the product. Nothing was watching scan time, so it
grew from minutes to over an hour against a 45-endpoint application without a
single number moving.

The card now carries seconds and requests per target and in total, and names
the slowest target rather than averaging it away - an average hides the one
target that took an hour behind eleven that took a minute.

A scanner nobody can afford to run is not accurate, it is theoretical.
The image COPYs lime.py, api.py and scorer.py into /opt/limed, and the
container runs from there. The repo is mounted, but at its own path, so editing
the source changed nothing until somebody remembered to rebuild.

The symptom is the worst kind there is: the code you just changed runs exactly
as it did before, silently, and the obvious conclusion is that the change was
wrong. That cost an hour on 2026-09-11, chasing a scorer edit that was correct
and simply never loaded.

Mounting the source read-only over /opt/limed means a restart is enough. The
image still carries a copy, so it runs standalone where the tree is absent.
A pass truncated by a deadline reports its misses exactly the way a pass
that finished and failed does. Tonight's Mutillidae measurement crossed
the harness's own 3600s limit, and the scorecard would have recorded the
endpoints it never reached as things the scanner cannot detect.

The scorecard now carries `truncated_targets` and
`recall_is_a_lower_bound` at the top of the totals, where they cannot be
read past, and marks the target itself.
43 of 48 in-scope entries, no false positives, on the run that produced
the backlog everything since has been working through. It belongs in the
repo rather than in a shell's scrollback: the next scorecard is only
meaningful next to this one.
A skipped site is not a clean site. Recall computed without saying how
many the engine could not get a baseline for is measuring the load the
pass applied as though it were the tool's ability to find things.
"The scan was slow" is not actionable. "File inclusion was half of it and
command injection was two percent" is, and it is the opposite of what
everyone assumed.
androgoat has no containers because an APK is installed on a device, not
started here. state_of reported that absence as "stopped", so a healthy
lab always looked like it had one target down, and the CLI disagreed
with both the panel and doctor, which have always called it what it is.

Every caller that compares against "running" is unaffected: a fixture was
never running under the old label either.
The README described every page of the control panel in a table and never
showed one of them. The first question anyone asks about a lab they are
deciding whether to run is what it looks like running, and a table of page
names does not answer it.

The image is the Targets page with the light fleet up, so it doubles as a
list of what is actually in the lab. It is referenced by its raw URL on
main rather than a relative path, so it also renders wherever the README
is read outside GitHub.
@kewmine
kewmine merged commit 80b297e into main Sep 12, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant