Skip to content

Latest commit

 

History

48 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

CausalLift

Leakage-controlled Skill optimization and paired evaluation for SkillsBench / BenchFlow.

CausalLift is an engineering project for developing reusable agent Skills without turning public benchmark tasks into an answer bank. It freezes task identity and data splits, records execution provenance, compares no_skill and with_skill runs explicitly, and treats candidate Skill changes as hypotheses that must survive regression and generalization checks.

Competition status: the project was submitted as a competition entry; results are still pending. This README reports repository-verifiable engineering properties and does not claim any placement or award.

Problem

Public benchmark tasks are useful for diagnosing agent failures, but optimizing directly against them creates an obvious failure mode: a Skill can appear to improve while actually encoding task-specific trajectories, artifacts, or answers.

CausalLift therefore treats public tasks as a trajectory laboratory, not an answer library. A rule is only worth keeping when it can be expressed as a reusable procedure, invariant, or routing boundary that remains meaningful across instances and task classes.

The project separates three concerns:

  1. Data / leakage control — build a stable task map and freeze development, holdout, and final-blind membership.
  2. Evaluation integrity — make execution identity, task identity, Skill identity, conditions, and result provenance explicit.
  3. Skill generalization — convert observed failures into minimal reusable procedures rather than public-task-specific patches.

Evaluation design

flowchart LR
    A[SkillsBench task roster] --> B[Task map + exposure ledger]
    B --> C[Frozen 60 / 17 / 10 split]
    C --> D[Evaluation Protocol v1 preflight]
    D --> E1[no_skill run]
    D --> E2[with_skill run]
    E1 --> F[Run records + provenance]
    E2 --> F
    F --> G[Explicit paired comparison]
    G --> H[Candidate review / regression gate]
    H --> I[Generalize, reject, or keep]
Loading

Frozen benchmark identity

Item Frozen value
SkillsBench roster skillsbench@1.1
Upstream snapshot 9a1f4dd5f7659f75707435da3ce854b6e48321d1
Split algorithm joint-stratified-v1
Dev 60 tasks
Holdout 17 tasks
Final blind 10 tasks
BenchFlow 0.6.6
Sandbox Docker

data/splits.json stores the membership plus metadata and exposure-ledger digests. scripts/evaluation_protocol.py independently checks the frozen identity and membership digest before evaluation.

What the protocol records

Each physical execution attempt is represented by a RunRecord containing, among other fields:

  • task ID and task-content digest;
  • split and trial index;
  • no_skill / with_skill condition;
  • model, provider, agent, BenchFlow, and sandbox identity;
  • Skill-content digest and optional Git commit evidence;
  • command, timing, result, config, trajectory, stdout, and stderr paths;
  • runtime Skill evidence;
  • reward, run status, and protocol error.

Paired comparisons are represented separately through PairRecord, so a lift is tied to two explicit run identities rather than inferred from unrelated aggregate scores.

The runner also requires an explicit experiment_id, writes an experiment manifest, and stores config and split provenance. Reusing the same experiment identity therefore has a concrete on-disk meaning instead of silently creating an unrelated run history.

Candidate Skill strategy

The current frozen candidate focuses on artifact-producing tasks, where success depends on both semantic correctness and preservation of document-native structure.

Document artifact workflow

skills/document-artifact-workflow/SKILL.md encodes a reusable procedure:

  1. establish the authoritative input, output path, requested semantic change, and preservation requirements;
  2. choose the highest-level representation that can express and verify the required invariant;
  3. inspect only structurally relevant regions;
  4. make the smallest sufficient mutation;
  5. checkpoint the canonical artifact;
  6. reopen it independently and verify the task-relevant postconditions;
  7. escalate to lower-level diagnostics only when evidence shows the current representation is insufficient.

The important constraint is that verification must establish the requested observable property; the Skill does not claim visual or structural correctness that available evidence did not prove.

Spreadsheet artifact workflow

The spreadsheet workflow applies the same principle to workbook semantics: classify the requested change, preserve unrelated workbook structure by default, and use targeted inspection / verification helpers when formulas, references, structure, or preservation requirements make a direct edit insufficient.

Generalization discipline

Candidate development follows a simple rule:

A public failure can motivate a hypothesis, but it cannot become a permanent rule merely because one public task failed.

Observed behavior is classified into reusable procedures and invariants or rejected as task-specific. This is why the repository preserves rejected and frozen candidate histories separately: failed ideas are evidence for model selection and causal attribution, not content to hide by rewriting history.

Validation evidence

The frozen A2.2 candidate is preserved at commit:

cfd026d4d7d7b3bcbfc2e6ca4b03df451fc93883

Its checked-in experiment log (docs/experiments/a22-codex-progress.md) records:

  • baseline full suite: 167 passed;
  • focused acceptance suite: 31 passed;
  • final full suite: 167 passed;
  • no forbidden runtime residue;
  • no Skill __pycache__ residue;
  • helper baseline diff clean;
  • git diff --check clean.

These numbers describe the recorded frozen-candidate validation run. They are not presented as a fresh CI result from the current README-only branch.

The test suite covers task-map construction, evaluation protocol invariants, runner lifecycle behavior, result summarization, document and spreadsheet Skill contracts, and candidate-routing / procedure constraints.

Repository structure

causallift/
├── configs/evaluation/        # pinned evaluation configuration
├── data/
│   ├── splits.json            # frozen 60 / 17 / 10 membership
│   ├── task_map.csv           # leakage-controlled task map
│   └── preexisting_exposure.json
├── docs/
│   ├── experiments/           # candidate and validation evidence
│   └── superpowers/           # protocol contracts, specs, and plans
├── scripts/
│   ├── build_task_map.py      # task-map / split construction
│   ├── evaluation_protocol.py # frozen identities and record schemas
│   ├── run_evaluation.py      # BenchFlow execution lifecycle
│   └── summarize_results.py   # paired result aggregation
├── skills/                    # reusable candidate Skills
├── tests/                     # protocol and Skill regression tests
├── pyproject.toml
└── uv.lock

Local static verification

Requires Python 3.12+ and uv.

uv sync --frozen
uv run --frozen pytest -q

A real BenchFlow evaluation additionally requires the pinned SkillsBench task snapshot and the configured runtime/provider credentials. The example config expects the SkillsBench task checkout at ../skillsbench/tasks.

A single evaluation request is explicit about experiment, task, condition, and trial:

uv run python scripts/run_evaluation.py \
  --config configs/evaluation/opencode_go_v4_flash_dev.yaml \
  --experiment-id example-citation-no-skill-t0 \
  --task-id citation-check \
  --condition no_skill \
  --trial-index 0

For a paired with_skill run, use the same task/trial identity with --condition with_skill and provide --candidate-skill-root.

Engineering focus

The main engineering work in CausalLift is not the prompt text of an individual Skill. It is the control system around Skill iteration:

  • stable benchmark and split identity;
  • leakage-aware development boundaries;
  • explicit execution and artifact provenance;
  • paired condition comparison;
  • deterministic preflight and protocol gates;
  • regression tests for Skill routing and procedures;
  • preservation of rejected, frozen, and experimental candidate history.

That makes Skill changes reviewable as engineering changes rather than one-off benchmark prompt tuning.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages