fix(learn): fail-closed episodic mirror before candidate publish - #71
diazMelgarejo wants to merge 7 commits into
Conversation
Manual stage() previously wrote evidence_ids then fail-open-mirrored into AGENT_LEARNINGS.jsonl, so a mirror write error left dangling evidence. Publish candidate via temp+os.replace only after append_jsonl mirror succeeds. No path_hygiene (PT-local stays out of upstream). Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: cyre <Lawrence@bettermind.ph>
Match evidence on a parsed timestamp. Keep the temp when the mirror line is already readable, and do not treat a read error as corrupt JSON. A directory fsync after rename cannot force another append. Co-authored-by: cyre <Lawrence@bettermind.ph>
Reuse the earliest manual-stage row for a candidate id under the existing episodic flock. Do not add a lock around candidate files. Identical resumable temps publish as one transaction. Co-authored-by: cyre <Lawrence@bettermind.ph>
Recovery must see a complete JSONL generation. has_jsonl_timestamp holds LOCK_EX for the scan and does not create a missing episodic file. A read OSError still propagates so a resumable temp is kept. Co-authored-by: cyre <Lawrence@bettermind.ph>
append_jsonl stays strict so learn.py can fail closed. post_execution and on_failure catch OSError, warn on stderr, and return the entry they built. Co-authored-by: cyre <Lawrence@bettermind.ph>
The same handoff file now lists the three fix commits on the stacked branch, including the hook catch that landed there instead of on PR codejunkie99#71. Both plans point at those SHAs so the next agent does not rebuild them. Co-authored-by: cyre <Lawrence@bettermind.ph>
| _resume_temp(leftovers[0], path) | ||
| for temp_path in leftovers[1:]: | ||
| _remove_or_raise(temp_path) |
There was a problem hiding this comment.
Distinct candidates can be discarded. If two staging calls produce temps with the same candidate ID but different claims or reviewer sources, they share a mirror timestamp. Recovery treats that timestamp as proof that the candidates are identical, publishes the first temp by filename, and deletes the other. A caller can receive a successful stage result even though its candidate content was discarded.
Prompt To Fix With AI
This is a comment left during a code review.
Path: .agent/tools/learn.py
Line: 238-240
Comment:
**Distinct candidates can be discarded.** If two staging calls produce temps with the same candidate ID but different claims or reviewer sources, they share a mirror timestamp. Recovery treats that timestamp as proof that the candidates are identical, publishes the first temp by filename, and deletes the other. A caller can receive a successful stage result even though its candidate content was discarded.
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
There was a problem hiding this comment.
Verified in a01fd1a30e9f9cdae380645c6e628a68ac3e5b55: recovery now compares complete loaded temp payloads, not just their shared evidence timestamp. Distinct claims or reviewers raise an ambiguity error and leave every temp intact, while only identical resumable temps are collapsed and published once. The regression coverage exercises the shared-timestamp/different-reviewer case. This resolves the finding.
Tip: You can customize Greptile's behavior for this repo with .greptile/rules.md and .greptile/config.json.
| File order is the canonical order. Blank lines and corrupt lines | ||
| are skipped. A row with an empty timestamp does not count. | ||
| """ | ||
| for line in raw.decode("utf-8").splitlines(): |
There was a problem hiding this comment.
Malformed bytes block staging. If an older episodic entry contains a non-UTF-8 byte, decoding the whole log raises
UnicodeDecodeError before the corrupt-line handler runs. Manual staging now scans that log, so an unrelated malformed entry stops staging with an uncaught traceback and no candidate is published.
| for line in raw.decode("utf-8").splitlines(): | |
| for line in raw.decode("utf-8", errors="replace").splitlines(): |
Prompt To Fix With AI
This is a comment left during a code review.
Path: .agent/harness/hooks/_episodic_io.py
Line: 59
Comment:
**Malformed bytes block staging.** If an older episodic entry contains a non-UTF-8 byte, decoding the whole log raises `UnicodeDecodeError` before the corrupt-line handler runs. Manual staging now scans that log, so an unrelated malformed entry stops staging with an uncaught traceback and no candidate is published.
```suggestion
for line in raw.decode("utf-8", errors="replace").splitlines():
```
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.There was a problem hiding this comment.
Confirmed fixed in 91c3dff09094b35ef04f9020989ea77d28e7f420. _first_matching_action now parses each JSONL line independently via _json_object_from_line, skips non-UTF-8/corrupt lines without replacement decoding, and the regression test covers a non-UTF-8 line before a valid mirror row. This resolves the malformed-byte staging failure.
Summary
stage()previously wroteevidence_ids: [now]then fail-open-mirrored intoAGENT_LEARNINGS.jsonl, so a mirror write error left dangling evidence.os.replaceonly afterappend_jsonlmirror succeeds (fail-closed).Found and hardened downstream in Perpetua-Tools; contributed back without PT-local
path_hygiene.Test plan
cd .agent/tools && python3 -m unittest test_learn_episodic_mirror -v(3/3)Made with Cursor
Note
Make learn candidate staging fail-closed on episodic mirror errors
learn.stageso the episodic mirror row is written first and durably, and the candidate JSON is published only after that evidence exists. The mirror append is now append-once under an exclusive lock, so retries reuse one canonical timestamp instead of adding duplicate rows.OSErrorfrom the episodic append and warn on stderr instead of failing.learncommand to exit with a failure status, and evidence validation uses exact parsed JSONL timestamp fields instead of raw substring matching, so substring-only matches no longer count as evidence.Macroscope summarized 4b56ff0.
The PR is not yet safe to merge because concurrent recovery can discard candidate content and a malformed episodic entry can stop manual staging.
Reviews (4) · Last reviewed commit: "Merge pull request #3 from diazMelgarejo..."