fix: don't flag an in-flight parallel duplicate as crashed (port of #443) - #462
Merged
Merged
Conversation
Port of #443, which landed on dev. V1.4.3 carries the identical guard code and therefore the identical bug: execute_parallel dispatches a whole batch before any call completes, so two identical irreversible calls issue two begin()s with no complete() between them. The second saw the first's INTENT row, read it as an interrupted prior attempt, downgraded it to FAILED and told the model the action "may have already happened" while the first was still running. Age cannot separate the two cases -- an immediate crash-then-restart leaves an equally young INTENT row, and that one does need the warning. Tracking in-flight keys on the guard instance distinguishes them, because a restart builds a fresh guard with an empty set. Verified on V1.4.3, not assumed: the new test fails against this branch's parent and passes with the fix. Full suite 1259 passed, 0 failed. The author's noted caveat -- a key stuck in the in-flight set if complete() never runs -- is narrower than stated here: manager.py wraps execution in try/except and calls complete() after it, including on the cancellation path, so only a BaseException outside Exception/CancelledError or process death could leak, and process death clears the set anyway.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Port of #443 (merged to
dev) onto V1.4.3, which carries the identical guard code and the identical bug.execute_paralleldispatches a whole batch before any call completes, so two identical irreversible calls issue twobegin()s with nocomplete()between them. The second reads the first's INTENT row as a crashed prior attempt, downgrades it to FAILED, and tells the model the action "may have already happened" — while the first is still running.Verified on V1.4.3 rather than assumed:
app/triggers/activity_log.pyon V1.4.3 is byte-identical todev's pre-fix version, so the same bug and the same patch applyOne note on the author's own caveat (a key stuck in
_in_flightifcomplete()never runs): it's narrower than they feared.agent_core/core/impl/action/manager.pywraps execution intry/except Exception, setsstatus = "error", and callscomplete()after the block — including on theCancelledErrorpath, which re-raises only after persistence. So a leak needs aBaseExceptionoutsideException/CancelledError, or process death, which clears the set anyway.