Repository navigation
Conversation
… no MCP) Ports the decision-policy core from PR thruwire#14 onto the responsibility architecture: tighten-only abstention policy, ForemanModel.decide(), FOREMAN_DECIDED audit events, and a first-class DecisionPolicyResponsibility registered alongside the builtins with a packaged TOML definition. Drops all MCP stdio server code: no foreman mcp, no mcp package, no docs/mcp.md. The decision path is a single async seam, resolve_decision(), kept lifecycle-agnostic (no hard-coded BEFORE_TOOL); the hook point remains an open maintainer question. Abstention routes to the human via ESCALATE and never terminates the worker.
|
I noticed two lifecycle gaps that make this look premature to merge even though the tests and Ruff pass:
Could we first choose the lifecycle event/contract, wire this through the production serve path with an end-to-end test, and either add a distinct non-terminal human-routing outcome or constrain the integration so ESCALATE cannot terminate the session? |
|
Thanks, both gaps are fair. Here's what I'd propose for the rework: Hook point: Abstention outcome: a distinct non-terminal human-routing outcome rather than reusing Arming: scope the If |
|
Thanks for waiting on the architectural answer. I think these should be two separate seams:
The existing hook contract cannot inject a model-selected answer back into the worker, and |
|
Thanks — the two-seam split makes sense. Happy to rework this as the focused integration once the non-terminal routing outcome and decision response contracts land. Would you like me to leave #25 open until then, or close it and reopen fresh against the new seams? |
|
Opened #42 to propose the non-terminal human-routing outcome and decision response contract discussed above, so we can agree on the seams before this gets reworked. |
|
I think this feature is showing promise. A worker being able to ask its supervisor a question feels relevant to Foreman. What would make it compelling is the supervisor bringing its existing responsibilities, configuration, and evidence to the answer. If it mostly proxies a caller-supplied question and options to Jev, with a separate threshold and risk policy, I’m less convinced that Foreman adds enough value. Could this integrate more directly with existing capabilities? A few possible starting points:
Ideally, these requests would use the normal responsibility routing, check definitions, thresholds, and supervisory rules, with the decision recorded in the audit trail. That would also help clarify whether a separate decision policy is needed. Would you be interested in proposing one focused, end-to-end workflow along those lines—showing how the worker submits the question, how existing Foreman capabilities determine the answer, and how the worker receives it? I think that would give us a stronger foundation for agreeing on the lifecycle and response contracts. |
Follow-up to #14, per your direction on that thread. Drops the stdio
ask_foremanMCP tool entirely (nomcppackage, docs, CLI command, or tests remain) and keeps the reusable core as a serve-side capability aligned with the responsibility architecture: bounded decision models, the tighten-only abstention policy, and the audit trail.What this adds
supervision.decision-policyresponsibility (src/foreman/responsibilities/decision.py+ packagedsupervision.decision-policy.toml): contributes thedecision-requiredcheck (floor 0.70, always armed in the packaged TOML), registered inbuiltin_registry()alongside the core responsibilities.src/foreman/models/decision.py):AbstainCategory(destructive/irreversible/external/credentials),DecisionRequest,Decision.src/foreman/decision_policy.py): server baseline merged with the request tighten-only (max of thresholds, union of denylists — a caller can never loosen the floor). An answer is returned only if the model didn't abstain, confidence >= threshold, the choice is one of the options, and the classification doesn't intersect the denylist.ForemanModel.decide()(base.pyprotocol, Jev Noul implementation,FakeForemanModeldeterministic stub) — now called by the responsibility/serve path instead of the MCP tool.FOREMAN_DECIDEDevents (models/events.py) appended to.foreman/runs/<session-id>/events.jsonl(persistence.py) with question, classification, choice, confidence, and effective policy;foreman inspectrenders them.FOREMAN_DECISION_THRESHOLD(default0.80) andFOREMAN_DECISION_ALWAYS_ABSTAINvia the existingFactoryConfigenv conventions.resolve_decision()— model verdict → tighten-only policy gate → audit →ESCALATEdirective on abstain. Abstention routes to the human; it never terminates the worker.directives()stays a synchronous no-op so no hook point is hard-coded.tests/test_decision_policy.pycarried over; newtests/test_decision_responsibility.pycovering route/checks/directives and the abstain-never-terminates invariant. 191 passed; Ruff clean.Open questions
resolve_decision()—BEFORE_TOOL(gate a risky tool call) or a dedicated decision event?ESCALATEthe right ask-human representation, or should serve gain a distinct non-terminal human-routing directive?destructive/irreversible/external/credentials)?Follow-up to #14 and #13.