Python: fix: parse Responses function_call_output so hosted tool results reach transports - #8078
Conversation
…each transports A hosted tool that executes server-side -- for example a Foundry Toolbox dispatching through its generic `call_tool` wrapper -- returns its result as a standalone `function_call_output` Responses item rather than on the originating call item. None of the three parse dispatch sites handled that item type, so it fell through to the `Unparsed ...` debug log and was discarded. No `Content.from_function_result` was produced, and AG-UI consequently emitted TOOL_CALL_END with no matching TOOL_CALL_RESULT, falling back to treating the call as declaration-only. The model still received the real output, so only the client lost the structured result. Add a `function_call_output` branch to all three sites -- the non-streaming `_parse_response_from_openai` and both streaming `response.output_item.added` / `.done` handlers -- sharing one `_parse_function_call_output_content` helper so the lists cannot drift again. `output` is a string or a list of input-content parts, so it is normalized through the existing `_stringify_mcp_output` rather than JSON-encoding provider models. The streaming handlers emit from whichever event first carries a populated `output` and record the item id in a per-request set, so the other event cannot produce a second result. Keyed on the item id rather than `call_id`, which the function-calling loop contract says must not be assumed unique forever. Reviewed against docs/specs/004-python-function-calling-loop.md, which covers provider serialization of function calls and results: no result is orphaned or duplicated, and the streaming and non-streaming paths agree. Fixes microsoft#8068
…parse overrides `RawFoundryChatClient` and `RawFoundryAgentChatClient` override `_parse_chunk_from_openai` to intercept oauth_consent items and then delegate to `RawOpenAIChatClient`. Both had the pre-change signature, so once the base started passing `seen_function_call_output_ids` every Foundry streaming call raised `TypeError: _parse_chunk_from_openai() got an unexpected keyword argument`. Accept and forward the new parameter in both overrides, and update the two delegation assertions that pin the forwarded argument list. Caught against a live Foundry Responses endpoint; `poe test -P foundry` also reproduces it, but that package is not in the validation command list in docs/specs/004-python-function-calling-loop.md even though it subclasses the OpenAI Responses client.
Review follow-up. `Content.from_function_result` does not validate `call_id`, so a `function_call_output` item carrying a blank one produced an orphaned result: transports drop it (`_emit_tool_result` returns early on a falsy `call_id`) and the outbound serializer would re-send it as an unpairable `function_call_output` input item on the next turn. These items are synthesized by the hosting layer, so a blank `call_id` is a realistic host-side defect rather than a theoretical one, and the function-calling loop contract requires that no result becomes orphaned. Extract the emission gate into `_function_call_output_has_result` so the populated-output and pairable-call_id checks are shared by all three dispatch sites instead of being repeated at each one.
There was a problem hiding this comment.
🟡 Changes recommended
Supported older SDK versions can crash, and rich output parts are not serialized into usable results.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Pull request overview
Adds parsing for hosted function_call_output items so tool results reach downstream transports.
Changes:
- Parses streaming and non-streaming function outputs.
- Deduplicates streaming results by item ID.
- Propagates parser state through Foundry clients and adds tests.
File summaries
| File | Description |
|---|---|
python/packages/openai/agent_framework_openai/_chat_client.py |
Implements output parsing and deduplication. |
python/packages/openai/tests/openai/test_openai_chat_client.py |
Tests parsing scenarios. |
python/packages/foundry/agent_framework_foundry/_chat_client.py |
Forwards deduplication state. |
python/packages/foundry/agent_framework_foundry/_agent.py |
Forwards deduplication state. |
python/packages/foundry/tests/foundry/test_foundry_chat_client.py |
Updates delegation assertion. |
python/packages/foundry/tests/foundry/test_foundry_agent.py |
Updates delegation assertion. |
Review details
Suppressed comments (1)
python/packages/openai/agent_framework_openai/_chat_client.py:2692
- For list output, the SDK supplies
ResponseInputText/ResponseInputImage/ResponseInputFilemodel instances, not the dictionaries used in the new test. Text happens to work via.text, but image/file parts fall through tojson.dumps(..., default=str), producing quoted Pydantic reprs (and concatenating multiple reprs) rather than preserving usable rich output. Map these provider parts toContentitems, or serialize the full list with_serialize_provider_payloadwhile retaining its boundaries before creating the function result.
return Content.from_function_result(
call_id=item.call_id,
result=self._stringify_mcp_output(item.output),
additional_properties=additional_properties,
- Files reviewed: 6/6 changed files
- Comments generated: 1
- Review effort level: Balanced
💡 Add a code-review agent skill for context-aware, tailored reviews. Learn more in the docs.
… SDK floor Address review on two counts. `name` is not on `ResponseFunctionToolCallOutputItem` in openai 2.25.0, the declared floor -- that version ships only call_id/id/output/status/type. Reading it as an attribute raised `AttributeError` out of the shared parse helper, which all three dispatch sites call, so on any supported SDK below the release that added the field the whole response parse failed rather than merely dropping the result. Read it with `getattr`. The other attributes touched here (type/status/id/call_id/output) are all present on the floor, and the two module helpers already used `getattr`. `output` may also be a list of input-content parts. Passing those provider models straight to `_stringify_mcp_output` fell through to `json.dumps(..., default=str)` and embedded a Python repr in the result text sent back to the model -- e.g. `"ResponseInputImage(detail='auto', ...)"`. Dump each part first so text extraction still works and non-text parts serialize as readable JSON. Both paths are now regression-tested, including a stub item shaped like the 2.25.0 field set.
|
Reproduced end-to-end on real Foundry infrastructure, not just in unit tests, so the before/after is observable rather than inferred. Setup
The host emits the item as reported Matching AG-UI events, same request, same host, only the parser differing
On Two incidental notes from the exercise, neither part of this change
|
|
Flagging a process point on myself before anyone spends review time here.
Read together with "Any change to the function-calling loop or its approval/history/serialization paths", this PR is inside that scope — it changes how a provider serializer parses a function-call result — and I did not check first. That's on me. Rather than just flag it, here is where the change stands against the seven requirements in the same section, so you can judge cheaply whether it is worth continuing or whether you would rather own it:
The ask. On requirement 1 I would rather not guess at the matrix rows unilaterally, since that is exactly the judgment the clause reserves for you. My reading is that this touches "History and provider serialization" and nothing else, but I may be wrong about whether it also implicates the streaming/non-streaming agreement rows. So: would you prefer to
Happy with any of those. One note in case it affects the call: the underlying gap is not specific to a toolbox — |
|
tagging Eduard van Valkenburg (@eavanvalkenburg) here to look at this. |
…ride signature and rich content Addresses all four open review threads on microsoft#8078. **Back-compat (`_chat_client.py:1`).** `seen_function_call_output_ids` is off `_parse_chunk_from_openai` entirely and travels in the parse `options` mapping under a private key, so the signature is byte-identical to the one released `agent-framework-foundry` wheels override. All four Foundry changes this PR made to accommodate the parameter are reverted. Re-derived the analysis as asked, because the dependency state moved: Foundry's floor rose from `agent-framework-openai>=1.10.0,<2` to `>=1.14.2,<2` in microsoft#8222. That does not close the break -- `agent-framework-foundry` 1.13.0 still overrides the method with the old signature, and `>=1.14.2,<2` still admits a later 1.14.x carrying this change -- so the signature still cannot grow. Two tests pin it: one structural on the parameter list, one behavioural through a subclass that forwards only the four arguments it knows about. **Rich content (`_chat_client.py:2715`).** List-shaped `output` now maps to canonical `Content` items instead of being flattened: text parts to text, image and file URLs to `from_uri` with the media type inferred, `file_id` to `from_hosted_file`, inline base64 to `from_data`. A hosted reference is kept addressable rather than stringified, since only the issuing provider can resolve it and text would destroy that. Parts are read field-by-field because transports and test doubles deliver mappings as well as SDK models, which the previous behaviour handled. **Duplicate serialization (`_chat_client.py:365`).** `_plain_function_call_output` is deleted rather than relocated. The degrade path reuses the existing `_serialize_provider_payload` and `_stringify_mcp_output`, and the mapper now lives beside them in the class instead of being a fourth mechanism. **`Any` (`_chat_client.py:369`).** Concrete SDK types throughout, including a named `ResponseFunctionCallOutputPart` union. Typing them is what surfaced the `Package Checks` CI failure: `call_id` became `Optional[str]` in the SDK bump, so the validated id is threaded from `_pairable_function_call_output_call_id` rather than re-derived where it cannot be narrowed. One behaviour change to weigh: for an output with no text parts, the flat `result` is now empty because `from_function_result` derives it from text items. A consumer reading only `result` and not `items` previously saw JSON there. Each part mapping is wrapped so a content factory rejecting provider data degrades that part to JSON instead of raising out of a streaming parse and failing the request; the old flattening path could not realistically raise.
|
Reworked and pushed as Eduard van Valkenburg (@eavanvalkenburg) — "there have been updates to the dependencies now, so please have another look." So I took Evan Mattson (@moonbox3)'s first option. The per-request seen-id set now travels in the parse I did not take the Eduard van Valkenburg (@eavanvalkenburg) — duplicate serialization handling. You were right, and it was worse than Eduard van Valkenburg (@eavanvalkenburg) — Evan Mattson (@moonbox3) — list-shaped output as canonical Two things I would rather you heard from me than found:
Live testing did settle one thing worth passing on. I flagged in the description that a locally Because the mapping now calls content factories that validate their input, each part is mapped Validation: openai 502, foundry 391, core 5110, ag-ui 1284, declarative 1016, foundry_hosting 292; The ownership question is still open, and it is the one thing I have not been able to answer |
|
Thanks for the contribution. This branch currently conflicts with |
Eduard van Valkenburg (eavanvalkenburg)
left a comment
There was a problem hiding this comment.
Two remaining rich-output round-trip cases need covering before this mapping is complete. Both reproduce with real OpenAI SDK content models on this head; details inline.
Preserve hosted file references and provider image/file semantics during replay. Add 60 streaming and non-streaming rich-output round-trip cases and keep dispatch type-safe at the supported OpenAI 2.25 SDK floor.
…unction-call-output-8068
|
Pushed follow-up fixes in 5b6ce00 and merged current main in 8cadaa4. The merge conflict is resolved and both rich-output review threads are addressed, with 60 parse-to-replay regression cases. The OpenAI 2.25.0 and 3.0.0 dependency-bound jobs now pass, as do code quality, source/test typing, samples/markdown, and coverage. I retried only the Windows 3.12 matrix job after an unrelated workflow-shutdown test hit its one-second timeout. There is a separate base-workflow blocker: |
…unction-call-output-8068
|
Two questions on the latest revision: 1. Rich results reaching AG-UI The parser now preserves image/file parts in Could we add a parser-to-AG-UI regression case and clarify whether displaying attachments is outside this PR's scope? That would distinguish "preserved in framework history" from "visible to the frontend." This is separate from the original text-result fix and the already-addressed OpenAI replay cases. 2. Streaming placeholders and deduplication Can a Could we cover that sequence, or document the provider guarantee that rules it out? This is a contract question, not a claim that this sequence has been observed on live traffic; a genuinely completed empty result should still remain valid. |
Keep streaming placeholders from consuming the final output's deduplication ID. Cover completed empty results and the parser-to-AG-UI text-only boundary, and clarify that frontend attachment display is outside the parser's scope.
|
Jose Alvarez (@jpalvarezl) addressed both points in c4194e4.
The affected core/OpenAI/AG-UI/Foundry suites pass (7,934 tests), and the OpenAI suite plus source typing pass at both SDK 2.25.0 and 3.0.0 (603 tests each). |
|
Thanks for finishing this off, Eduard van Valkenburg (@eavanvalkenburg) — and for pushing the fix rather than sending it back I had validated the mapping thoroughly inbound: unit tests first, then real I went through your commit before this merged and checked the round trips independently rather than Also noted that you hit the Two takeaways I am keeping: a parse change is half a change, so round-trip it; and inference that |
Motivation & Context
A Foundry hosted agent using a Foundry Toolbox dispatches its knowledge-base tool through the toolbox's generic
call_toolwrapper. The Toolbox executes the inner tool server-side andResponsesHostServerserializes the result as a standardfunction_call_outputResponses item. None of the three parse dispatch sites inagent_framework_openaihandled that item type, so it fell intocase _:/ the end of theelifchain, was logged asUnparsed event of type: ...at debug level, and discarded.No
Content.from_function_result(...)was produced, soagent_framework_ag_uiemittedTOOL_CALL_ENDwith no matchingTOOL_CALL_RESULTand its declaration-only fallback took over (_agent_run.py:3100-3108). The model still received the real output — only the client lost the structured result. The same tool called directly is anmcp_calland works, which is what isolates this to the client parser.Description & Review Guide
What are the major changes?
A
function_call_outputbranch at all three dispatch sites — the non-streaming_parse_response_from_openaiand both streamingresponse.output_item.added/.donehandlers — sharing one_parse_function_call_output_contenthelper and one_pairable_function_call_output_call_idgate, so the three item-type lists cannot drift apart on this type again. The gate returns the validatedcall_idrather than a bool, so the parse is handed the narrowed value instead of re-deriving one the SDK types asOptional[str].outputisstr | list[ResponseInputText|Image|File]. The list form maps to canonicalContentitems -- text to text, image and file URLs tofrom_uriwith the media type inferred,file_idtofrom_hosted_file, inline base64 tofrom_data-- so a returned image stays addressable instead of arriving as JSON inside a text item. A hosted reference is kept as a reference because only the issuing provider can resolve it and text would destroy that. Parts are read field-by-field rather than byisinstance, because transports and test doubles deliver plain mappings as well as SDK models. Each part is mapped defensively: a content factory rejecting provider data degrades that part to JSON rather than raising out of a streaming parse and failing the request.One behaviour change to weigh: for an output with no text parts the flat
resultis now empty, becausefrom_function_resultderives it from the text items. A consumer reading onlyresultand notitemspreviously saw JSON there.Per-request dedup state stays off the signature.
_parse_chunk_from_openaiis byte-identical to the signature releasedagent-framework-foundrywheels override; the seen-id set travels in the parseoptionsmapping under a private key.optionsthere is the validated parse context, not the outbound request body -- that is built from a separaterun_optionsdict -- so the key cannot reach the service. Two tests pin the contract: one structural on the parameter list, one behavioural through a subclass that forwards only the four arguments it knows about.Re-derived against the current dependency state, since that was the ask. Foundry's floor rose from
agent-framework-openai>=1.10.0,<2to>=1.14.2,<2in Python: Bump package versions for 1.18.0 release #8222, which does not close the break:agent-framework-foundry1.13.0 still overrides the method with the old signature, and>=1.14.2,<2still admits a later 1.14.x carrying this change. A published constraint cannot be retracted, so the signature cannot grow. All four Foundry changes this PR previously made to accommodate a new parameter are reverted.The two streaming handlers defer outputs explicitly marked
in_progress, including empty or partial.addedplaceholders, without claiming their item IDs. The first eligible output is emitted and its item ID is recorded in a per-request set so the other event cannot emit a second result. Completed empty outputs remain valid. This does not depend on a provider guarantee that empty in-progress placeholders never occur. Deduplication is keyed on item ID rather thancall_id, which the function-calling loop contract says must not be assumed unique forever.Two follow-up commits are separated deliberately so they can be read on their own: forwarding the new parameter through the two
RawFoundryChatClient/RawFoundryAgentChatClientoverrides, and rejecting a blankcall_idso no unpairable result is emitted.What is the impact of these changes?
Hosted-toolbox tool results now reach transports as
function_resultcontent, so AG-UI emits the fullTOOL_CALL_START→TOOL_CALL_END→TOOL_CALL_RESULTlifecycle. No AG-UI change was needed:_emit_tool_resultalready handlesfunction_result.Rich-content scope: preserving image/file parts in framework
Content.itemsand OpenAI replay does not make them visible as frontend attachments. The ordinary AG-UI result emitter reads only the text-onlyContent.result: mixed outputs emit their text, while image/file-only outputs emit empty result content. Parser-to-AG-UI regression cases cover this boundary. Frontend attachment display is a separate transport/frontend capability and remains outside this PR's scope.Reviewed against
docs/specs/004-python-function-calling-loop.md, which covers provider serialization of function calls and results. No result is orphaned or duplicated, and the streaming and non-streaming paths agree. I have not edited the spec — its checklist asks that the matrix name a regression test for each affected scenario, and I would rather you decide whether this warrants a row than edit a cross-package contract in a bug fix. Glad to add one.Verified against a live Foundry Responses endpoint in addition to the unit tests. Two things that measurement settled:
.addedand.done(confirmed forfunction_call,message,reasoning). Emitting from both handlers without the seen-id set would have produced duplicate results; emitting from only one would have been a guess about which event carries the payload.function_call_outputis re-sent inline on the next turn underprevious_response_id, and the service accepts it. Since the outbound serializer emits onlycall_id/type/outputand no item id, a server-generated result is indistinguishable on the wire from a locally-executed one, so this needs no outbound companion change.What do you want reviewers to focus on?
Three things I could not settle myself, all raised deliberately rather than left for you to find:
file_idrepresentation. Aninput_image/input_filepart carrying afile_idbecomesfrom_hosted_filerather than text. I chose addressable over readable: only the issuing provider can resolve it, but flattening it to text destroys the information entirely, whereas a reference lets a consumer decide. Say if you would rather have text and I will change it -- it is one branch.resultabove. It follows from the shape you asked for, but it is caller-visible, so it is your call whether it needs a release note or a text summary of the non-text parts.packages/foundryis missing from spec 004's minimum validation commands (docs/specs/004-python-function-calling-loop.md), even though it subclasses the OpenAI Responses client the spec governs. That omission is exactly why my first sweep missed the regression above;uv run poe test -P foundryreproduces it. Worth adding for the next contributor.Separately, and not proposed here: the
case _:default at all three sites drops any unknown item type at debug severity, which is why the reporter hit three such types in one session (function_call_outputplus two SharePoint preview ones). Silently discarding a tool result is a correctness event rather than a diagnostic one. I deliberately kept this PR to the reproduced type — the SharePoint types are preview and I have no repro, so guessing their shape risks a wrong parser. If useful I will open a separate issue proposing either a shared dispatch table across the three sites or a warning for unknown*_outputitems.Related Issue
Fixes #8068
Contribution Checklist
breaking changelabel (or add "[BREAKING]" to the title prefix, before or after any language prefix) — a workflow keeps the label and the title prefix in sync automatically.