Repository navigation
Fix tool result images and optional parameters for GPT, Gemini and Ollama - #222
Merged
Merged
Conversation
Both Responses providers sent only the text of a tool result, so screenshots never reached the model. function_call_output.output now becomes a list of input_text/input_image when the result carries images (verified live against the ChatGPT subscription WebSocket endpoint). The WebSocket provider also keeps an image's media type instead of labelling every image as PNG.
Without an explicit strict flag the Responses API normalizes tool schemas into strict mode, which makes every property required. GPT models then filled optional parameters with placeholders (tab_id "", duration 0, even coordinate [0, 0] next to a ref) that the tools cannot tell from real values. Both Responses providers now send strict: false through one shared helper, so every request (including side requests) keeps an identical tool prefix.
Models fill optional parameters with "", [] or null instead of omitting
them: a tab_id "" then fails as an unknown tab, a coordinate [] as a
malformed one. The schema coercion pass now drops such values from
properties the schema does not list as required; 0, false and {} stay,
since they are as likely to be meant. Browser batch steps go through
the same coercion against the inner tool's schema, so a step behaves
like the direct call.
The Vertex/Gemini provider sent only the text of a tool result, so the model never saw screenshots and made up their content. Images now travel as functionResponse.parts with inlineData (multimodal function responses), verified live with Gemini 3.5 Flash.
Ollama messages carry images on any role; tool messages were sent with images: None, so screenshots were dropped. Verified live with a local vision model (qwen3.8:27b).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes two problems the browser tools had with GPT models over the OpenAI Responses providers (seen with "GPT-6.1 Sol (ChatGPT)" on the ChatGPT subscription WebSocket provider). Images in tool results were also dropped by the Gemini and Ollama providers; this PR fixes those too.
Images in tool results never reached the model
Both Responses providers sent only the text of a
ToolResultasfunction_call_output, so screenshots were dropped.outputis now a list ofinput_text/input_imagewhen the result carries images (the Responses API accepts an array there); text-only results stay a plain string. Verified live against the ChatGPT subscription WebSocket endpoint: the model described the screenshot correctly. The WebSocket provider also keeps an image's media type instead of labelling every image as PNG.The same loss happened in two more providers, fixed natively without extra user messages:
vertex.rs): images travel asfunctionResponse.parts[].inlineData(multimodal function responses). Verified live with Gemini 3.5 Flash.imagesfield. Verified live withqwen3.8:27b.In each live check, the model described the probe screenshot correctly with the fix. Without the fix it made up the content.
GPT filled optional tool parameters with placeholders
Without an explicit
strict, the Responses API normalizes tool schemas into strict mode, which makes every property required. The model then senttab_id: "",coordinate: [],duration: 0, and even a made-upcoordinate: [0, 0]next to aref."strict": false. Live, the sameclickrequest came back as just{"ref": "ref_7"}. Every request (including side requests like compaction) renders tools the same way, so the prompt-cache prefix stays identical; running sessions lose their cached prefix once."",[]andnullfrom properties the schema does not list asrequired(0,false,{}stay). This also covers other models and MCP tools.browser_batchsteps now go through the same coercion against the inner tool's schema.Behaviour change:
get_session_contentwithtool_names: []used to hide all tool calls and now shows all of them, which matches the documented "omit for all tools".Not covered
openai.rs(Chat Completions, also used by Groq, Cerebras, Mistral, Moonshot, OpenRouter and Z.ai) still drops images in tool results. Its spec allows only text parts in tool messages, so a fix would need a separate user message; that is left for later.Testing
cargo test --release,cargo clippy --all-targets --all-features -- -D warnings,cargo fmt --checkpassa_hung_page_fails_fastfailed once under full-suite load (timing-based, unrelated) and passed on rerun