Skip to content

Fix tool result images and optional parameters for GPT, Gemini and Ollama - #222

Merged
stippi merged 5 commits into
mainfrom
fix/responses-tool-images
Oct 5, 2026
Merged

stippi merged 5 commits into
mainfrom
fix/responses-tool-images

Conversation

@stippi

@stippi stippi commented Oct 5, 2026 •

Copy link
Copy Markdown
Owner

Fixes two problems the browser tools had with GPT models over the OpenAI Responses providers (seen with "GPT-6.1 Sol (ChatGPT)" on the ChatGPT subscription WebSocket provider). Images in tool results were also dropped by the Gemini and Ollama providers; this PR fixes those too.

Images in tool results never reached the model

Both Responses providers sent only the text of a ToolResult as function_call_output, so screenshots were dropped. output is now a list of input_text / input_image when the result carries images (the Responses API accepts an array there); text-only results stay a plain string. Verified live against the ChatGPT subscription WebSocket endpoint: the model described the screenshot correctly. The WebSocket provider also keeps an image's media type instead of labelling every image as PNG.

The same loss happened in two more providers, fixed natively without extra user messages:

  • Gemini (vertex.rs): images travel as functionResponse.parts[].inlineData (multimodal function responses). Verified live with Gemini 3.5 Flash.
  • Ollama: the tool message carries the images in its images field. Verified live with qwen3.8:27b.

In each live check, the model described the probe screenshot correctly with the fix. Without the fix it made up the content.

GPT filled optional tool parameters with placeholders

Without an explicit strict, the Responses API normalizes tool schemas into strict mode, which makes every property required. The model then sent tab_id: "", coordinate: [], duration: 0, and even a made-up coordinate: [0, 0] next to a ref.

  • Opt out of strict mode: both Responses providers render tools through one helper that sets "strict": false. Live, the same click request came back as just {"ref": "ref_7"}. Every request (including side requests like compaction) renders tools the same way, so the prompt-cache prefix stays identical; running sessions lose their cached prefix once.
  • Tool layer: the schema coercion pass drops "", [] and null from properties the schema does not list as required (0, false, {} stay). This also covers other models and MCP tools. browser_batch steps now go through the same coercion against the inner tool's schema.

Behaviour change: get_session_content with tool_names: [] used to hide all tool calls and now shows all of them, which matches the documented "omit for all tools".

Not covered

openai.rs (Chat Completions, also used by Groq, Cerebras, Mistral, Moonshot, OpenRouter and Z.ai) still drops images in tool results. Its spec allows only text parts in tool messages, so a fix would need a separate user message; that is left for later.

Testing

  • New red/green unit tests for all four providers, the coercion pass and batch step parsing
  • cargo test --release, cargo clippy --all-targets --all-features -- -D warnings, cargo fmt --check pass
  • a_hung_page_fails_fast failed once under full-suite load (timing-based, unrelated) and passed on rerun
  • Not yet exercised end to end in the app

stippi added 5 commits October 5, 2026 07:40
Both Responses providers sent only the text of a tool result, so
screenshots never reached the model. function_call_output.output now
becomes a list of input_text/input_image when the result carries images
(verified live against the ChatGPT subscription WebSocket endpoint).
The WebSocket provider also keeps an image's media type instead of
labelling every image as PNG.
Without an explicit strict flag the Responses API normalizes tool
schemas into strict mode, which makes every property required. GPT
models then filled optional parameters with placeholders (tab_id "",
duration 0, even coordinate [0, 0] next to a ref) that the tools cannot
tell from real values. Both Responses providers now send strict: false
through one shared helper, so every request (including side requests)
keeps an identical tool prefix.
Models fill optional parameters with "", [] or null instead of omitting
them: a tab_id "" then fails as an unknown tab, a coordinate [] as a
malformed one. The schema coercion pass now drops such values from
properties the schema does not list as required; 0, false and {} stay,
since they are as likely to be meant. Browser batch steps go through
the same coercion against the inner tool's schema, so a step behaves
like the direct call.
The Vertex/Gemini provider sent only the text of a tool result, so the
model never saw screenshots and made up their content. Images now travel
as functionResponse.parts with inlineData (multimodal function
responses), verified live with Gemini 3.5 Flash.
Ollama messages carry images on any role; tool messages were sent with
images: None, so screenshots were dropped. Verified live with a local
vision model (qwen3.8:27b).
@stippi stippi changed the title Fix browser tools with GPT over the Responses providers Fix tool result images and optional parameters for GPT, Gemini and Ollama Oct 5, 2026
@stippi
stippi merged commit 7060936 into main Oct 5, 2026
5 checks passed
@stippi
stippi deleted the fix/responses-tool-images branch October 5, 2026 07:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant