Versions: docker-agent v1.142.0 and v1.148.0 (linux/amd64, inside a Docker Sandbox), config version 16, provider chatgpt (ChatGPT subscription login), model gpt-6.1-sol.
What happens
Images never reach the model when the agent uses the chatgpt provider:
docker-agent run agent.yaml --exec --attach screenshot.png "Describe the image" → the model answers that no image is attached.
- Images returned by MCP tools (e.g. a screenshot tool whose result carries
images: [{data: <base64 png>}]) are also dropped: the next request's new input is only ~230 tokens.
The model itself accepts images: models.dev lists openai/gpt-6.1-sol with input modalities text, image, pdf, and the same account and model see images through Codex CLI (codex exec -i).
Minimal config
version: '16'
agents:
root:
model: sol
instruction: Describe the attached image.
models:
sol:
provider: chatgpt
model: gpt-6.1-sol
capabilities:
image: true
Debug log (--debug), same with and without the capabilities block:
level=DEBUG msg="Failed to resolve model capabilities for message transforms" model=chatgpt/gpt-6.1-sol error="provider \"chatgpt\" not found"
level=DEBUG msg="strip_unsupported_modalities: stripped media part" kind=image role=user reason="model does not support image input"
level=DEBUG msg="Stripped media content from message" role=user original_parts=2 remaining_parts=1
Analysis (from reading the source)
LocalRuntime.prepareMessagesForModel (pkg/runtime/transforms.go) looks the model up in models.dev as chatgpt/gpt-6.1-sol. models.dev has no chatgpt provider, so catalogModel is nil.
- It then uses
cfg.CapsOverride() (pkg/model/provider/contracts), which reads ModelConfig.Capabilities. Because the image is still stripped with capabilities: {image: true} set, the override appears to be nil for the chatgpt provider by the time the runtime sees it. We did not find where it is lost.
- With no catalogue entry and no override,
providerFallbackCaps returns empty capabilities, and strip_unsupported_modalities removes every image part before convertDocumentToResponseInput (which would have produced a correct input_image) is reached.
Expected
- The
chatgpt provider resolves capabilities from the openai models.dev entry (it is the same model family on the Responses API), and/or
- an explicit
capabilities: override is honoured for the chatgpt provider, as the warning text in modelinfo.warnCapsLookupMiss suggests.
Versions: docker-agent v1.142.0 and v1.148.0 (linux/amd64, inside a Docker Sandbox), config version 16, provider
chatgpt(ChatGPT subscription login), modelgpt-6.1-sol.What happens
Images never reach the model when the agent uses the
chatgptprovider:docker-agent run agent.yaml --exec --attach screenshot.png "Describe the image"→ the model answers that no image is attached.images: [{data: <base64 png>}]) are also dropped: the next request's new input is only ~230 tokens.The model itself accepts images: models.dev lists
openai/gpt-6.1-solwith input modalitiestext, image, pdf, and the same account and model see images through Codex CLI (codex exec -i).Minimal config
Debug log (
--debug), same with and without thecapabilitiesblock:Analysis (from reading the source)
LocalRuntime.prepareMessagesForModel(pkg/runtime/transforms.go) looks the model up in models.dev aschatgpt/gpt-6.1-sol. models.dev has nochatgptprovider, socatalogModelis nil.cfg.CapsOverride()(pkg/model/provider/contracts), which readsModelConfig.Capabilities. Because the image is still stripped withcapabilities: {image: true}set, the override appears to be nil for thechatgptprovider by the time the runtime sees it. We did not find where it is lost.providerFallbackCapsreturns empty capabilities, andstrip_unsupported_modalitiesremoves every image part beforeconvertDocumentToResponseInput(which would have produced a correctinput_image) is reached.Expected
chatgptprovider resolves capabilities from theopenaimodels.dev entry (it is the same model family on the Responses API), and/orcapabilities:override is honoured for thechatgptprovider, as the warning text inmodelinfo.warnCapsLookupMisssuggests.