ZCode may classify an auto-discovered custom OpenAI-compatible model as text-only even when the underlying model and gateway support Vision.
After attaching a screenshot, ZCode responds with a message similar to:
I'm unable to view images. The selected model does not support image input.
ZCode stores custom-provider metadata in:
~/.zcode/v2/config.json
An auto-discovered model may be recorded with an explicit text-only input modality:
{
"modalities": {
"input": ["text"],
"output": ["text"]
}
}Setting only "supportsImages": true is not enough. ZCode uses modalities.input as the effective capability gate, so the working model entry needs both declarations:
{
"modalities": {
"input": ["text", "image"],
"output": ["text"]
},
"supportsImages": true
}First verify that the model accepts an OpenAI-compatible multimodal request through the same gateway used by ZCode. A valid message uses a content array containing text and image_url parts:
{
"model": "glm-5.2",
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": "Describe this image."},
{
"type": "image_url",
"image_url": {"url": "data:image/png;base64,..."}
}
]
}
]
}If this request fails, fix the model/gateway path first. This ZCode patch only corrects client-side capability metadata.
- Quit ZCode completely (
Cmd-Qon macOS; Quit/Exit on Windows or Linux). Closing only the window is insufficient because ZCode caches model capabilities. - Download
fix-zcode-custom-model-vision.pyfrom this gist. - Run it with the model ID used by the custom provider:
python3 fix-zcode-custom-model-vision.py glm-5.2If multiple custom providers contain the same model ID, list the candidates:
python3 fix-zcode-custom-model-vision.py glm-5.2Then rerun with the intended provider UUID:
python3 fix-zcode-custom-model-vision.py glm-5.2 --provider-id PROVIDER_UUIDThe script:
- ignores ZCode's built-in providers;
- creates a timestamped backup;
- preserves credentials and unrelated settings;
- atomically updates the configuration;
- verifies the saved result;
- never prints API keys or provider options.
- Reopen ZCode.
- Select the patched custom model.
- Start a new task/chat.
- Attach a screenshot.
- Ask for objectively verifiable OCR, such as:
Transcribe the largest visible heading exactly.
A successful HTTP request alone is not enough—the answer must demonstrate that the model saw the image.
The script prints the backup path, for example:
~/.zcode/v2/config.json.before-image-modalities-20260728-170000
Quit ZCode and restore that file over ~/.zcode/v2/config.json.
Editing or refreshing the custom provider—or upgrading ZCode—may regenerate the model metadata and restore "input": ["text"]. If screenshot support disappears, inspect the model entry and rerun the patch.
The durable upstream fix is for ZCode to preserve custom input modalities, expose them in the custom-model UI, or import reliable capability metadata during model discovery.