# Vision

> Send images alongside text to multimodal models.

Models whose input modalities include **image** accept images in user messages, in the OpenAI content-parts format.

```python
completion = client.chat.completions.create(
    model="Qwen/Qwen3.8-27B",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "What does this chart show?"},
            {"type": "image_url", "image_url": {"url": "https://example.com/chart.png"}},
        ],
    }],
)
```

## Sending images

- **URL** — a publicly reachable `https://` URL. We fetch it once to process the request.
- **Base64** — a data URL such as `data:image/png;base64,...`. Use this for private images.

Supported formats are PNG, JPEG, WebP, and non-animated GIF. Images are counted as input tokens; larger images cost more.

## Video

Models whose input modalities include **video** — for example Kimi K3, MiniMax M3, and GLM 5.3 Flash — accept a video
content part in the same way:

```python
{"type": "video_url", "video_url": {"url": "https://example.com/clip.mp4"}}
```

Video is sampled into frames and counted as input tokens. Filter the [Model catalog](/docs/getting-started/models) by
"Image / video input" to see which models accept images or video.

Images, like all inputs, are processed in memory and not stored — see [Zero data retention](/docs/data-security/zero-data-retention).
