openinstinctdocs

Images

Put screenshots, photos, document pages and charts into the state.

An object of this shape, anywhere in the state, is read as an image and not as text:

{
  "type": "image",
  "source": { "type": "base64", "media_type": "image/png", "data": "<base64>" },
  "max_pixels": 1048576
}
FieldMeaning
source.typeAlways "base64"
source.media_typeimage/png, image/jpeg or image/webp
source.dataThe image file, base64-encoded
max_pixelsOptional. A pixel budget for this image below the server's own: fewer tokens, less detail

It can be the state itself, the value of a field, or an item of a list.

Where the image stands

The model reads the image at the place where it stands in the state, and the field name is its label:

{
  "goal": "open the settings page",
  "screen": { "type": "image", "source": { "type": "base64", "media_type": "image/png", "data": "..." } },
  "elements": ["[1] button Settings", "[2] link Help"]
}

The model reads goal: open the settings page, then screen:, then the image, then the list of elements.

Several images

A state can hold up to 8 images, mixed with text in any order. Give each one a name the questions can refer to:

{
  "state": {
    "question": "Which chart shows the larger growth?",
    "chart_a": { "type": "image", "source": { "type": "base64", "media_type": "image/png", "data": "..." } },
    "chart_b": { "type": "image", "source": { "type": "base64", "media_type": "image/png", "data": "..." } }
  },
  "questions": {
    "larger": {
      "type": "choice",
      "instructions": "Which chart shows the larger growth?",
      "criteria": { "chart_a": "The first chart.", "chart_b": "The second chart." }
    }
  }
}

Size, tokens and cost

An image is resized to fit a pixel budget and counted in tokens. A full 1920 × 1080 screenshot is about 2,000 tokens. The response reports them:

{ "usage": { "input_tokens": 2163, "output_tokens": 58, "image_tokens": 2040 } }

image_tokens is present only when the request had images, and those tokens are also counted inside input_tokens: text tokens are input_tokens - image_tokens.

To spend less:

  • Send a smaller image. A JPEG scaled to 1280 pixels on its longer side is enough for most screens, and it uploads much faster than a full-size PNG. On a slow connection the upload, not the model, is most of the wait.
  • Set max_pixels on images where detail does not matter (a thumbnail, a second screenshot for context).
  • Crop to the part the question is about when you know where it is. Small text is read better from a crop than from a whole screen scaled down.

Limits

LimitValue
Images in one state8
One image, encoded16 MB
One image, decoded40 million pixels
Request body32 MB
FormatsPNG, JPEG, WebP

A request over a limit is refused with invalid_request and a message that names the image.

What is not supported

  • URLs and file paths. The server does not open an address a caller names. Send the bytes.
  • Images in questions. An image in instructions or criteria is refused. Put the image in the state and refer to it by its field name.
  • Video. Send frames as separate images.

On this page