Reading screens
Find an element, check a state, and verify an action on a screenshot.
Screens are what Instinct One was built for first: a screenshot and, when you have it, the list of elements on it go into the state; the questions ask what is there.
A helper used by the examples on this page:
import base64
import os
import requests
def image(path, media_type="image/png"):
with open(path, "rb") as f:
data = base64.b64encode(f.read()).decode()
return {"type": "image", "source": {"type": "base64", "media_type": media_type, "data": data}}
def ask(state, questions):
response = requests.post(
"https://api.openinstinct.dev/v1/systemone",
headers={"Authorization": f"Bearer {os.environ['OPENINSTINCT_API_KEY']}"},
json={"state": state, "questions": questions},
timeout=120,
)
response.raise_for_status()
return response.json()["answers"]Which element does the instruction mean?
List the elements with a number each (from the accessibility tree, the DOM, or your own detector), and ask a
choice over the numbers. The model answers with a probability for every element.
elements = {
"1": 'push button "Cancel"',
"2": 'push button "Save"',
"3": 'check box "Wrap lines", checked',
"4": 'text field "File name", value "notes.txt"',
}
answers = ask(
state={
"instruction": "Save the file",
"screen": image("screenshot.png"),
"elements": [f"[{n}] {text}" for n, text in elements.items()],
},
questions={
"target": {
"type": "choice",
"instructions": "Which element does the instruction act on?",
"criteria": {**elements, "none": "None of the listed elements."},
},
},
)
target = answers["target"]
if target["choice"] != "none" and target["probabilities"][target["choice"]] >= 0.9:
click_element(target["choice"])- The
noneoption gives the model a way to say the element is not on the screen. Without it, the most likely of the listed elements is returned even when nothing fits. - The list alone, without the screenshot, is often enough and much cheaper: a text-only request is a few hundred tokens, a screenshot adds about two thousand. Send the image when the names do not settle it (icons without labels, several buttons with the same name).
Is the screen in this state?
noul questions about what is visible. Several of them cost one request.
answers = ask(
state={"screen": image("checkout.png")},
questions={
"form_error": {"type": "noul", "instructions": "Is the form showing a validation error?"},
"dialog_open": {"type": "noul", "instructions": "Is a modal dialog open over the page?"},
"loading": {"type": "noul", "instructions": "Is the page still loading (a spinner or a progress bar)?"},
},
)
if answers["loading"]["noul"] > 0.5:
wait_and_look_again()
elif answers["form_error"]["noul"] > 0.5:
report("the form has an error")Did the action work?
Take a screenshot before and after, name them, and ask about the difference. Both images go into one state.
answers = ask(
state={
"action": 'Clicked the "Dark mode" switch',
"before": image("before.png"),
"after": image("after.png"),
},
questions={
"worked": {
"type": "noul",
"instructions": "Did the action have its expected effect on the screen?",
"criteria": {
"true": "`after` shows the interface in a dark theme.",
"false": "`after` looks the same as `before`, or shows an error.",
},
},
},
)Saying in criteria what success looks like for this action makes the question one of perception, which is what
the model does best.
Tips for screenshots
- Scale down before sending. 1280 pixels on the longer side, as JPEG, is enough for most desktop screens and uploads several times faster than a full-size PNG.
- Crop for small text. If the question is about one area, send that area.
- Ask what is there, decide in code. "Is the Pay button enabled?" and "Is the card number field empty?" are reliable. "What should be done next?" asks the model to plan; see what to expect.
- Measure on your screens. Collect a few dozen of your own cases with the right answers and look at the probabilities before you pick a threshold.