openinstinctdocs

How it works

The state, the questions, and how the model turns them into probabilities.

A request has two parts:

  • state: everything the model should know. A string, a JSON object, a list, with or without images.
  • questions: what you want to know about the state. Each question names its possible answers.

The model reads the state once, then scores every option of every question in the same forward pass. No text is generated and nothing is sampled, so the response time does not grow with the length of an answer.

The state

The state is any JSON value. Objects and lists are flattened into text the model reads top to bottom, and field names are kept as labels:

What you send
{
  "goal": "open the settings page",
  "elements": ["[1] button Settings", "[2] link Help"]
}
What the model reads
goal: open the settings page
elements:
  - [1] button Settings
  - [2] link Help

This has two practical consequences:

  • Name your fields well. refund_policy, customer_message and order_total tell the model what each value is. a, b, data do not.
  • Order matters. The model reads from left to right. Put the context first and the thing to be judged after it, the way you would write it for a person.

An image in the state is read at the place where it stands, and its field name is its label.

Put into the state only what the questions need. A shorter state is cheaper, faster, and leaves less for the model to be distracted by.

The questions

questions is an object: the names are yours, and the answers come back under the same names.

{
  "questions": {
    "is_spam": { "type": "noul", "instructions": "Is this message spam?" },
    "language": {
      "type": "choice",
      "instructions": "What language is the message written in?",
      "criteria": { "en": "English", "de": "German", "uz": "Uzbek", "other": "Any other language" }
    }
  }
}

There are three types, described in Question types:

TypeAnswersReturns
noulyes or nothe probability of yes
choiceone of the options you name (up to 255)the chosen option and a probability for each
scorea position on an ordered scalethe expected level and a probability for each

Ask everything you need about a state in one request rather than one request per question: the state is read once for all of them.

Probabilities are the product

The model's answer is a distribution, and what to do with it is your code's decision:

answer = answers["target"]
if answer["confidence"] >= 0.9:
    click(answer["choice"])
else:
    ask_the_planner_again()
  • Put a threshold on the number that fits the cost of a mistake. A wrong click that can be undone tolerates a lower threshold than a payment.
  • A flat distribution is information too: it says the state does not settle the question. Send such cases to a slower path (a larger model, a person) instead of taking the top option.

Writing good questions

  • Ask about what is in the state. "Is the form showing an error?" can be answered from a screenshot. "Will the customer churn?" cannot be answered from one message.
  • Describe every option. In a choice, the description after the name is what the model compares against. Say when the option applies, not only what it is called.
  • Give a way out. If none of the options may be right, add one that says so ("none": "None of the listed elements match."). Without it the probabilities still sum to 1 over the options you gave.
  • One question, one fact. "Is the form valid and is the button enabled?" is two questions. Ask them separately and combine the answers in code.
  • Keep policy in your code. The model is at its best when it reports what is there. See what to expect.

On this page