How it works
The state, the questions, and how the model turns them into probabilities.
A request has two parts:
state: everything the model should know. A string, a JSON object, a list, with or without images.questions: what you want to know about the state. Each question names its possible answers.
The model reads the state once, then scores every option of every question in the same forward pass. No text is generated and nothing is sampled, so the response time does not grow with the length of an answer.
The state
The state is any JSON value. Objects and lists are flattened into text the model reads top to bottom, and field names are kept as labels:
{
"goal": "open the settings page",
"elements": ["[1] button Settings", "[2] link Help"]
}goal: open the settings page
elements:
- [1] button Settings
- [2] link HelpThis has two practical consequences:
- Name your fields well.
refund_policy,customer_messageandorder_totaltell the model what each value is.a,b,datado not. - Order matters. The model reads from left to right. Put the context first and the thing to be judged after it, the way you would write it for a person.
An image in the state is read at the place where it stands, and its field name is its label.
Put into the state only what the questions need. A shorter state is cheaper, faster, and leaves less for the model to be distracted by.
The questions
questions is an object: the names are yours, and the answers come back under the same names.
{
"questions": {
"is_spam": { "type": "noul", "instructions": "Is this message spam?" },
"language": {
"type": "choice",
"instructions": "What language is the message written in?",
"criteria": { "en": "English", "de": "German", "uz": "Uzbek", "other": "Any other language" }
}
}
}There are three types, described in Question types:
| Type | Answers | Returns |
|---|---|---|
noul | yes or no | the probability of yes |
choice | one of the options you name (up to 255) | the chosen option and a probability for each |
score | a position on an ordered scale | the expected level and a probability for each |
Ask everything you need about a state in one request rather than one request per question: the state is read once for all of them.
Probabilities are the product
The model's answer is a distribution, and what to do with it is your code's decision:
answer = answers["target"]
if answer["confidence"] >= 0.9:
click(answer["choice"])
else:
ask_the_planner_again()- Put a threshold on the number that fits the cost of a mistake. A wrong click that can be undone tolerates a lower threshold than a payment.
- A flat distribution is information too: it says the state does not settle the question. Send such cases to a slower path (a larger model, a person) instead of taking the top option.
Writing good questions
- Ask about what is in the state. "Is the form showing an error?" can be answered from a screenshot. "Will the customer churn?" cannot be answered from one message.
- Describe every option. In a
choice, the description after the name is what the model compares against. Say when the option applies, not only what it is called. - Give a way out. If none of the options may be right, add one that says so (
"none": "None of the listed elements match."). Without it the probabilities still sum to 1 over the options you gave. - One question, one fact. "Is the form valid and is the button enabled?" is two questions. Ask them separately and combine the answers in code.
- Keep policy in your code. The model is at its best when it reports what is there. See what to expect.