Models and limits
The models the API serves, what they are good at, and where they fall short.
Models
| Name | What it is |
|---|---|
instinct-one-latest | The default. The 4B model: text and images |
instinct-one-flash | The smaller 0.8B model. Faster and less accurate; an early version |
- Leave
"model"out of a request and the default model answers. GET /v1/modelslists the models enabled right now.- Prices per million tokens are shown in the console.
What a request can hold
| State | Text or JSON of any shape |
| Images in the state | Up to 8; PNG, JPEG, WebP |
| Questions | One or more, of types noul, choice, score |
| Options in one question | Up to 255 |
| Request body | Up to 32 MB |
Speed
The model's own time for a decision is about 100 ms, reported as latency_ms in the response. What you measure
from your side adds the network and, for images, the upload.
The model runs on GPUs that are released when idle. The first request after a quiet period waits for a GPU to
start, which can take a minute or two; it may come back as 503 model_loading or 504 timeout. Retry with a
pause. See Errors.
Repeated states are faster: when consecutive requests carry the same state (an agent asking several rounds of questions about one screen), the server reuses its reading of the state and computes only the questions.
What to expect
Instinct One is trained on text decisions (classification, routing, checks, rating) and on images of desktop screens, charts, tables and synthetic scenes.
It is strongest at perception: questions whose answer is in the state.
- Which of the listed elements is the "Save" button?
- Is the form showing a validation error?
- Which bar is the tallest? What does the row for March say?
- Which team does this ticket belong to?
It is weaker at decisions that need reasoning toward a goal. A question like "what should the agent do next?" asks the model to combine a goal, the state of the screen and a plan. It can see that a form shows an error and still rank "press Pay" above "fix the card number". For such cases:
- Ask the perception questions (
form_error,card_number_filled,pay_button_enabled) and write the rule that combines them in your code; or - let a larger model plan, and use Instinct One for the fast questions inside each step. See Agent loop.
Other things to know:
- It does not generate text. It cannot extract a value, write a summary or explain its answer. It chooses among the answers you supply.
- It always answers. The probabilities sum to 1 over your options even when none is right. Add a "none of these" option where that can happen, and read the confidence.
- It does not compute. Arithmetic over dates in the state (how many days between two of them) is unreliable. Compute in your code and put the result into the state.
- Screens outside its training look different. Accuracy measured on desktop applications does not carry over unchanged to every website or mobile UI. Measure on your own cases before relying on a threshold.
Compatibility
The request and the response follow TypeSafe's System One API. A client written for Jev or Kev works against
https://api.openinstinct.dev with an OpenInstinct key; requests without images get a response of the same shape.
The server also echoes the x-typesafe-request-id header.
Instinct One is not a reproduction of Jev: Jev's architecture and training are not published.