openinstinctdocs

Models and limits

The models the API serves, what they are good at, and where they fall short.

Models

NameWhat it is
instinct-one-latestThe default. The 4B model: text and images
instinct-one-flashThe smaller 0.8B model. Faster and less accurate; an early version
  • Leave "model" out of a request and the default model answers.
  • GET /v1/models lists the models enabled right now.
  • Prices per million tokens are shown in the console.

What a request can hold

StateText or JSON of any shape
Images in the stateUp to 8; PNG, JPEG, WebP
QuestionsOne or more, of types noul, choice, score
Options in one questionUp to 255
Request bodyUp to 32 MB

Speed

The model's own time for a decision is about 100 ms, reported as latency_ms in the response. What you measure from your side adds the network and, for images, the upload.

The model runs on GPUs that are released when idle. The first request after a quiet period waits for a GPU to start, which can take a minute or two; it may come back as 503 model_loading or 504 timeout. Retry with a pause. See Errors.

Repeated states are faster: when consecutive requests carry the same state (an agent asking several rounds of questions about one screen), the server reuses its reading of the state and computes only the questions.

What to expect

Instinct One is trained on text decisions (classification, routing, checks, rating) and on images of desktop screens, charts, tables and synthetic scenes.

It is strongest at perception: questions whose answer is in the state.

  • Which of the listed elements is the "Save" button?
  • Is the form showing a validation error?
  • Which bar is the tallest? What does the row for March say?
  • Which team does this ticket belong to?

It is weaker at decisions that need reasoning toward a goal. A question like "what should the agent do next?" asks the model to combine a goal, the state of the screen and a plan. It can see that a form shows an error and still rank "press Pay" above "fix the card number". For such cases:

  • Ask the perception questions (form_error, card_number_filled, pay_button_enabled) and write the rule that combines them in your code; or
  • let a larger model plan, and use Instinct One for the fast questions inside each step. See Agent loop.

Other things to know:

  • It does not generate text. It cannot extract a value, write a summary or explain its answer. It chooses among the answers you supply.
  • It always answers. The probabilities sum to 1 over your options even when none is right. Add a "none of these" option where that can happen, and read the confidence.
  • It does not compute. Arithmetic over dates in the state (how many days between two of them) is unreliable. Compute in your code and put the result into the state.
  • Screens outside its training look different. Accuracy measured on desktop applications does not carry over unchanged to every website or mobile UI. Measure on your own cases before relying on a threshold.

Compatibility

The request and the response follow TypeSafe's System One API. A client written for Jev or Kev works against https://api.openinstinct.dev with an OpenInstinct key; requests without images get a response of the same shape. The server also echoes the x-typesafe-request-id header.

Instinct One is not a reproduction of Jev: Jev's architecture and training are not published.

On this page