openinstinctdocs

Rate limits and usage

Limits per workspace, the headers that report them, and how usage is counted.

Limits belong to the workspace

A limit is not per key. All keys of a workspace, and its playground, share one count.

LimitMeaning
Requests per minuteA 60-second window. Each request inside a batch counts as one
Concurrent requestsRequests being answered by the model at the same moment. A batch takes one slot

Limits grow with the workspace's tier, which is set by the total it has paid; credit that was granted does not count toward it. A workspace can also have limits agreed individually.

Your current values are in every response:

Header
x-ratelimit-limit-requestsRequests per minute
x-ratelimit-limit-concurrentConcurrent requests
x-ratelimit-remaining-requestsRequests left in the current window

When a limit is reached

The request is refused with 429 before it reaches the model, and a retry-after header says how many seconds to wait:

code
rate_limitedThe per-minute limit
concurrency_limitedToo many requests in flight

To stay under the limits:

  • Ask all questions about a state in one request instead of one request per question.
  • Use a batch for lists of items: it takes one concurrent slot.
  • Keep a small pool of workers (at most your concurrent limit) rather than firing all requests at once.

Usage and billing

Usage is counted in tokens, from the usage of each response:

Counted as
Text tokensinput_tokens - image_tokens
Image tokensimage_tokens
Output tokensoutput_tokens: the tokens of the serialised answers

Each kind has its own price per million tokens, per model; the console shows the prices, your usage by day and by key, and the workspace's balance.

  • Only requests that reach the model are counted. A request refused for its key, shape, limit or balance costs nothing.
  • When the balance of a workspace is zero or below, requests to paid models get 402 insufficient_balance. Free models are not affected.

On this page