openinstinctdocs

Errors

Error responses, their codes, and which ones to retry.

Shape

An error is a JSON body with the HTTP status of the failure:

{
  "error": { "code": "insufficient_balance", "message": "..." },
  "detail": "..."
}
  • error.code is stable: branch on it.
  • error.message is for people: log it, do not parse it.
  • detail repeats the message in the form TypeSafe and Kev clients read. For a request the model refused, it is the model server's own detail, unchanged.

Codes

StatuscodeWhenRetry?
400invalid_jsonThe body is not JSONNo. Fix the request
400, 422invalid_requestThe body has the wrong shape, or the model refused it (an unreadable image, too many options)No. Fix the request
401invalid_api_keyThe key is missing, wrong or revokedNo
402insufficient_balanceThe workspace's balance is zero or below. Not checked for free modelsAfter topping up
404model_not_foundNo such model, or it is disabledNo
413request_too_largeThe body is over 32 MBNo. Send smaller images or a smaller batch
429rate_limitedThe workspace's requests-per-minute limitYes, after retry-after
429concurrency_limitedToo many requests in flightYes, after retry-after
502upstream_error, unreachableThe model server failed or could not be reachedYes
503model_loading, overloadedThe model is starting, or busyYes, after a few seconds
503not_configuredThe model has no serverNo
504timeoutNo answer in time. Often a cold startYes

Cold starts

The model's GPUs are released when idle. A request that arrives then starts one, and waits. If the wait is longer than the API's timeout you get 504 timeout or 503 model_loading, and the model keeps starting: the same request a little later is answered.

So treat 502, 503 and 504 as "try again", with a pause that grows:

import os
import time
import requests

RETRY = {429, 502, 503, 504}


def ask(body, attempts=6):
    for attempt in range(attempts):
        response = requests.post(
            "https://api.openinstinct.dev/v1/systemone",
            headers={"Authorization": f"Bearer {os.environ['OPENINSTINCT_API_KEY']}"},
            json=body,
            timeout=120,
        )
        if response.status_code not in RETRY:
            response.raise_for_status()
            return response.json()
        wait = float(response.headers.get("retry-after", 2**attempt))
        time.sleep(min(wait, 30))
    response.raise_for_status()

Set your client's own timeout generously (a minute or more) for the first request, and do not retry 400-class errors: the same request will fail the same way.

Reporting a problem

Every response carries an x-request-id header. Send it with your report: it finds the request in the logs.

On this page