Errors
Error responses, their codes, and which ones to retry.
Shape
An error is a JSON body with the HTTP status of the failure:
{
"error": { "code": "insufficient_balance", "message": "..." },
"detail": "..."
}error.codeis stable: branch on it.error.messageis for people: log it, do not parse it.detailrepeats the message in the form TypeSafe and Kev clients read. For a request the model refused, it is the model server's own detail, unchanged.
Codes
| Status | code | When | Retry? |
|---|---|---|---|
| 400 | invalid_json | The body is not JSON | No. Fix the request |
| 400, 422 | invalid_request | The body has the wrong shape, or the model refused it (an unreadable image, too many options) | No. Fix the request |
| 401 | invalid_api_key | The key is missing, wrong or revoked | No |
| 402 | insufficient_balance | The workspace's balance is zero or below. Not checked for free models | After topping up |
| 404 | model_not_found | No such model, or it is disabled | No |
| 413 | request_too_large | The body is over 32 MB | No. Send smaller images or a smaller batch |
| 429 | rate_limited | The workspace's requests-per-minute limit | Yes, after retry-after |
| 429 | concurrency_limited | Too many requests in flight | Yes, after retry-after |
| 502 | upstream_error, unreachable | The model server failed or could not be reached | Yes |
| 503 | model_loading, overloaded | The model is starting, or busy | Yes, after a few seconds |
| 503 | not_configured | The model has no server | No |
| 504 | timeout | No answer in time. Often a cold start | Yes |
Cold starts
The model's GPUs are released when idle. A request that arrives then starts one, and waits. If the wait is longer
than the API's timeout you get 504 timeout or 503 model_loading, and the model keeps starting: the same
request a little later is answered.
So treat 502, 503 and 504 as "try again", with a pause that grows:
import os
import time
import requests
RETRY = {429, 502, 503, 504}
def ask(body, attempts=6):
for attempt in range(attempts):
response = requests.post(
"https://api.openinstinct.dev/v1/systemone",
headers={"Authorization": f"Bearer {os.environ['OPENINSTINCT_API_KEY']}"},
json=body,
timeout=120,
)
if response.status_code not in RETRY:
response.raise_for_status()
return response.json()
wait = float(response.headers.get("retry-after", 2**attempt))
time.sleep(min(wait, 30))
response.raise_for_status()Set your client's own timeout generously (a minute or more) for the first request, and do not retry 400-class
errors: the same request will fail the same way.
Reporting a problem
Every response carries an x-request-id header. Send it with your report: it finds the request in the logs.