openinstinctdocs

Agent loop

Use Instinct One for the fast decisions inside an agent that operates a computer.

An agent that operates a UI spends most of its time on small questions: which element is "Save"? did the dialog close? is the page still loading? Answering each with a turn of a large generative model costs seconds. Instinct One answers them in about a hundred milliseconds, with a number the loop can check.

The division of work:

Does
A large model (the planner)Understands the task, writes the plan: steps with an instruction and an expected result
Instinct OneFor each step: finds the element, checks whether the step worked
Your codeApplies thresholds, performs the clicks, goes back to the planner when unsure

One step

Each step asks two things in one request: where to act, and whether the goal of the step is already reached.

def step_request(instruction, expected, screenshot, elements):
    return {
        "state": {
            "instruction": instruction,
            "expected_result": expected,
            "screen": screenshot,                                   # an image object
            "elements": [f"[{n}] {text}" for n, text in elements.items()],
        },
        "questions": {
            "done": {
                "type": "noul",
                "instructions": "Does the screen already show `expected_result`?",
            },
            "target": {
                "type": "choice",
                "instructions": "Which element does `instruction` act on?",
                "criteria": {**elements, "none": "None of the listed elements."},
            },
        },
    }

The loop

ACT = 0.9        # act only when this sure
DONE = 0.8


def run(plan, desktop, ask):
    for step in plan:
        for attempt in range(3):
            screenshot, elements = desktop.look()
            answers = ask(step_request(step["instruction"], step["expected"], screenshot, elements))

            if answers["done"]["noul"] >= DONE:
                break                                   # next step

            target = answers["target"]
            choice = target["choice"]
            if choice == "none" or target["probabilities"][choice] < ACT:
                return {"stopped_at": step, "saw": answers}     # back to the planner

            desktop.click(choice)
        else:
            return {"stopped_at": step, "saw": answers}
    return {"finished": True}

What makes this work:

  • The planner states the expected result of every step. "The Save dialog is open" turns "did it work?" into a question about what is on the screen.
  • The loop stops when it is not sure, and returns what it saw. The planner reads that and corrects the plan. A wrong click costs more than a question.
  • Thresholds are yours. Use a higher one before anything that cannot be undone, or ask a person first.
  • The same screen is cheap to ask about twice. When the state of a request repeats the previous one, the server reuses its reading of it.

Do not ask the model for the plan

"What should be done next?" with a list of actions is a planning question. The model can see that a form shows an error and still prefer "press Pay". Keep the plan with the planner and the policy in your code; ask Instinct One what is on the screen. See what to expect.

A ready-made implementation

computer-use is this loop as a tool: it gives an agent its own mouse and keyboard on a real desktop, lists the elements on the screen, and asks Instinct One which one an instruction means. Its run-plan command executes a JSON plan and verifies every action from a fresh look.

On this page