Agent loop
Use Instinct One for the fast decisions inside an agent that operates a computer.
An agent that operates a UI spends most of its time on small questions: which element is "Save"? did the dialog close? is the page still loading? Answering each with a turn of a large generative model costs seconds. Instinct One answers them in about a hundred milliseconds, with a number the loop can check.
The division of work:
| Does | |
|---|---|
| A large model (the planner) | Understands the task, writes the plan: steps with an instruction and an expected result |
| Instinct One | For each step: finds the element, checks whether the step worked |
| Your code | Applies thresholds, performs the clicks, goes back to the planner when unsure |
One step
Each step asks two things in one request: where to act, and whether the goal of the step is already reached.
def step_request(instruction, expected, screenshot, elements):
return {
"state": {
"instruction": instruction,
"expected_result": expected,
"screen": screenshot, # an image object
"elements": [f"[{n}] {text}" for n, text in elements.items()],
},
"questions": {
"done": {
"type": "noul",
"instructions": "Does the screen already show `expected_result`?",
},
"target": {
"type": "choice",
"instructions": "Which element does `instruction` act on?",
"criteria": {**elements, "none": "None of the listed elements."},
},
},
}The loop
ACT = 0.9 # act only when this sure
DONE = 0.8
def run(plan, desktop, ask):
for step in plan:
for attempt in range(3):
screenshot, elements = desktop.look()
answers = ask(step_request(step["instruction"], step["expected"], screenshot, elements))
if answers["done"]["noul"] >= DONE:
break # next step
target = answers["target"]
choice = target["choice"]
if choice == "none" or target["probabilities"][choice] < ACT:
return {"stopped_at": step, "saw": answers} # back to the planner
desktop.click(choice)
else:
return {"stopped_at": step, "saw": answers}
return {"finished": True}What makes this work:
- The planner states the expected result of every step. "The Save dialog is open" turns "did it work?" into a question about what is on the screen.
- The loop stops when it is not sure, and returns what it saw. The planner reads that and corrects the plan. A wrong click costs more than a question.
- Thresholds are yours. Use a higher one before anything that cannot be undone, or ask a person first.
- The same screen is cheap to ask about twice. When the state of a request repeats the previous one, the server reuses its reading of it.
Do not ask the model for the plan
"What should be done next?" with a list of actions is a planning question. The model can see that a form shows an error and still prefer "press Pay". Keep the plan with the planner and the policy in your code; ask Instinct One what is on the screen. See what to expect.
A ready-made implementation
computer-use is this loop as a tool: it gives an agent its own mouse and keyboard on
a real desktop, lists the elements on the screen, and asks Instinct One which one an instruction means. Its
run-plan command executes a JSON plan and verifies every action from a fresh look.