openinstinctdocs

computer-use

A mouse and a keyboard of its own for an AI agent on your desktop, with Instinct One for finding elements.

computer-use is an open-source command-line tool and a skill for Claude Code. It lets an agent work in the windows of a real desktop while you keep working at it: the agent has its own cursor, marked "AI", and its own keyboard.

It works without any server. With an Instinct One key, it can also answer "which element does this instruction mean?" in one fast request instead of a turn of a large model.

Platform
LinuxX11 sessions (not Wayland)
macOSWith the Accessibility and Screen Recording permissions
WindowsNot supported yet

Install

git clone https://github.com/OpenInstinct/computer-use && cd computer-use
./install.sh --check      # what is there and what is missing
./install.sh

Or as a Claude Code plugin:

claude plugin marketplace add OpenInstinct/computer-use
claude plugin install computer-use@openinstinct

Connect it to the API

Nothing is sent anywhere until you configure a server.

computer-use config --url https://api.openinstinct.dev --key -    # the key is typed, not shown
computer-use status --server                                      # does the server answer; its models
computer-use config --model instinct-one-latest                   # optional

Ask about the screen

computer-use start                                    # makes the agent's devices
computer-use elements                                 # [12] push button "Save"  @187,239

computer-use find "the Save button"                   # ranked elements with probabilities, no click
computer-use act "Press Save"                         # find, then click if the probability is at least 0.9
computer-use ask 'Is check box "Wrap lines" checked?' # the probability of yes
CommandRequest it makes
findA choice over the elements on the screen
actThe same, then a click when the top element reaches the threshold
askA noul about the screen
doSteps by element name; asks only when a name does not settle the element
run-planExecutes a JSON plan, verifying each action from a fresh look

What is sent

ModeSent with each question
text (default)The instruction, and the list of elements: roles, names, states, the text fields hold
hybridThe same, and a screenshot of the monitor
visionThe instruction, a screenshot, and the elements' boxes

The text of a type command is not sent. The key is never printed and is not sent over plain http to another machine.

Safety

The tool acts on a real desktop, so every command checks itself: it does not click on a screen that changed since the agent looked, types only into a window its own click selected, refuses window-system shortcuts, does nothing on a locked screen, and computer-use stop removes its devices at once.

These guards narrow what can go wrong; they do not make an agent's judgment safe. find answers with one of the listed elements and has no "none of these" answer, so read its answer before clicking.

Everything else, including all commands, configuration and limits, is in the repository's README.

On this page