Self-Operating Computer lets a vision-capable AI model control a desktop by reading the screen and choosing mouse and keyboard actions to carry out a goal. It's a Python framework for developers and researchers exploring AI agents that work through application interfaces. The code uses the MIT license.
The framework runs on macOS, Windows and Linux, with an X server required on Linux. On macOS, the terminal needs Screen Recording and Accessibility permissions. Desktop control happens on your machine; where the model processes screen input depends on the backend you choose. LLaVA runs locally through Ollama. Cloud models require provider API access and send model inputs to the selected service.
Supported models include GPT-4o, GPT-4.1 and o1, along with Gemini Pro Vision, Claude 3 and Qwen-VL. That choice lets you compare different models on computer-use tasks within the same framework. The local LLaVA route has very high error rates, so it's a base for experimentation rather than a dependable desktop assistant.
Beyond reading screenshots, the framework supports optical character recognition to help models select interface elements by their text. Its Set-of-Mark mode uses YOLOv8 button detection to mark visual targets, and developers can replace the detector with their own trained model. Voice input lets you speak the objective you want the agent to attempt.
Claim this page and we'll verify you by hand. Self-Operating Computer gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Self-Operating Computer?Promote it
Something wrong or outdated on this page?
12.4KUpdated 4 weeks agoApache-2.0
macOS · Windows · Linux#Code execution#Multimodal input#Tool calling
Agent S is an open-source AI agent framework that controls ordinary desktop and web applications through the screen, mouse and keyboard. It runs on macOS, Windows and Linux and is aimed at developers, computer-use researchers and people building automation for their own desktops. You describe the task in natural language; the agent clicks, types and scrolls without requiring an API integration or a separate script for each application.
116.8KUpdated 4 days agoMIT
#MCP#Ollama integration#Structured output
6.4KUpdated 2 years agoApache-2.0
Web · Browser Extension#Code execution
LaVague is a Python framework for developers building AI agents that carry out tasks in a web browser. An agent takes a plain-language objective, examines the current page, and generates and executes browser actions across multiple steps. It's open source under the Apache 2.0 license.
10.5KUpdated 7 months agoApache-2.0
macOS · Windows · Linux · Android · Web#Code execution#MCP#Multimodal input
4.4KUpdated 19 hours agoMIT
macOS · Windows · Linux · Web · JetBrains#Agent Client Protocol#Code execution#Git integration
28.6KUpdated 1 day agoMIT
macOS · Windows · Linux#LM Studio integration#MCP#Multi-agent workflows
Browser Use is an MIT licensed browser agent for developers who want AI to carry out tasks on websites. You can run the open source agent on your own machine from Python, choose a model, and use either a local or cloud browser. A CLI is available for browser tasks too.
aichat brings Ollama and cloud AI services into the same terminal interface for developers and people who work at the command line. It runs locally on macOS, Linux and Windows, with Android support through Termux. Model processing happens through the backend you choose: Ollama supports local models, while providers such as OpenAI, Claude and Gemini process requests in the cloud.
gptme is a self-hosted AI agent that works directly in your terminal, with access to your files and installed tools. It's for developers who want a coding assistant in their own environment, and people who want an agent for data analysis or other knowledge work. The software is free under the MIT license and doesn't require a gptme account.
Semantic Kernel is an MIT-licensed SDK for developers adding AI agents to their applications. It supports local models through Ollama, LMStudio and ONNX, alongside cloud services such as OpenAI and Azure OpenAI. You choose the model backend. The SDK runs on Windows, macOS and Linux and supports C#, Python and Java.