Case studyAI agents and automation
Rouh: a computer-use agent for the desktop
By Sagar DavaraPublished
PrototypeIn short
Mirai Minds built the prototype of Rouh, a desktop AI agent that carries out tasks on a user's own computer. The app runs a browser agent that works from screenshots, connects to MCP servers such as Google Workspace, Telegram and WhatsApp, runs scheduled tasks and keeps long-term memory with Mem0. It is a working prototype, not a released product.
The story
The problem
Rouh's founder wanted an assistant that does things instead of only answering: open the browser, find the page, fill the form, send the message, all on the user's own computer. Many of those actions have no API, and the ones that do are spread across different apps.
What we built
- 2024: a browser-extension agent that turned instructions into actions such as opening tabs and searching the web.
- 2025: a desktop app built with Tauri v2, a Next.js interface and a Python FastAPI sidecar packaged as a single executable. Inside it:
- a browser agent that works from screenshots, scaled down for the model and mapped back to real screen coordinates;
- an MCP hub where users add servers, including Google Workspace, Telegram and WhatsApp, or their own;
- a manager for scheduled tasks and a listener for notifications;
- long-term memory, first as an MCP memory server and later with Mem0.
- Model experiments on screen understanding with PyTorch Lightning, screen data and synthetic data.
How the prototype works
A request goes to the agent loop, which chooses between a browser action and an MCP tool. Browser actions run through Playwright; after each one the agent gets a fresh screenshot and decides the next step. Scheduled tasks run on their own thread so they don't block the interface. Memories are written after conversations and read back when they are relevant to a new request. The approach behind it is described under AI agents and automation.
What we'd change
- Start with five tasks, not an open-ended assistant. Open-ended computer use is where agents are slowest and least reliable. A short list of tasks, each with its own test, would reach real users sooner.
- Make confirmation part of the first release. Any irreversible step, such as sending a message or paying, should show the user exactly what will happen and wait for a yes.
Stack
What it runs on
- Desktop
- Tauri v2Next.jsPython FastAPI sidecarPyInstaller
- Agent
- ClaudeGPT-4oPlaywrightModel Context Protocol
- Memory and scheduling
- Mem0Scheduled-task manager
- Model experiments
- PyTorch LightningScreen-understanding dataSynthetic data
- Web
- VueFastAPIMongoDB
Awesome work!!! Everything works as expected. I don’t have any major feedback other than keep going.