Skip to content
Mirai Minds

Case studyAI agents and automation

Rouh: a computer-use agent for the desktop

By Sagar DavaraPublished

Prototype

In short

Mirai Minds built the prototype of Rouh, a desktop AI agent that carries out tasks on a user's own computer. The app runs a browser agent that works from screenshots, connects to MCP servers such as Google Workspace, Telegram and WhatsApp, runs scheduled tasks and keeps long-term memory with Mem0. It is a working prototype, not a released product.

The problem

Rouh's founder wanted an assistant that does things instead of only answering: open the browser, find the page, fill the form, send the message, all on the user's own computer. Many of those actions have no API, and the ones that do are spread across different apps.

What we built

  • 2024: a browser-extension agent that turned instructions into actions such as opening tabs and searching the web.
  • 2025: a desktop app built with Tauri v2, a Next.js interface and a Python FastAPI sidecar packaged as a single executable. Inside it:
    • a browser agent that works from screenshots, scaled down for the model and mapped back to real screen coordinates;
    • an MCP hub where users add servers, including Google Workspace, Telegram and WhatsApp, or their own;
    • a manager for scheduled tasks and a listener for notifications;
    • long-term memory, first as an MCP memory server and later with Mem0.
  • Model experiments on screen understanding with PyTorch Lightning, screen data and synthetic data.

How the prototype works

A request goes to the agent loop, which chooses between a browser action and an MCP tool. Browser actions run through Playwright; after each one the agent gets a fresh screenshot and decides the next step. Scheduled tasks run on their own thread so they don't block the interface. Memories are written after conversations and read back when they are relevant to a new request. The approach behind it is described under AI agents and automation.

What we'd change

  • Start with five tasks, not an open-ended assistant. Open-ended computer use is where agents are slowest and least reliable. A short list of tasks, each with its own test, would reach real users sooner.
  • Make confirmation part of the first release. Any irreversible step, such as sending a message or paying, should show the user exactly what will happen and wait for a yes.

What it runs on

Desktop
Tauri v2Next.jsPython FastAPI sidecarPyInstaller
Agent
ClaudeGPT-4oPlaywrightModel Context Protocol
Memory and scheduling
Mem0Scheduled-task manager
Model experiments
PyTorch LightningScreen-understanding dataSynthetic data
Web
VueFastAPIMongoDB
Awesome work!!! Everything works as expected. I don’t have any major feedback other than keep going.
Humaid AlfalahiFounder, Rouh.aiThe project →

Asked on the first call

What can the agent do on a computer?

In the prototype it drives a browser: it takes a screenshot, decides where to click or what to type, acts and looks again. It can also call connected apps through MCP servers, such as mail and calendar through Google Workspace, and run tasks on a schedule the user sets.

How is it kept safe on a personal computer?

It runs locally in a desktop app and acts only through the browser and the MCP servers the user has connected. Every task comes from the user, either directly or as a schedule they set, and each app connection is added by the user.

Why is it still a prototype?

Computer-use agents are slow and can fail on long, multi-page tasks. The build proved the approach and the desktop packaging. A production release needs a narrower task list, tests for each task and a confirmation step before anything irreversible.

Have a system in mind? Let's scope it.

A 30-minute call with an engineer who has shipped this before. You leave with a plan, a rough timeline and what it would take — whether or not we build it.