Skip to content
← Blog

Inside Muse’s browser stack

This is Part 3 of digging through the Muse filesystem, after my two previous articles (Part 1, Part 2) hit #1 on Hacker News last week.

I’ve seen a few posts about how good Muse and Instinct are at browser use.

So I’ve dug through the Muse filesystem and some of the machinery behind what makes it feel “magical”.

The browser itself is a near-stock Chromium installation (if you’re a dev, this might be a sign to just download it, not fork it). The interesting code really comes from how Meta provisions it, routes its browser traffic, and gives the agent control.

The Muse logo beside a stack of browser windows, captioned Inside the browser stack

A red herring

I started my journey in /opt/meta-chromium/chrome and went looking for what made this “meta”-chromium, and found something called kHatchBitmapData. Hatch is the internal runtime behind Muse, so this looked promising. I’ve got them... doh

After about 30 minutes I realized this was a red herring and was actually just PDFium XFA hatch-fill patterns that have been in Chromium for years. (“Hatch” the graphics term vs “Hatch” the product... Shite.)

Onwards.

The broker

The better lead was found in /opt/hatch/bin/browser-broker.

This was an 18 MB Rust binary sitting in front of Chromium, and really the central controller of the operation. The build paths point to Meta’s internal hatch-engine workspace, with crates called browser-broker, hatch-browser-lease, hatch-networking, and hatch-media.

From what I can tell, the Chrome DevTools Protocol (CDP) layer is hand-rolled on Tokio, and whoever built this put. in. some. time.

(A lot of my findings here come from runtime files, embedded strings, build paths, and documentation. These tend to reveal the shape of the system Meta is using, but they don’t establish every implementation detail or how often each path runs in production, etc.)

A pool of browser machines

browser-broker requests a browser machine from a remote pool of pre-warmed VMs built specifically to run browser tasks. This is different from my initial understanding of Muse, which was that the browser tool was just like every other tool: the agent works on its own computer, your own VM, the Muse differentiator.

The runtime files reference a dedicated image at /hatch_browser_pool/hatch-browser-image. Keeping a pool of these machines warm greatly reduces startup time, and running Chromium on a separate VM keeps its memory consumption away from the agent’s own VM. Both improve latency and performance in a meaningful way.

But the most interesting reason to separate the machines is what I found next: IP routing.

From browser-broker I found an internal Meta control plane called Stefi. Stefi’s job is to confirm the machine lease and its egress configuration, connected to the VM through a Unix socket. When this fails, because of a bad egress profile, a bad machine or some other reason, there’s a fallback that runs the browser tool inside the agent’s own VM.

So where is all this going, and why the separate VMs, really?

Residential egress

The remote browser can have something the local fallback doesn’t have in the config: residential internet access.

The Muse config points to Oxylabs for this. Instead of sending browser traffic directly from a datacenter IP, the system routes it through a residential exit. This is an essential part of the lease machinery.

A request includes a requested_egress_profile, with values such as default, residential, and unknown. The grant returns an endpoint IP and port, plus a PASETO token. There’s telemetry called public_egress_changed and failure names such as ResidentialEgressUnavailable.

So why all the smoke and mirrors?

Well, websites care about where traffic comes from. An agent request arriving from cloud infrastructure can be treated differently from one arriving through a residential network. That can mean more challenges, blocked pages, or a checkout that never gets far enough for the agent to do anything useful and complete its task.

A residential IP gets you past most of the blocking from websites that don’t take kindly to headless Chromes. It doesn’t make an automated browser undetectable or guarantee passage through Cloudflare, but it works pretty damn well.

The config also takes the user’s location into account.

I found references to location_refresh.rs, which describes client location metadata, including client_ip, longitude, client_timezone, and a freshness stamp. That context feeds into the provisioning path to pick the right residential exit.

This gives the system a way to select an IP appropriate to the user’s geo, which matters for a lot of tasks on the web.

After all of this runaround, the agent has a clean residential IP instead of a datacenter IP and can go to town.

Meta also appears to measure the results of Stefi: the schema and documentation include bot-detection outcomes alongside session-assignment and egress information. I imagine those fields let Meta compare residential exits against default exits, which you could do a few things with.

Clicking the right thing

So, once the browser finally reaches a page, how does it actually click anything?

The Rust controller contains names like AxSnapshot, AxResolve, and AxRead. These point to an accessibility-tree approach, where the agent gets a structured view of the page with references, roles, and names for elements.

The agent can ask to “click ref 42.” The broker handles finding the element and working out where to send the input next. That saves the model from having to stare at a screenshot and guess where the button is using computer vision or some other approach.

Visual clicking is still available when the page doesn’t expose a useful element target, and some interfaces still need it.

This is my reconstruction from the strings:

  1. Find the element. AxResolve resolves the reference to a node. There’s an expected_content_quad, which describes the region the element is expected to occupy.
  2. Check whether it can be clicked. Actionability checks appear to cover things like visibility, enabled state, and stability. If something is wrong, there are fields for actionability_reason and actionability_detail.
  3. Check consent. There’s an action_boundary_consent gate in the action machinery. Finding a button doesn’t automatically mean the agent has permission to press it.
  4. Send the input and check it. CDP dispatches the action, with InputDispatchVerifier providing a separate verification step. Sending a click is only part of the job.
  5. Watch what happens. InputClickAndWatch and post_click_frame_tree_ms point to monitoring the frame tree after the click, alongside document changes and lifecycle events. There’s explicit synchronization machinery here, beyond just waiting half a second and hoping.
  6. Work out what changed. There are mutation watermarks and a mutation_ambiguous outcome. These suggest the broker tracks page changes and has a way to report when their relationship to the action is unclear.

Warm browsers and profiles

Browsers seem to live in a slot pool and carry persistent profiles. This means cookies and local storage can preserve authenticated sessions between tasks, and the agent can stay logged in. The profile also appears to contain history and bookmarks.

So we’ve got a warm browser that reduces startup work and a persistent profile that preserves state. Together they save the (world) agent from repeatedly booting an environment and asking you to log in again, and again, and again.

Onwards.

Three brokers

Beyond browser-broker there are three distinct broker runtime directories:

  • /run/hatch/browser-broker/
  • /run/hatch/cron-browser-broker/
  • /run/hatch/proactivity-browser-broker/

Each has its own viewer socket, which suggests separate browser infrastructure exists for interactive work, scheduled jobs, and proactive tasks. The browser machinery appears to extend beyond the conversation you currently have open, which allows for deep research and long-running background tasks. (A couple of magic sprinkles here.)

The proactivity-browser-broker is a pretty interesting directory to find, and definitely deserves some sprinkles for the magic category as well.

Handoff, credentials and the black hole

Less exciting, but when you take over the browser from the agent, the broker first tries to fence off downloads. If that fails, it revokes read access and continues the handoff. There’s explicit handling of downloads and read permissions as control passes back to you.

Credential handling seems pretty solid, and it’s present throughout the broker’s lifecycle vocabulary. There are separate events for credential_task_retired and renderer_task_retired, and for credential_browser_recycled and renderer_browser_recycled.

Those names suggest distinct credential and rendering roles, but they don’t establish whether the separation is between browser instances, processes, or something else. Password handling is one of the next paths I want to trace, and I can tell you it takes a separate path by design.

There are references to sensitive-site policies, navigation consent, and approval gates for credentials and purchases. Its central_blackhole policy machinery includes central_blackhole_version, central_blackhole_manifest_sha256, and last_heartbeat_at.

Those fields sound super scary and interesting at the same time. They don’t tell me how updates are delivered or where every rule is enforced, but I’d love to take a further look into the black hole.

There’s also a path named challenge_capture. That raises a question I didn’t address earlier: what happens when residential egress works but the page still throws a challenge? I haven’t established whether challenge_capture collects evidence for the agent, hands the challenge to the user, feeds another process, or solves the CAPTCHA. Honestly, it could be all four.

So this is, in a nutshell, the browser breakdown for Muse in one narrow instance.

The lightweight path

BUT, not every task has to go through the browser process above.

Our buddy browser-broker also has a lightweight path via its little brother, browser-service. Here you have exposed connectors like Firecrawl, Perplexity, a page cache, and a direct-HTTP path that calls direct_http_observer to fetch.

A simple fetch can go to Firecrawl, rather than a task that needs an authenticated, interactive VM browser session.

There’s a lot I’m leaving out of this article, and frankly too much to cover in one go.

The last tidbit came from the search blocklists under /opt/hatch/assets/blocklist, which point to who Meta and Muse have news and IP deals with. A file called mase_contract_domains.fst lists 295 “contract” publishers (WSJ, Barron’s, Fox, People, Reuters, CNN, Le Monde, El País and a long tail of USA Today local papers), while the 2.67-million-host safety blocklist next to it, mostly adult and phishing sites, also includes nytimes.com, washingtonpost.com, bbc.com, theguardian.com, bloomberg.com, reddit.com and x.com. It also potentially shows how results from the app’s web search tool will be skewed by this approach.

Thoughts

I went looking for what made the browser feel good, and ended up spending most of my time on the broker, learning what makes these systems tick and how much is actually involved.

Getting a browser ready, preserving a session, choosing an exit, identifying an element, checking an action, and handing control back to you is a task that spans 300 files in the Muse repo.

pete at mouse dot dev

-Pete

@heypeterjames