Parker Rex
All writing
ai

How Grok Bot works: I took apart version 0.47

· 15 min read

Thanks to GPT-6 Astra Max for helping with this teardown.

Grok Bot is an AI agent app. You talk to a bot and it does things for you: runs commands, uses a browser, texts someone. I wanted to know how it actually works, so I took apart the macOS app.

Someone already did a great version of this. shadown's teardown covers version 0.18.0 in serious detail, and I used it as a checklist. The copy on my Mac was 0.47.0. I checked 29 of that article's claims against the newer build. 12 still held, 11 had changed, 4 no longer matched at all, and 2 I couldn't settle. That isn't a knock on the article. This app moves fast.

Here's the short version. The app on your Mac is not the agent. The agent runs somewhere else. The Mac app is a chat window, a remote control for a cloud computer, and a guard that decides whether that remote agent can touch your machine. Most of the interesting engineering is in the guard.

How I looked#

  • I copied the app and confirmed that all 658 files, folders and links in the copy matched the original.
  • I unpacked app.asar, the archive Electron apps ship their JavaScript in. That gave me 480 files. I parsed and indexed them: 124,871 functions, no source maps, so the original names are gone.
  • I opened the native helpers in Ghidra. I decompiled all 28 functions in the 1Password launcher and indexed about 4,600 functions in the Computer Use helper.
  • I did not run anything. No app, background process or helper was launched.

So when I say "the app does X," I mean the code says it does X. I didn't watch it happen.

There's also a hard limit. The model loop, the cloud computer, the scheduler and the server storage aren't in the app. I can see what the Mac sends and what it expects back. I can't see what the server does with it.

The pieces#

YOUR MAC                                      REMOTE (not in this app)
Grok Bot.app
  main process       (windows, secrets,       Cursor's API (api2.cursor.sh)
                      updates, supervision)     accounts, agents, messages
  renderer           (React UI, sandboxed)
  coordinator        (routes traffic)  <---->  cloud computer ("box")
                                                where the agent does its work
  local-exec daemon  (runs approved work  <----  "can I run this on the Mac?"
                      on your Mac)
  Computer Use helper (Swift: screen, clicks, Messages)
  1Password launcher  (C)
  WebAuthn signer     (Rust: security keys)

A quick detour: the name tags#

The product says Grok. The code has other names on it.

  • The bundle ID is com.anysphere.sand. Anysphere is the company behind Cursor.
  • package.json calls the package sand, with author SpaceXAI and homepage https://cursor.com.
  • The default API is api2.cursor.sh. Sign-in and metrics go to other cursor.sh hosts.
  • The update channels are named sand, sand-nightly and sand-dogfood.
  • The internal message channels start with dune-rpc:. The base error class is SandDomainError. Someone there likes deserts.
  • The Computer Use helper's bundle ID is co.anysphere.grok-bot-computer-use. Note co, not com.
  • That same helper still contains the string "Cursor is controlling this Mac."
  • The default chat model string is grok-4.5. Voice calls connect directly to wss://api.x.ai/.
  • The app includes the generated API definitions for Cursor's background agents: merging pull requests, Linear and Slack follow-ups. I found no Grok Bot code calling them. My guess is shared code generation in one big repository.

None of that is a scandal. It does tell you this was built fast on top of an existing platform.

Sending a message without sending it twice#

The first real problem the app solves is boring and important. You hit send. The connection drops. Did the server get your message?

If the app just resends, you might get two messages and two agent runs. If it doesn't resend, your message might be lost.

Grok Bot handles this with a nonce, a random ID the app creates for each message. The server replies with one of four outcomes: ACCEPTED_BOX, ACCEPTED_TEMPORAL, DUPLICATE or REFUSED. If the send fails in a way that leaves it unclear, the app doesn't resend right away. It asks the server whether it has a record of that nonce, and only resends if the answer is no.

send(message, nonce)
  accepted or duplicate?   -> done
  failed, but unclear?     -> ask server: "do you have this nonce?"
    yes                    -> don't resend
    no                     -> send once more

When your message echoes back in the transcript, the app matches it by the same nonce so it doesn't show twice.

The two "accepted" outcomes point to two places an agent can run: a box and something called temporal. The name suggests Temporal, the workflow engine, but that's a guess. The server code isn't here.

Four processes, each with a job#

The main process owns windows, secrets, updates and the other processes. In 0.18 it was one 18.5 MB file. In 0.47, main.cjs is an 11,615-byte loader. It runs main-core.cjs, waits for startup to reach a checkpoint, then runs main-app.cjs, using V8's code cache so later launches are faster.

The renderer is the React 19 UI. The 0.18 article said the main window ran without Chromium's sandbox. In 0.47 the window is created with sandbox: true, context isolation on and Node integration off. Everything privileged goes through a small preload bridge, and sensitive calls are only accepted from the app's own top-level frame.

The coordinator is an Electron utility process that routes traffic between the UI, the servers and the local daemon. It talks to the main process over three message ports. If it crashes, main restarts it with a policy that's easy to read:

{ name: "sand-coordinator-relaunch", initialDelayMs: 250,
  backoffFactor: 2, maxDelayMs: 1e4, mode: "until-signal" }

Start at 250 ms, double each time, stop growing at 10 seconds, keep trying.

The local-exec daemon is the interesting one. It's the Grok Bot binary launched with ELECTRON_RUN_AS_NODE=1, so it runs as plain Node. It's detached from the app. When you quit normally, the app sends SIGTERM, checks 40 times at 100 ms intervals, then sends SIGKILL. When the app restarts for an update, it writes an update-lease file first, and the daemon keeps running.

When the agent wants your Mac#

The agent lives remotely. Sometimes it needs your actual machine: run tests in your repo, read a file, send an iMessage. Those requests go to the daemon. There are seven kinds: exec, upload, download, Messages operation, Messages consent, retire approval and cancel.

There are two ways requests arrive. The older one is Server-Sent Events, a long-lived HTTP connection the server pushes events down: GET /local-exec/requests in, POST /local-exec/responses out. The newer one is a set of methods on aiserver.v1.GrokBotService: get a credential, watch for work, poll for requests, submit results. The limits are written right into the code:

  • 8 requests running at once, 64 waiting
  • polls return up to 64 requests
  • result batches cap at 256 frames or 8 MiB
  • files move in 4 MiB chunks
  • a watch connection counts as stalled after 35 seconds, with a backup poll every 30

Permission is not an on/off switch#

The settings show three choices: Always allow, Ask every time, Never allow. The default is ask. But the decision the daemon makes combines a lot more than that setting:

server policy says never?             -> deny
approved explicitly by the server?    -> allow
standing permission covers it?        -> allow
Messages grant for this recipient?    -> allow
your setting is never?                -> deny
your setting is always?               -> allow
exact approval for this action+target -> allow
otherwise                             -> ask

(That's a simplified version of what the code does, not the code itself.)

Server policy acts as a ceiling. A team admin can switch off "always allow," and the UI has a message for that while the policy loads.

My favorite design choice is what "ask" does. It doesn't pause a half-started process while waiting for your click. The daemon runs nothing and reports back. You approve in the app, the approval goes through the server, and later a new request arrives already marked as approved. Nothing sits half-run.

Approval records hold an ID, the action, the target, and optionally a resource path and account. There's no timestamp. The 0.18 article said approvals expire after ten minutes. I couldn't find an expiry in 0.47's daemon. The server can revoke an approval, though, so the accurate claim is "no expiry on the Mac side that I could find," not "approvals never expire."

Files are contained. The shell isn't.#

File tools are limited to a root folder: SAND_LOCAL_EXEC_ROOT, then SAND_AGENT_PROJECT_DIR, then your home directory. Paths are resolved with realpath to block the common symlink escapes, and files cap at 100 MiB.

The shell is different. A command can run in any existing absolute directory, and the root doesn't restrict it. That isn't a bug. It's what the approval system exists for. But "file tools are sandboxed" and "the agent is sandboxed" are very different sentences.

There's also a sandbox policy named insecure_none. It shows up 15 times in the daemon, including as a default policy object, plus an error message suggesting you use it when a helper binary is missing. I couldn't prove from reading the code which policy a real command gets. That one needs a live test.

A small shell trick#

The agent's shell keeps its state between commands. After each command, the daemon saves aliases, functions and shell options, then restores them before the next one. Except for three:

grep -vE '^set [-+]o (errexit|nounset|pipefail)$'

If one command turns on "exit on the first error," that setting doesn't carry into the next command. My guess: one strict script shouldn't quietly break every command the agent runs after it.

Where it's fragile#

The daemon separates getting a request from finishing it. An ACK means "I received this," not "I did this." That's the right idea. The details have gaps.

  • Duplicates are only caught while running. If the same request arrives twice while the first copy is running, the second is ignored. If it arrives after the first one finished, it can run again. Messages keeps a memory of 64 recent sends, so an older duplicate could send again after enough new messages.
  • Results live in memory. The daemon removes a batch of results from its outbox, tries to send it three times, 200 ms apart, then logs that it's dropping the batch and moves on. If the daemon restarts, unsent results are gone too.
  • Partial acceptance isn't retried. If the server accepts part of a batch, the rest isn't put back in the queue.

So in the bad case, the agent runs a command on your Mac and never hears the result. Or it runs it twice. I'm reading this from code, not from a failure I watched, but the paths are there. If I built this, results would sit in SQLite until the server confirms them, and duplicate detection would survive restarts.

The native helpers#

Computer Use#

This is a separate signed Swift app with no Dock icon. Only this helper asks for Accessibility, Screen Recording, Apple Events and Contacts permissions. The main app's Info.plist doesn't request any of them. That's a clean split: the Electron app never holds the powerful macOS permissions itself.

The app talks to it over a Unix socket that only your user can open. The helper checks that the caller runs as the same user before it checks any code-signing policy. It exposes 14 tools: screenshot, click, drag, scroll, type, press key and so on.

The best detail is how it handles clicks. A click based on an old screenshot is dangerous, because the window might have moved and you'd click the wrong thing. So each screenshot comes with a coordinate token. Before a click, the helper compares the window's frame, display and scale to what they were at screenshot time, within 1 point. If anything changed, it refuses with "Coordinate token is stale."

The symbol names also describe a remote-control lease. When control is revoked, any keys or mouse buttons it was holding get released. You don't want an agent to lose control with Shift still held down.

Messages go out through osascript, not by writing to the Messages database. The helper checks for Full Disk Access before it reads that database.

The 1Password launcher#

This is a 136 KB C program. Its only job is to run the 1Password CLI safely. In order, it:

  1. Resolves the path and opens the file without following symlinks.
  2. Requires a regular file.
  3. Checks the code signature through the open file (/dev/fd/N): identifier com.1password.op, 1Password's team ID.
  4. Accepts only five exact command shapes: version, account list, vault list, vault create, and creating a read-only service account.
  5. Validates every argument. Service account lifetimes must end in s, can't start with zero, and can't exceed 365 days.
  6. Removes every inherited OP_* environment variable and sets three of its own.
  7. Applies a sandbox profile that blocks reading 1Password's app data folder.
  8. Rechecks the file's device, inode, size and timestamps.
  9. Calls execve with the path.

That last step is the catch. It verified the open file, then runs the path. There's a tiny window where the file could be swapped. Someone would already need to be able to replace that binary, and macOS doesn't offer fexecve to run an already-open file, so this may be the best practical option. It's still worth knowing where the gap is.

The WebAuthn signer#

This is a Rust binary that talks to USB security keys over IOKit. It reads one JSON request on stdin, writes the result on stdout, and reports events like pin-required and presence-required on stderr.

Why is it here? The agent's browser runs on a remote computer. Your security key is plugged into your Mac. When a site in that remote browser asks for a passkey, the desktop app acts as the USB port. I didn't fully decompile this one, because Ghidra's Rust support wasn't installed.

What changed since 0.18#

  • main.cjs went from one 18.5 MB file to a small loader and split bundles.
  • The main window is now sandboxed.
  • GPU acceleration is on by default on macOS.
  • The desktop app no longer ships the remote host code (host-main.cjs). The 0.18 article could read the cloud side from the app. You can't anymore.
  • Message channels moved from sand-rpc: to dune-rpc:.
  • Deep links went from three routes to eight, across grokbot:// and sand://.
  • Secrets moved to user-secrets.json, encrypted with Electron's safeStorage. When needed, the app decrypts them and sends the whole set to the cloud computer: up to 100 variables, 96 KiB total.
  • A second, credentialed transport for local requests was added.
  • I couldn't find the ten-minute approval expiry.

What I couldn't see#

  • How the server builds prompts, picks a model or runs the tool loop
  • What the cloud computer's image looks like and how tenants are isolated
  • How scheduled automations handle overlaps, missed runs or retries
  • Where connector OAuth tokens are stored
  • Which of the two local transports a normal account actually uses
  • What the model actually sees: screenshots, accessibility data or both

Answering these would take a live test in a throwaway macOS account with the vendor endpoints blocked, or the server code. Searching the desktop app more won't get there.

What I'd build#

After the teardown I wrote a plan for my own version. It doesn't copy the private code. It copies the boundaries.

  • The server owns the durable records. Messages, runs, requests, approvals and results live in Postgres. The desktop app can be replaced.
  • The local daemon stays narrow. Every action on the Mac is a typed request with a unique ID, an explicit authorization decision, a deadline and a structured result.
  • Receiving a request and finishing it are separate steps, and both are saved. Results stay in a local SQLite outbox until the server confirms them.
  • Approvals are tied to an account, a machine, an action and an exact target, and they expire.
  • File access goes through a pre-opened root folder, not string checks on paths.
  • Server policy is the ceiling. Local settings can only take permissions away.

The first milestone uses no model at all. A fake worker streams progress and a result, and the tests prove that nonce dedup, recovery after a dropped connection, cancellation and a service restart all work. If that isn't solid, adding a model just gives you a more expensive way to run a command twice.

The model gets all the attention. The code that decides whether it can run rm on your laptop is the part I'd want to get right.

Related