Watch the browser
A panel that shows exactly what the agent's browser shows, live. When the agent hits something only a person can do, like a login or a text code, the owner can take the controls and then hand them back.
Where you see it
When an agent gets stuck in a website, its row turns to the "needs you" color and shows Open the browser. There is also a Watch the browser button in the header of the Office and the Chats drawer.
The panel slides up over the page. Inside is a live picture of the agent's whole browser window. The buttons are Take over, Hand back, and a back button. A small caption explains that you see exactly what the agent sees.
On a narrow phone screen the panel may use an older, simpler page view instead.
What happens, step by step
- An agent is signing in to a supplier's website. The site opens a "Continue with Google" pop-up and then asks for a code.
- The agent's row turns to the "needs you" color. The Office shows Open the browser.
- The owner taps it. The panel opens with a live view of the browser.
- The owner taps Take over. Now the owner's clicks and typing reach the real browser. Anything the agent tries to type is dropped while a person drives.
- The owner clicks "Continue with Google" and types the code from the text message.
- The owner taps Hand back. The agent picks up where the owner left off.
- Closing the panel stops nothing. The browser keeps running behind it.
Each tier (a kind of user, like the owner or a client) has one shared browser. A browser started from any chat shows the same red dot in every chat of that tier.
What powers it
| Part | What it does |
|---|---|
browser_panel.js | The slide-up panel. One shared part, loaded by the Office and the Chats drawer. |
browser_relay.py | Starts the browser in its own screen on the box and picks the right browser tab to watch. |
| noVNC | A screen viewer that runs in the page. It draws the browser's whole screen. |
/api/browser/vnc/... | A live connection that carries the picture one way and the owner's clicks the other. Guarded by the sign-in check. |
_RfbInputGate | Drops the agent's own input while a person is driving. |
watch_browser.py | The agent's own tool for driving and looking at the browser. It can grab a picture of the whole screen to confirm a pop-up appeared before it asks for help. |
Why it works this way
It shows the whole screen, not just the page. Dropdown menus and pop-up windows (like "Continue with Google") live outside the web page. The owner can only help with what the owner can see, so the view shows the whole browser window, pop-ups and all.
Take over is for the steps only a person can do: a login, a text code, a checkbox. The owner does that one step in the real browser, then hands the rest back to the agent.
Connected to
- The Office: where the Open the browser button appears.
- How an agent asks the owner: the other way an agent asks for help.
- Secrets and logins: how logins are kept.
- Services: the sites and accounts agents sign in to.