The Browser, and Agents That Try What They Built
Freddy's search box works. He knows because the agent said so, and because the tests pass. He has not seen it work, and neither has the agent. This chapter closes that gap: a browser next to the conversation, for him and for the agent.
The problem
Passing tests tell you the code does what the tests expect. They say nothing about a button that sits too low, a heading in the wrong color, or a flow that needs a click the tests never make. Agents have the same blind spot, with one addition: they cannot see the page at all unless you describe it.
The hard way
A browser in another window, a screenshot saved and attached, the console errors copied by hand, and a description of the problem that starts “the second button, the blue one…” Each piece is fine. Doing all of them for every change is how visual checking quietly stops happening.
A browser in every thread
Each thread has its own web browser in the Browser
tab of the tools panel (⇧⌘B). Use it to
look at the app you and the agent are building, next to the
conversation, without switching windows.
-
Type an address and press
Enter.localhost:3000,myapp.testand other local addresses open overhttp; other addresses with a dot open overhttps; anything else searches DuckDuckGo. - Before a page is open, the tab lists the development servers running from the thread's folder. Click one to open it. Clicking a server address under Local servers in the Context tab opens it here too.
- The toolbar has back, forward, reload and Open in your browser, which hands the page to your default browser.
It uses Safari's engine (WebKit). Pages share cookies and storage across threads, as in one Safari window with several tabs, and closing a thread's tab closes its browser.
Pointing at an element
Press the target button in the toolbar and click
the thing on the page. While picking, the element under the
pointer is outlined and the page does not react to clicks;
Esc or the button again stops. The element is
attached to your message as a picture of it with a little of its
surroundings, and its details: a selector that matches only it,
its position and size, its text, its computed styles (layout,
spacing, colors, fonts) and its HTML. Write what is wrong,
this button sits too low, and send. It works on any page,
since it only adds to the message you send, and password field
values are left out.
This is the end of “the second button, the blue one.” The agent does not need to guess which element you mean; it has the selector.
Errors on the page
While a local page is open, Workspace watches it for errors:
anything written with console.error, exceptions
nothing caught, and requests that fail or return an error status.
When new ones appear, a chip above the message box says so:
3 new errors in the browser. Add to
message attaches them to your next message, with the page
address, so you do not have to copy them from the console; the
× dismisses them. Either way they are not
offered again.
If Elyra Grove runs the project as an app, the chip also counts server errors: requests Grove recorded with a 5xx status since the thread was opened, from any browser. Adding them attaches Grove's explanation of each (up to three): the request and its body, the SQL it ran and the mail it sent, and the error log with its stack trace. The page in the Browser tab does not have to be open for these.
And then the part that answers “did the fix work?” When
the turn after such a message ends, Workspace sends those requests
again (grove replay --same-data, so a request that
writes starts from the same data each time) and says in the
conversation whether the fix worked: Replayed POST /checkout:
was 500, now 200 ✓, or still 500. A turn that
fails keeps them for the next one.
Pictures before and after a turn
When the Browser tab shows a local page while you send a message, Workspace keeps a picture of the page from before the turn. When the turn ends it reloads the page, waits a moment for the dev server, and takes another. Both appear under the turn in the conversation, so you see what the change did to the page. Click a picture to open it full size.
This only happens while the Browser tab is on screen, since the
page can only be pictured then. The pictures are kept in
~/.elyra/snapshots and deleted with the thread.
Letting the agent look at and use the page
So far you have used the browser. With the agent gateway on (chapter 14), agents get tools to open a page in their thread's browser and look at it: its structure, elements and their styles, the console, network calls and a screenshot. They can also use it like you would: click (by selector or visible text), fill in fields (by selector, label or placeholder; selects and checkboxes too), press keys, and wait for something to appear. You can ask things like:
“Open localhost:5173, check why the cart total is wrong, fix it, then add two items, apply the code SUMMER and check the total.”
That is the difference between “I changed the code” and “I tried it.” The guard rails are the interesting part:
- The first time an agent wants to click or type in its thread's browser, Workspace asks: Don't allow, Allow once or Allow for this thread. While it is allowed, a bar above the page says so; Take over there withdraws it. Threads in Full access are not asked.
-
Local pages only. Agents can only open, read
and use pages served from this Mac:
localhost,127.0.0.1,[::1], and names ending in.localhost,.testor.local. If the thread's browser shows any other site, the tools refuse it. - No passwords, no scripts. Agents never read password fields, and cannot run their own scripts in the page.
The console and network calls are recorded from when a local page loads, and a screenshot needs the Browser tab to be on screen.
Journeys: a flow that is replayed like a test
A journey is a flow through your app that an agent walked through
in the browser and saved: open a page, click, fill in, press keys,
wait, and what must be on the page when it works. Ask for one, for
example “try the checkout with the code SUMMER and save
it as a journey.” The agent saves the steps it just took
in .elyra/journeys/<name>.json in the project, a
small JSON file you can read, edit and commit with the code.
- Replay. The Context tab lists the project's journeys; ▶ replays one in the thread's browser and Run all every one. Agents replay them too, to confirm a change did not break a flow. Journey passed or Journey failed (with the step and why) appears in the conversation.
- Replay after each turn. Switch it on in the Context tab, and after every turn that changed files (and after the checks from chapter 9 passed) the journeys run. A broken flow goes back to the agent like a failing check, with the same number of automatic fixes. It needs the dev server running and the window open.
-
Addresses are replayed on the site the thread's browser shows,
so a journey saved on
shop.testalso runs on a worktree's own site (chapter 7).
Replaying clicks and types on the page, like the agent's own actions, so an agent asks first unless the thread is in Full access. Notice what has been built across these chapters: checks for the code, journeys for the flow, and both go back to the agent when they fail. Freddy saves the search flow once and stops re-testing it by hand.
What you learned
- How the browser in each thread works, and how to point at an element
- How page errors, and Grove's server errors, reach the agent without copy and paste
- How the before and after pictures work, and when they do not
- What agents can do in the browser, and the guard rails: asked first, local pages only, no passwords or scripts
- How a flow becomes a journey, and how it is replayed after each turn