The Browser Agent That Picks Instead of Writes

Jev Ultrafast turns web automation into a multiple-choice problem — one decision model, one numbered list, one network round trip per step — and the speed follows.
Every few weeks a browser agent demo goes viral, and the recipe is usually the same: a big vision model stares at screenshots, emits some JSON, a driver executes it, repeat until the token budget or your patience runs out. Jev Ultrafast, the new project from the browser-use organization, took a different bet and got the same attention spike — reportedly more than 14,000 GitHub stars within five days of launch, according to a LinkedIn post that walked through the demo. The bet: stop asking a language model to write browser actions, and start asking it to pick them.

The headline number is doing a lot of the hyping. A one-way Google Flights search — Zürich to London, one adult, economy — completes in 7.1 seconds of wall-clock time, including model calls, generated text, browser work, and Google’s own loading waits. The same policy opened a requested Wikipedia article in 2.8 seconds and passed a local hotel search-and-filter task in 1.9 seconds. At a reported cost of $0.0039 per flights run, with a median decision time of 178 milliseconds, this stops being a demo and starts being infrastructure you could plausibly run all day.
The idea: shrink the question until the answer is cheap
The core move is architectural, and it’s worth slowing down for because most of the speed is hiding in it.
At each step, the agent reads the visible DOM and produces an element table — a numbered list of every interactive control, with its role, name, and current value. Against that table, a decision model (TypeSafe’s Jev) answers a single request: which operation, and which target. The operations are a closed set — click, type text, select, scroll up, scroll down, wait, done, blocked — and only operations the target actually supports are offered. If the operation is a click, only clickable elements appear in the click-target head. If it’s typing, only the typeable ones show up.
This is the part that displaces the incumbent design. A screenshot-driven agent asks a vision model an open-ended question — “here’s a picture, what should happen next?” — and pays for the generality in latency, tokens, and hallucinated coordinates. Jev Ultrafast asks a classification question with a bounded answer space. The model isn’t composing a selector or a snippet of JavaScript; it’s choosing an operation and an index. Model output never becomes selectors, coordinates, shell commands, or executable code, which is as much a safety property as a speed property.
Text is the one place generation is unavoidable — you can’t pick “Zürich” out of a list of page elements — so it’s quarantined. A small LLM (the demo uses Mercury 2.5 with reasoning disabled) writes text only when the chosen operation is TYPE_TEXT, and its output must parse as a small JSON object before anything gets typed. Everything else is selection.
The request structure is also speculative in a way that saves a round trip: the operation and target heads share one observed state and one network call. Two decisions, one request. On a task where the loop runs dozens of times, that adds up.
The boring part is where the speed actually lives
A fast decision model with a slow executor is a sports car on a gravel road, and the repository’s own accounting makes this plain. In six alternating runs with identical models and settings, both the old and new versions passed three out of three attempts — but median task time dropped 25 percent (9.45 seconds to 7.09), while median browser protocol calls collapsed from 1,092 to 101. An order of magnitude fewer conversations with Chrome is where the real engineering went.
The mechanics are unglamorous and mostly invisible in the demo video. One browser call per snapshot reads visible controls, their names, values, and text atomically, keeping references to the actual DOM nodes rather than re-querying. Clicks are validated against the document, form values, target, and nearby context — animation alone doesn’t force a new prediction. Geometry is resolved at execution time, and covered controls are rejected before input. After typing into a combobox, the agent waits for visible suggestions, capped at 200 milliseconds; other interactions get at most two animation frames or 50 milliseconds. Background tabs keep rendering through focus emulation, dodging Chrome’s throttling of hidden tabs. Offscreen article bodies and footers never enter the model’s context.
None of this is clever in the way demo videos reward. It’s the accumulated discipline of not doing redundant work, and it’s the difference between an agent that is fast on paper and one that is fast on a real page with real loading spinners.
The security angle nobody planned for
The constrained action space also lands at an interesting moment for the browser-agent field. A 2025 arXiv paper on the security of browsing agents — with Browser Use itself as the white-box case study — documented prompt injection, domain validation bypass, and credential exfiltration, including a disclosed CVE and a working proof-of-concept. The paper’s proposed defenses read like a wishlist: input sanitization, planner-executor isolation, formal analyzers, session safeguards.
Jev Ultrafast accidentally implements a chunk of that wishlist by construction. A model that can only emit an operation and an index cannot be talked into emitting a selector, a URL, or a script by hostile page content. The executor rechecks page freshness and click occlusion; every executed target is resolved from an observed node. That’s planner-executor isolation in practice, arrived at for performance reasons.
It’s not a complete answer. The paper’s threat model extends to the execution environment and stored credentials, and the Hacker News thread raised its own concern: one commenter who looked under the hood of browser-use flagged default-on telemetry to PostHog, claiming the browser-harness component in particular could leak credentials. That’s a community claim, not a verified finding, but it’s exactly the class of issue the academic work warns about, and it deserves a straight answer from the project.
What the skeptics are saying
The Hacker News discussion was notably more divided than the star count suggests, and the objections are worth taking seriously.
The sharpest one is about the timing boundary. The 7,073-millisecond figure starts after initial page observation — and as one commenter put it, isn’t that the part that takes the most time? The README is honest about the boundary and about what the measurement includes, but a headline number that excludes the first, often slowest, step invites exactly this question.
Then there’s the benchmark itself, which the README concedes is “three repeats of one task on one browser profile, not a general reliability benchmark.” A commenter on the LinkedIn thread asked the right follow-up: what is the success rate across a task set, and what verifies that DONE is true — that the results on screen actually match the goal? The project’s answer is that a DONE choice still requires independent outcome verification, and the flights example does supply a fresh check of the route, date, and visible results. But that verification is per-example scaffolding, not a property of the agent.
The dependency question cuts deeper. Jev is a cloud model, and several commenters wanted it local — one wished aloud for something running on consumer hardware, complaining about paying “AI tax to gate keepers for each use.” Another dismissed the whole approach as “almost like a smart switch statement,” which is unkind but not entirely wrong as a description of what a bounded classifier is. The counterargument, made by another commenter, is that fast, cheap classification is precisely the point — applicable to robotics and real-time decision-making, an order of magnitude faster and cheaper than generative alternatives. Both readings are true; they just value different things.
There were also practical complaints: users reporting the demo button erroring, a suggestion that the Google Flights task could be done by constructing a protobuf-encoded link directly (true, and missing the point — the demo is a stand-in for sites with no API), and a note from an AI Agent Store profile that the repository snapshot showed just three commits. This is a young project wearing a viral moment, and it shows.
Where it sits in the bigger picture
The timing is not accidental. Enterprise interest in agentic AI is surging — IBM’s research has 80 percent of executives increasing agentic AI investment, with spending projected to nearly triple by 2027, and BCG frames agents as the observe-plan-act loop finally delivering end-to-end transformation. But most of that energy has gone into planning and reasoning — the expensive, open-ended parts. Jev Ultrafast is a bet on the opposite end: the act step, made so cheap it disappears from the budget.
That reframes the use cases. At $0.0039 a run, browser automation stops being a heavyweight orchestration problem and becomes something you can fire off casually — search on real sites, form filling, working with sites that expose no API, prototyping agents cheaply enough to run them often. The AI Agent Store profile categorizes it as a development framework aimed at engineers studying goal-driven browser automation, which is the honest framing: this is a research artifact with a commercial shadow, not a product.
And the commercial shadow is explicit. The README’s most prominent call-to-action is a waitlist for “ultrafast browser agents in the cloud” — Browser Use’s cloud, running this architecture as a service. The open-source repo is simultaneously a good-faith artifact (the whole loop fits in a handful of readable files, with offline tests and published measurement boundaries) and a marketing surface for the hosted version. That tension is normal now, but it explains some of the HN grumbling about gatekeepers.
The open questions
Three, roughly. First, generalization: the element-table approach handles common HTML and ARIA controls, but shadow roots, frames, canvas, uploads, pop-up tabs, nested scrolling, and arbitrary keyboard widgets are all explicitly outside this MVP. The web is mostly the messy parts. Second, dynamism: as one LinkedIn commenter noted, the interesting test is how well a re-enumerated action space holds up on pages where the available controls change after every interaction — Google Flights is comparatively polite. Third, trust: whether the DONE signal can ever be verified by the agent itself rather than per-task scaffolding, and whether the telemetry concerns get addressed with the same rigor as the click validation.
The repo’s own “Evidence and limits” section is unusually candid for a project riding a hype wave, and that candor is the best reason to take it seriously. Jev Ultrafast may or may not be the future of browser agents. But it has demonstrated something concrete: that most of what we’ve been paying vision models to do on web pages, a classifier with a numbered list can do in 178 milliseconds. That’s not a revolution. It’s just arithmetic, and it’s compelling.
Sources
- AI Agents Landscape
- AI Agent Use Cases
- What Makes Jev Ultrafast So Fast at Browser Automation
- The Hidden Dangers of Browsing AI Agents
- AI Agents: What They Are and Their Business Impact
- Jev Ultrafast Browser Agent for Fast Web Tasks
- Best Open Source Autonomous Web Browser AI Agents for ...
- Adoption & Usage of Open-Web AI Agents | by Cobus Greyling
- Jev Ultrafast - AI Agent | Pricing, Reviews & Alternatives
- Autonomous Web Agents Landscape Map
- Exploring AI Agent Adoption: How Enterprises Can Identify ...
- Jev Ultrafast: A browser agent with a dynamic, indexed ...