Then the three newest browser agent papers on arXiv, each one confirmed on its own abstract page instead of trusting the search listing.
17 steps, 256k tokens.
Opus 5 is now available in Hyperbrowser.
It comes close to Fable 5's intelligence at half the price, and that it verifies its own work and keeps iterating until it succeeds.
We paired that with stealth browsers and residential proxies.
Just got scary good↓
We had one Kimi K3 browse 30 pages at once, on Hyperbrowser.
Browserswarm spins up a swarm of cloud browsers that read the web in parallel and stream every page into a single shared context. One brain, many eyes.
428k tokens into one window, zero failures, one cited answer in just a few minutes.
We gave Gemini 3.5 Flash-Lite a multi-hop research task in a Hyperbrowser cloud browser: dig through GitHub's docs, find the REST API rate limits, check whether GraphQL differs, and summarize both.
24 steps. Multiple pages, navigation decisions at every hop, no hand-holding. It finished:
Gemini 3.5 Flash-Lite is cheap and genuinely capable at computer use. Google's most affordable model.
A computer-use model is only as good as the browser you give it. It's now live on Hyperbrowser:
Stealth, proxies, captcha solving, and a live view you watch it work through. The perception is cheap now, we gave it the environment.
We benchmarked GPT-5.6 Sol (high), Grok 4.5 (high), and Claude Fable 5 on 24 browser tasks, run on Hyperbrowser with an identical harness,
3 trials per task.
Fable 5 led overall at 78%, Sol 75%, Grok 4.5 69% and each specialized: Fable was perfect on reading, Sol was best at navigation (94%), and Fable led on forms and logins (89%).
Full task set and harness are open source. The breakdown below ↓
Fable cost 6x more per completed task than Grok ($0.06 vs $0.01).
Sol sat in the middle.
Setup: live websites, one system prompt, one action space, no per-model tuning. Fresh cloud browser per trial, interleaved so no model saw a different web.
Environment and API failures were tagged separately and excluded from model scores.
We put grok 4.5 and opus 4.8 head to head on the same browser task, on Hyperbrowser Sandboxes.
We then asked it to open a page in a real sandboxed browser and pull the title. Grok build on Grok 4.5, Claude code on Opus 4.8, identical setup.
Grok 4.5 came out ahead ↓
GPT-5.6 sol is insane when paired with cloud browsers.
Sol, Terra and Luna are now available on Hyperbrowser.
Openai's most capable models yet, running computer use tasks on cloud browsers and sandboxes.
here's how it performs ↓
7 Followers 116 FollowingA hostname that resolves and a certificate that stays current, in one API call. No domain to buy, no zone to configure, no person in the loop.
61 Followers 2K FollowingCurated AI SEO tools + playbooks for small biz. We also build websites, SEO, and managed AI employees. Tested picks, not hype.
2 Followers 53 FollowingMCA'26 | DSA
Solve DSA on LeetCode, highest Rating 1411
ranked 7th in GFG in my College peers
top 27% in LeetCode Contest 440
Build GenAI and MERN stack
23K Followers 4K FollowingLead for Chrome DevRel @ Google. Progressive Web Apps, Mr Web Intents. Dev of many things: Twollo, Twe2, FriendDeck and https://t.co/C1nmx67K3t https://t.co/jHwOp3rPWH
11K Followers 964 FollowingBuilding @recaplyai | Sharing insights on AI, tools & practical ways to apply AI to your daily work | 📧 Inquiry: [email protected]
32 Followers 2K Followinga desi punk operator and venture builder.
2x founder and built digital literacy products for a 100Mn+ user base
currently building a better for planet tradeco
294 Followers 2K FollowingCo-Founder & Chief Strategist @conquer365 Live Chat Platform for #legalmarketing Data Hacker, problem solver & son, brother, husband, father & tennis player
76 Followers 233 FollowingA new journey in the land of X.
Applying the principles of the “How to get rich” tweet storm by Naval. Using Media and Code as Leverage.
#X #Blockchain #AI