←Ben’s Bites Waybackbensbites.com/p/how-to-see-through-ai-marketing9 Oct 2025sent to 144,013opened 37.5%clicks 8,708unsubs 150← 7 Oct · 14 Oct → · random

The newsletter for the technically curious. Updates, tool reviews, and lay of the land from an exited founder turned investor and forever tinkerer.


Hey folks,

I was up late last night in SF at the meet-up I hosted, and it was so awesome to meet and speak to a bunch of you. I’m saving my full SF thoughts for a post next week.

In the meantime, Google released a preview of its first computer-use model271 based on Gemini 2.5, in partnership with Browserbase124. It’s a good model—it scores decently better than Sonnet 4.5 and much better than OpenAI’s computer use model on benchmarks64.

But benchmarks and evaluations can be misleading, especially if you only go by the official announcement posts. This one is a good example to dig into:

  • This is a model optimised for browser usage, so it’s not surprising that it does better than the base version of Sonnet 4.5

  • OpenAI’s computer use model used in this comparison is 7 months old—a version based on 4o. (side note: I had high expectations for a new computer use model at Dev Day)

  • The product experience of the model matters. ChatGPT Agent, even with a worse model, feels better because it’s a good product combining a computer-using model, a browser and a terminal.

I don’t mean to say that companies do it out of malice. Finding the latest scores and implementation of a benchmark is hard, and you don’t want to be too nuanced in a marketing post about your launch. But we, as users, need to understand the model cycle and the taste of the dessert being sold to us.

Even with all these factors, the new Gemini model definitely passes the smoke test.

The smoke test is just one way we decide what makes it into every newsletter post and what doesn’t. Shanice wrote about it in detail in this post, A day in the life of Ben’s Bites.

With Retool167, you can turn prompts into full-stack internal tools—connected to your data, hosted in your cloud, and secured by your rules. Build easier, deploy safely, and move from idea to production in minutes.*


🌐 What I’m consuming

Image

⚙️ Tools and demos

  • Scout Monitoring’s MCP93 - AI-native monitoring. It feeds performance issues and slow endpoints directly into your AI coding assistant.*

  • Google AI Studio155 now lets you use your voice as an input for vibe coding.

  • ElevenLabs220 launched Agent Workflows and an open source UI library for building voice agents.

  • Grok Imagine72 now uses xAI’s Imagine 0.9 model with audio generation.

  • Opal307, Google’s experimental product for chaining AI steps together with a visual builder, is now open globally (with MCP support coming soon).


🥣 Dev dish

  • Playwright204, the browser automation library, has agents now. One to plan tests, another to generate them and a healer to debug and fix failing tests.

  • Recall137 - Redis-powered persistent memory for Claude (usable as an MCP server).

  • sora-mcp113 - An MCP server to use Sora video generation APIs.

  • You can use any open-source model in Factory AI’s Droid74.

  • Repobench124 - Ranking models for large context reasoning, file editing precision, and instruction adherence for coding tasks. I met Eric yesterday and chatted about how he built this56.


🍦 Afters


That’s it for today. Feel free to comment and share your thoughts. 👋

* marks sponsors that make this newsletter possible :)
Wanna partner with us35? Last few slots left for the rest of the year.