←Ben’s Bites Waybackbensbites.com/p/any-model-in-your-cli9 Sep 2025sent to 143,353opened 38.6%clicks 68,886unsubs 188← 4 Sep · 11 Sep → · random

The newsletter for ai builders of all levels. Mini-tutorials, tool reviews, and lay of the land from an exited founder turned investor and forever tinkerer.


Hey folks,

Over the weekend, the 'evals’ debate really took off on Twitter. Debates like this have a ton of “is a burger a sandwich?” nonsense, but the big question is: should you, as an AI product builder, have a testing process for how well your product performs on certain tasks? Of course, the answer is yes, and of course, the usual caveat applies: both too little or too much of this testing are bad. I plucked some snippets from people who wrote good stuff about evals this weekend:

  • Shreya (In defence of evals145): Expertise lets you avoid static metrics. Dogfooding (using your own product) as an expert regularly and updating it based on the vibes is evals. To paraphrase this: “we don’t do evals” is mostly a misnomer like “our sci-fi movie has no CGI.”

  • Alex (Evals are a scam)139: You don’t outsource evals to someone with zero expertise about your product. You can use their tooling to measure something, but don’t let them tell you what to measure. Most evals companies sell logging, observability and complexity beyond that.

  • other posts:

You can now Branch chats in ChatGPT167. Click on the three dots after any response to create a copy of your chat up until that message, where you can talk about something different. I assume this will be very useful for hashing out a product idea in a chat and then branching it once you want the model to code/create PRDs, etc.

We keep hearing of the mass movements from Claude Code to Codex - but no doubt that’ll reverse with another Claude model. But if you want to get your hands on a new CLI tool that supports gpt/claude/gemini models - add a reply in this thread294 and I’ll get you access.

An early tester had this to say “am i allowed to tweet about the cli yet? it definitely feels much nicer than codex for gpt-5 agents”.

Kimi K2 has a new variant161 that’s better than Opus 4.1 on Terminal-Bench and as good as Sonnet 4 across other software engineering benchmarks.

OpenRouter has two new stealth models106 - Sonoma Sky Alpha and Sonoma Dusk Alpha. 2M context window and likely as good as Opus and Sonnet. Guess is that they are from xAI.

MCPs now have an official Registry476 - An open catalogue and API for publicly available MCP servers to improve discoverability and implementation. Smithery132 (portco) now supports MCPs hosted from anywhere. Listen to Henry (founder) talk about MCP95.

How are you evaluating your AI outputs? Learn how the experts quickly and accurately evaluate AI using LLM judges250. Enjoy 70 pages of content on how to automate evaluations using advanced techniques, including practical frameworks for building your own LLM judges. Get the free eBook250!*

*sponsored

🌐 What I’m consuming

Image239

⚙️ Tools to tinker with

  • Fenic86 - OSS PySpark-inspired DataFrame library for LLMs. Run semantic joins, batch inference, and transform markdown & transcript to insight.*

  • Operate251 - Precision-built CRM designed for sales and built for founders.

  • Whisper479 - Desktop AI that sees your screen and delivers everything proactively.

  • Notte189 - Build and deploy agents that work on the web without breaking.

  • Compound by Groq154 - Use open models with a complete set of agentic tooling, including web search, code execution, browser automation and more.

  • Oasis 2.0164 - Re-skin Minecraft in real-time, 1080p, 30 fps.

  • NotebookLM186 got Flashcards, Quizzes, more report templates and new voices for Audio Overviews.

  • Rork for iOS121 - make apps from your phone (launched on Product Hunt37)

  • List of mini tools for everyday work443 (by Simon Willison)

*sponsored

🥣 Dev dish


🍦 Afters


That’s it for today. Feel free to comment and share your thoughts. 👋

📷 thumbnail creds: @keshavatearth29