The newsletter for ai builders of all levels. Mini-tutorials, tool reviews, and lay of the land from an exited founder turned investor and forever tinkerer.
Hey folks,
Over the weekend, the 'evals’ debate really took off on Twitter. Debates like this have a ton of “is a burger a sandwich?” nonsense, but the big question is: should you, as an AI product builder, have a testing process for how well your product performs on certain tasks? Of course, the answer is yes, and of course, the usual caveat applies: both too little or too much of this testing are bad. I plucked some snippets from people who wrote good stuff about evals this weekend:
Shreya (In defence of evals145): Expertise lets you avoid static metrics. Dogfooding (using your own product) as an expert regularly and updating it based on the vibes is evals. To paraphrase this: “we don’t do evals” is mostly a misnomer like “our sci-fi movie has no CGI.”
Alex (Evals are a scam)139: You don’t outsource evals to someone with zero expertise about your product. You can use their tooling to measure something, but don’t let them tell you what to measure. Most evals companies sell logging, observability and complexity beyond that.
other posts:
What are evals, and who needs them?251 - Ben Hylak, CTO Raindrop (A/B testing product)
A/B testing can’t keep up with AI159 - Ankur Goyal, CEO Braintrust (evals product)
You can now Branch chats in ChatGPT167. Click on the three dots after any response to create a copy of your chat up until that message, where you can talk about something different. I assume this will be very useful for hashing out a product idea in a chat and then branching it once you want the model to code/create PRDs, etc.
We keep hearing of the mass movements from Claude Code to Codex - but no doubt that’ll reverse with another Claude model. But if you want to get your hands on a new CLI tool that supports gpt/claude/gemini models - add a reply in this thread294 and I’ll get you access.
An early tester had this to say “am i allowed to tweet about the cli yet? it definitely feels much nicer than codex for gpt-5 agents”.
Kimi K2 has a new variant161 that’s better than Opus 4.1 on Terminal-Bench and as good as Sonnet 4 across other software engineering benchmarks.
OpenRouter has two new stealth models106 - Sonoma Sky Alpha and Sonoma Dusk Alpha. 2M context window and likely as good as Opus and Sonnet. Guess is that they are from xAI.
MCPs now have an official Registry476 - An open catalogue and API for publicly available MCP servers to improve discoverability and implementation. Smithery132 (portco) now supports MCPs hosted from anywhere. Listen to Henry (founder) talk about MCP95.
How are you evaluating your AI outputs? Learn how the experts quickly and accurately evaluate AI using LLM judges250. Enjoy 70 pages of content on how to automate evaluations using advanced techniques, including practical frameworks for building your own LLM judges. Get the free eBook250!*
How to code with Droids423 - step-by-step guide for how artists, designers, writers, and more can create software.
How we built an interpreter for Swift108.
Build an AI life co-pilot719 with Claude Code in 25 minutes.
In the age of AI, young founders aren’t waiting to grow up.266
The bear and bull case for local models239 in just 4 basic graphs.
239Fenic86 - OSS PySpark-inspired DataFrame library for LLMs. Run semantic joins, batch inference, and transform markdown & transcript to insight.*
Operate251 - Precision-built CRM designed for sales and built for founders.
Whisper479 - Desktop AI that sees your screen and delivers everything proactively.
Notte189 - Build and deploy agents that work on the web without breaking.
Compound by Groq154 - Use open models with a complete set of agentic tooling, including web search, code execution, browser automation and more.
Oasis 2.0164 - Re-skin Minecraft in real-time, 1080p, 30 fps.
NotebookLM186 got Flashcards, Quizzes, more report templates and new voices for Audio Overviews.
Rork for iOS121 - make apps from your phone (launched on Product Hunt37)
List of mini tools for everyday work443 (by Simon Willison)
BuildKit 2.0197 - shadcn for AI tools. Build AI tools and MCPs in minutes.
Twiggy118 - Let your cursor agent see your entire codebase's structure in real-time.
OpenAPI to MCP158 - Convert any server described with OpenAPI into an MCP endpoint!
Open-source example of an end-to-end vibe-coding platform.131 (demo58)
SemTools70 - a toolkit for parsing and semantic search in the CLI. (read more34)
Stagehand Agent66 (browser automation) can now use MCP tools. (examples46)
Cursor will soon support custom /slash commands54.
Codex CLI now has web search85. Enable it with --search flag.
Story Arc Engine204 - Tweak parts of a story to see how plot changes trickle down to the rest of the story.
Nano-banana browser380 - Generate websites (screenshots) based on the URLs.
OpenAI is
a) planning to build a job platform139 to match talent with businesses using AI
b) backing a full-length animated film97 to tap into Hollywood.
sfcompute is hiring48 for systems & networking engineers.
I hate my friend215 - review of the “friend” pendant by Wired.
A new startup, Alterego155, claims it can capture silent speech (mouthing) and let you take notes, reminders or talk to other people using their device.
That’s it for today. Feel free to comment and share your thoughts. 👋
📷 thumbnail creds: @keshavatearth29
Any model in your CLI
MCP tools, secret models and an evals debate
3,150 readers clicked at least once. 53 links, 53 with clicks. Heat is relative to the most clicked link in this issue.
Substack · essay