←Ben’s Bites Waybackbensbites.com/p/caught-cheating23 Jul 2026sent to 168,867opened 32.6%clicks 12,933unsubs 184← 21 Jul · 28 Jul → · random

Hey folks,

Another day in the Vercel vs Cloudflare feud: this time they are fighting over whose AI gateway is faster299.

Here’s the result of last week’s poll:

I guess everyone likes Fable more.


Ben’s Bites is brought to you by Metatate181

Most agentic data work is quietly propped up. The agent returns something plausible, and every answer gets checked in case it's plausibly wrong. Metatate gives agents the rules they're missing: which revenue definition to use, which policy applies, which records to trust. Try it for free181.


Headlines

OpenAI’s models hacked Hugging Face202 - by accident. OpenAI was testing its models (Sol and an unreleased one—GPT-6??) on a cybersecurity benchmark with safety refusals switched off. The models found an unknown bug in the test environment, and a few more, and eventually broke into Hugging Face’s production servers.

And why? To steal the answers to the test.

Both security teams caught it, the bug has been reported, and both139 sides124 have published what they know. Hugging Face says open models were a key part of its defence - its team fought back with GLM-5.2198.

Simon’s write-up630 is always a good read.

Google released some new Gemini models158 - Gemini 3.6 Flash gives you the same 3.5 Flash performance with a) more efficient token usage and b) a slightly lower cost for output tokens. Gemini 3.5 Flash Lite116 is a big upgrade over 3.1 Flash Lite, but again, comes at a ~30% price increase. And 3.5 Flash Cyber is a security model for governments and trusted partners only.

There are two things you’d want to use the Gemini Flash models for:

  1. Fast speed - if you want a chat-only model with a good context window, it’s a good model to use in your apps.

  2. Vision - if you want to use the model for a lot of visual analysis, these are the models you need to pick.

That’s it tbh.

Substack will now tell you what’s AI-written213. It’s adding AI detection through Pangram - you can scan posts, replies and comments in the app for an estimate of how much was written by a human.

Pangram has sent the claim of “AI detectors don’t work” for a toss—it works wayyy better than most. But I’m still unsure about how reliable it is. I tested it on some pieces of 100% AI-written content (though that content was a result of a complex pipeline built over months), and I got 100% human scores on most of them.

— Keshav

A relevant experiment155: given access to Pangram’s API, Grok 4.5 rewrote an essay 14 times until it passed as human-written, then built a website211 showing off all 14 attempts. GPT-5.6 Sol and Fable 5 refused146 to game the detector.

Cursor also launched a router186 - it picks which model handles each request, claiming 60% lower cost with similar quality of responses. The router lets you select between three options: “cost”, “intelligence” or “balance”.

Routers also have a history of poor performance in real usage. OpenAI’s router, which routed requests between GPT-5’s no-thinking and thinking variants, didn’t fare well. Since then, most companies that build these routers are the ones who sell inference to devs (like OpenRouter), so you don’t really get the feedback loud and clear. With Factory, Ramp and now Cursor making these available directly to users, I hope we’ll get more feedback on whether the routers actually help or if they add too much latency/degrade performance by a lot.

Router or not, we might see more companies adopting this cost/balance/intelligence trio to minimise the headache of choosing the “correct” model for a task.

Claude can now learn a skill by watching you402. Record your screen while you do a task, talk through it as you go, and Cowork turns it into a skill Claude can run again - same idea as Codex’s Record & Replay from last month. It’s under “Record a skill” in the desktop app, on Pro, Max and Team plans.

Also: Claude Code got an iOS simulator panel204 and a security plugin161, plus you can now ask Claude201 about how people actually use AI at work.


Quick links

  • BUZZ337 - Jack Dorsey’s open-source group chat for teams of people and agents, aimed squarely at Slack and GitHub. (tweet134)

  • Replit’s mobile app160 got a full redesign - build and ship from your phone on iOS and Android.

  • Slate277 is a voice journal where the AI never leaves your iPhone - transcription, reflection and storage all happen on-device.

  • AFK285 - macOS app to transcribe multiple-hour recordings without sending any of them to a cloud server.

  • OpenAI Presence173 - voice and chat agents for enterprises that answer questions, use company systems and hand over to people when needed.

  • Fable found a 15-30% memory improvement183 in Next.js’s bundler, nearly autonomously.

  • Obliterate, don’t automate250 - USV on backing AI companies that replace markets entirely instead of making them a bit more efficient.

  • YC’s new startup wishlist336† - AI moving into the physical world: education, healthcare, defence, finance and factories.

  • Dana155 - Applied Intuition’s agentic development environment for physical AI: cars, robots and machines.

  • The Claude Code team on how Claude Code gets built239 - annotated interview.

  • Why the team at Factory refunded its first few customers154 (millions in revenue) before Droid CLI took off.

  • Never enough292 - short post on why AI makes the work rat race feel faster but not more satisfying.

  • Devin Outposts176 lets you run Devin on your own machines - a Mac mini, a GPU box in your lab, or a cluster inside your private network.

  • Language Model Builder239 - build a tiny language model yourself, then chat with the thing you made. (tweet115)

  • How the Exe team built a distributed DNS server in ~a week137 and got zero incidents in a month.


Skills section…

  • Frontend Textbooks334 - skill to turn AI’s wall-of-text deep dives into nicely designed HTML books with covers and diagrams. (repo130)

  • /pick-ui-library216 - your agent picks a UI library Emil trusts instead of hand-rolling a toast component or installing an abandoned package. (repo106)

  • and a one-shot prompt183 to build your own "codex-style" app for Pi, accessible from desktop and mobile.


Afters

I think loops were a short-lived patch for models that couldn't reliably keep working on long problems until they hit a defined goal Fable and GPT-5.6 (and probably Kimi K3 as well) can just do that out of the box, unassisted

@simonw184

I need a magic lens that extracts colors 🔎

@tldraw213


* sponsors who make this newsletter possible :)
Wanna partner with us for the next quarter?
Email us at shanice@bensbites.com100 or k@bensbites.com85

Caught cheating
models, writers, and routers

less more clicks · number = clicks † = site gone, link opens the Wayback Machine

Most clicked in this issue

630simonwillison.net/2026/Jul/22/openai-cyberatt…
402x.com/claudeai/status/2079595988998554047
337buzz.xyz
336ycombinator.com/rfs
334x.com/tareqismail/status/2079544423688425699
299x.com/dok2001/status/2079961408490381433
292dark.ronacher.eu/2026/7/21/never-enough
285optionafk.com
277x.com/elirousso/status/2079594911637094442
250x.com/mignano/status/2079935862372958326
239languagemodelbuilder.com
239simonwillison.net/2026/Jul/21/cat-and-thariq

2,960 readers clicked at least once. 48 links, 48 with clicks. Heat is relative to the most clicked link in this issue.

Substack · essay