←Ben’s Bites Waybackbensbites.com/p/1-billion-chatgpt-users30 Jul 2026sent to 169,258opened 29.9%clicks 15,260unsubs 128← 28 Jul · 4 Aug → · random

Hey folks,

Following on from Tuesday’s messing around building, I built a few more widgets on my canvas. It pulls my X bookmarks, emails, and todos.

im not looking at files again! building personal software with @tldraw offline is dope

@bentossell559

Before making this video I didn’t have Loom installed (since it’s crap after being acquired), so just told Codex to build me one.

I drew the image on the left, then just sent that prompt and it worked straight away.

So as software and mini tools are getting easier to create, the tools to create them are not...

i dunno why anyone thinks its complicated

@bentossell191

I use Codex as my default app because it’s better than all the others and works the best on mobile. But I just downloaded t3490 because they basically copied the interface and features, but it lets you choose between different agents; codex, claude, cursor (pi + droid soon). It’s made by a reputable developer that you may have seen mentioned here before, so I trust it’s built well.

At the moment, I ask Codex/ChatGPT to be the orchestrator and to go ask claude about something that is design-related. Which is fine-ish but not great as user experiences go.

What I’m noticing by doing this is how well Codex creates prompts for other agents to follow. I just have a normal chat in my session, then say ‘Use X agent to implement this. Watch its progress and send screenshots every time new work has landed.’ To which it sets up its own monitoring every 5 mins or so and updates me in the same thread, with screenshots.

I’m going to try and put together a ‘bites of the week’ email over the next few days to try and summarise all the stuff going on, what themes people are talking about (loops?!, software factories?!, etc) and explain them.

Let me know if there’s anything specific you need to wrap your head around (I may need to too).


Ben’s Bites is brought to you by Brief311

Struggling with fragmented context, slow product decisions, and rework? Brief distills your critical product context into an opinionated graph, then puts a PM agent everywhere you work, (e.g. Slack, Claude Code, email) reducing alignment tax and accelerating cycles. Learn more.311


Headlines

OpenAI used Sol to optimise Sol itself204, cutting serving costs by 20% and making it 15%+ more efficient at generating tokens. And turns out, it also tops the ARC-AGI-3146 benchmark.

Well, there’s a catch: OpenAI says the official ARC-AGI harness hurts Sol’s performance by “forgetting” its reasoning every turn and disabling compaction. Fixing these two things triples Sol’s score from 13.3% to 38.3%, with 6x fewer output tokens.

Re: last week’s fiasco of an OpenAI model hacking Hugging Face - HF published a full replay192 of roughly 17,600 actions taken by the model. METR and Redwood Research155 will also independently review what happened.

Though OpenAI is not out of trouble just yet, a Reuters report154 claims that the same model broke into a customer account at another company (Modal Labs), with rumours suggesting that even more companies were affected.

Anthropic also claimed that Claude Mythos found better attacks on two cryptographic algorithms161, though neither affects systems in use today.

Separately (not at all as a reaction to this general trend, right?), ~1300 people working at leading AI companies (OpenAI, Anthropic & others) want the US government to help “pace the frontier157” of AI development. Kinda expected when the pace picks up, but this time a lot of the “model makers” themselves are in favour of this pause/slowdown.

btw, The Information reports ChatGPT is nearing one billion weekly users198 - a milestone OpenAI originally hoped to hit seven months ago. More from OpenAI this week: Codex Security CLI150, free frontier access for Academic Researchers139, and two new transcription models165.

Grok app builder164 - Grok has a vibe coding interface inside its app now. Create games and apps that can be shared directly to the X timeline. Also see: Drawesome184 - a zero-dependency drawing toolbar for React, built over a weekend with Grok Build.

Pangram 4237 claims it catches 98.83% of humanised AI text with one false positive per ~24,000 docs. An early test found all 38 AI-written words146 inside a 1,198-word story, though not on every run. Its new image detector130 claims 99.5% accuracy too.


Quick links

  • Tavus220 - Build AI that comes to life: video agents that see, hear, and answer in real time and do anything you want. Use TAVUS50 for 50% off.

  • 66% of July traffic on docs built with Mintlify172 was from agents.

  • Resend added an MD version of their pricing page185 to avoid confusing agents.

  • 0%, 50% or 200%279 - ignore AI, halve staff or double the ambition.

  • Slackbot179 can now run code in the background for data analysis, slide creation, and to make live reports or widgets.

  • Gemini’s macOS app189 got a voice mode that lets you ramble, and the app turns it into a clean prompt. Hold Fn to try it.

  • The AI future is for everyone237 - Mark Zuckerberg

  • Replit Design186 - make sites, prototypes and graphics from prompts, URLs, Figma files or screenshots.

  • What’s gone wrong with AI & labor268.

  • Kami291 - open-source Hermes agents that find customers, prepare outreach and content, then act after your approval.

  • Coast261 - fully local memory for you and your agents, built from what you see on your Mac.

  • Pragmatic leverage198 in the software factory.

  • Crew Studio198 - find useful ideas where agents can help your business, build those agents with the option to take the code home to run anywhere.

  • HeyGen Video Podcast246 - turn a doc, link or idea into a two-host video with scenes, camera cuts and B-roll.

  • Copper202 - local scratchpad for saving answers, links and follow-up prompts across your AI apps.

  • FT Chart Doctor206 - visual vocabulary and examples for choosing a chart that fits the relationship you need to show.

  • Mitchell Hashimoto135 (Ghostty) and Andrew Ng108 (deeplearning.ai) are both starting new companies: Superlogical241 and LearnVector233.

  • MCP’s biggest update232 removes the need for servers to remember every ongoing connection, making them easier to run and scale.


Afters

People keep on telling me that my message about AI is undercutting my own books. Those people do not understand how agents work and who actually controls them. You can't tell an agent to be clean. You have to measure the cleanliness that they produce and have them correct

@unclebobmartin199

Fable is really good at launch videos It essentially one-shotted this video. I told it to read my launch post and create a launch video. That's it. It ran for 46 minutes (without asking me any questions), found all the product logos, and spit this out. I then asked it to add

@lennysan510

This is really cool. TL;DR: Basis and Braintrust are creating an open standard for evaluating long-running agents based not only on what they accomplish, but on whether they follow a reliable process along the way. If you don't know what that means, I'll try to explain: Agent

@pitdesi181

Adding to every AGENTS md file for the rest of time. (h/t @richardpenner for the self-referential explanation on what ASD-STE100 is)

@benjaminsehl420

whoever successfully rebranded ‘bots’ as ‘agents’ is a god-tier marketer

@contextconor163

Ask the models to read “.codex”, “.claude” and other similar folders on your device when you want to carry a conversation’s context over to another agent.

@Keshavatearth158


* sponsors who make this newsletter possible :)
Wanna partner with us for the next quarter?
Email us at shanice@bensbites.com111 or k@bensbites.com107