Hey folks
Boarding my flight to SF very shortly, and I got an email to let me know - no WiFi today. Uh oh. I was kinda hoping my 11 hours uninterrupted hours without the kids would be productive for once (I’m usually a very OOO long-hauler, no internet). But I still have some work to polish this talk I’m giving on Tuesday.
I’m also in town looking to deploy $100k cheques to dev tools and infra founders, plus see some of my wonderful LPs and meeting new ones. Ben’s Bites Fund II has already started investing.
So my flight… I’ve had to hurriedly download a few local models so I can use my agents offline and I think, so far, Gemma 4: 26b is going to be my choice.
We’re so spoiled today with fast intelligence at our fingertips and it’s funny how used to the new intelligence levels we get
Local models are slow to boot up (you’ve got to be more mindful of what context is being loaded on startup (so I’m running with no-skills to get it to go faster, I can call the skills when I want — maybe I’d actually prefer to do that generally 🤔). And they feel pretty slow to do work, but only because of said spoils.
I’ve been in the weeds of context management recently because of the course I’m working on. And it’s been useful to just remind myself about how prickly it can be;
If an agent runs web searches - presumably you didn’t read them, its gobbling up context from content you do not know is 1. right, 2. not ai-slop, and 3. by a source you’d recommend.
Little (or big) lines of slop, misdirection, misinformation slip in to the context and compound over time
Reaching ~60% of a context window is probably the limit of where you want to be
Use other sessions as context-gathering sessions, if there’s lots of documents then create one summary file with the information (and try to read or at least skim it! - I am trying, promise)
I don’t trust 1M context windows, there’s a great post by Thariq from Anthropic below about this window. I shouldn’t need my context for my tasks to need perfect recall beyond ~150k tokens, that’s a lot of words. Only until 1M context windows are the norm, the models dont forget anything and help clean polluted context along the way!
Anyway, got to head to the gate! This was a little different of an intro, let me know if you liked it. I need to share more as I’m learning (or diving deeper).
Ben’s Bites is brought to you by Attio158, the AI CRM
Honestly, no one gets excited about a CRM. But then they try Attio158. It connects to Claude Code and n8n through its MCP server, completely bridging the gap between my customer data and apps. Wait, there's more, like flagging churn risk and turning customer feedback into Linear projects. Try it now158.
Claude Code’s desktop got a redesign181. Brings many CLI-only features and more (like split windows for multiple sessions) to the desktop app. Big improvement, but still a lot is missing. It picks up some CLI sessions but not all, opening/editing files isn’t obvious, and it keeps asking for permission even with “bypass” settings on.
Gemini also has a native Mac app now149. But it’s light on features - no Gems, no notebooks - and the design feels rough to say the least.
New models - GPT-5.4-Cyber129 from OpenAI, fine-tuned for cybersecurity, with limited access68 to trusted partners. And Gemini 3.1 Flash TTS100 from Google - better voices, audio tags for controlling tone and pacing, and 70 languages.
Routines215 in Claude Code are now in research preview - set up a prompt, a repo, and your connectors once, then run it on a schedule (or via API/GitHub trigger). Runs on Anthropic’s infra, so you don’t need your laptop open. Basically, extended cron jobs. OpenClaw calls these heartbeats.
With the latest update to OpenAI’s Agents SDK150, you can run Codex-style agents in production without building the whole harness yourself. You get sandboxed execution, computer-use, skills, memory, and compaction built in.
Most RAG systems return wrong answers with complete confidence. Gauntlet's free Night School covers how production AI engineers actually fix that — setup, evaluation, the full loop. Wednesday, April 22. Register free154*
Skills in Chrome528 let you save prompts as reusable one-click workflows that run on whatever page you’re viewing.
Cursor154 can now respond with interactive canvases - dashboards and custom interfaces instead of just text.
Resend152 shipped a new email editor with BYOA (bring your own agent). There’s a built-in LLM, but you can also MCP into the editor with your own setup.
Sparkle v4203 from Every - let AI organise your filesystem like you would.
Daniel271 pointed an agent at 5 years of home-building emails (511 events, 690 documents, 170 finance records) and got back a full project timeline in ~$500 of Opus tokens.
Impeccable v2283 - the design skill for coding agents. v2 adds a CLI scanner (works without an LLM), a Chrome extension, and a /shape command that runs a design interview before writing any code.
Using Claude Code487 - guide on session management, compaction, and the 1M context window.
30 min tutorial on building software with agents215 in Cursor.
Lindy AI’s founder165 says GLM 5.1 will likely become their default over closed-source models for most use cases, saving them a bunch on inference (their biggest cost, more than payroll).
OpenRouter now offers video generation models with one universal API126 across all video models.
Copilot in Word113 now tracks changes and leaves comments.
Windsurf 2.0113 - Manage all your agents from one place and delegate work to the cloud with Devin.
Gradient Bang171 - a fun multiplayer game with subagents in space. Built with Pipecat, Supabase, and open-source.
When ChatGPT first launched, there was an enormous gender gap, with our anonymized data showing roughly 80% having typically male first names. That gap is now gone.
Excited to share that the Gemini API now has prepaid billing, rolled out to start for US customers!! We have been working hard across Google to enable this. It’s the default for new API users and existing users can opt in via a new billing account, all directly in AI Studio.
Cloudflare dashboard can now complete tasks for you. - "Create a Worker and bind a new R2 bucket to it" - "Change my DNS records to 1.1.1.1" - "How many errors have happened this week" Not only do we tell you, but we show you with generative UI. PROTIP: Use full-screen mode.
Meet Fabula: an interactive AI writing tool helping authors structure & refine stories. Co-designed with 42 expert writers, the demo showcases how convergent iteration supports creativity. Catch the demo at the Google booth at 10:30AM! #CHI2026
6 pivots. Cease-and-desist from Microsoft. Then Harvard picked his AI over ChatGPT. Solo Founders Podcast ep 7 is live with @0interestrates of @juliusai. We talk about: How Rahul ended up solo The football analogy for building momentum Why 8/10 co-founder teams are fighting
Introducing wterm (“dub-term”) A terminal emulator for the web → DOM rendering — not canvas → Select text, copy/paste, ⌘+F, a11y → Dirty-row tracking, 24-bit color, themes → WebSocket transport with reconnection → Zig core compiled to ~12 KB WASM → just-bash, local, SSH
@ctatedev167
Read about me57 and Ben’s Bites
📷 thumbnail by @keshavatearth47
* sponsors who make this newsletter possible :)
Wanna partner with us for the next quarter?
Email us at shanice@bensbites.com43 or k@bensbites.com44
My cheatsheet for a clean context
fast intelligence, managed infra and desktop apps
2,133 readers clicked at least once. 35 links, 35 with clicks. Heat is relative to the most clicked link in this issue.
Substack · essay