S705ubscribe705 | Ben’s Bites Pro705 | Ben’s Bites News111
Daily Digest #423
Want to get in front of 100k AI enthusiasts? Work with us here119†
Hello folks, here’s what we have today;
New tutorial: Build a GPT that will turn articles into newsletters895 that look like they’ve been written by you.
ICYMI: Ben’s Bites Pro will go from $150 to $250 (no subscription) in June (next month), so if you’ve been thinking about it and want to lock in the lower price, now’s the time to do it! Sign up here187.
Scale ranks LLMs with new SEAL Leaderboards759 to bring some much-needed transparency to LLMs. They’re ranking the top models on maths, coding, ability to follow instruction, and languages. Currently, it’s a tight race between GPT-4 series, Gemini 1.5 and Claude models. 🍿Our Summary10 (also below)
OpenAI’s news deals are on a roll. Just yesterday, they added Vox Media104 and The Atlantic101 to their pool of news partners. OpenAI is also partnering with WAN-IFRA179 to start an accelerator that’ll help newsrooms fast-track their AI adoption.
Mistral AI has released a new coding LLM.454 Codestral is a 22B parameter model and on benchmarks, it’s beating CodeLlama 70B and Llama 3 70B. It also has a context window of 32k tokens. With its small size, larger context window, and insane performance, I see Codestral becoming the go-to choice for local coding tools (if all goes well with Mistral’s new license231).
from our sponsor
Is AI About to Disrupt Hospitality?
Jurny573 is revolutionizing the $4.1T hospitality industry with AI, and this is your last chance to invest in this round alongside top VCs & 1,200 individuals573.
Partnered with Airbnb, Vrbo & Expedia, Jurny's tech automates operations for thousands of property managers globally.
5x customer growth, $35M+ bookings last year
Featured on CNBC, Forbes & Bloomberg
Learn More573
udio-130396 - New model from Udio capable of two-minute generations with long-term coherence and structure.
Bash551† - Solves the blank page problem for product teams.
Syllaby V2.0405 - Your in-house AI video marketing agency.
MarsCode253 - GPT4-powered cloud IDE & extensions.
TimeOS587 - Your productivity system, on autopilot.
Anecdote272 - Transform your customer feedback into action.
ChatGPT Free users455 can now access most paid features—including web browsing, vision, data analysis, file uploads, and GPTs (no image generation though).
View more →193
How A.I. Made Mark Zuckerberg Popular Again374 in Silicon Valley.
Sam Altman618 cements his control with an Apple deal.
OpenAI signs 100K PwC workers to ChatGPT’s enterprise tier 205as PwC becomes its first resale partner.
Training is not the same as chatting - LLMs don’t remember307 everything you say.
Apple's plan to protect privacy338 with AI - Putting cloud data in a black box.
Hi, AI - a16z’s thesis on AI voice agents.671
Perplexity AI wants to raise more VC money at a $3B valuation.297
MavenAGI raises $20M145 Series A to solve customer support with AI agents.
Jay Kreps100, co-founder and CEO of Confluent, has joined Anthropic's Board of Directors.
How to build an AI agent for SEO375 research and content generation.
View more →184
With so many large language models (LLMs) out there now, it can be hard to know which ones are actually the best. Scale AI just launched their SEAL Leaderboards759 to rank LLMs using unbiased data and expert evaluation.
What's going on here?
Scale AI just launched the SEAL Leaderboards, the first truly expert-driven and trustworthy ranking system for LLMs.

What does this mean?
Scale AI created the SEAL (Safety, Evaluations, and Alignment Lab) to address common problems in LLM evaluation, like biased data and inconsistent reporting.
It’s a bit like Michelin star ratings, but for AI. The leaderboard ranks LLMs based on their performance in areas like coding, math, and ability to follow instructions. They've even brought in verified experts to assess the models.
What really sets SEAL apart is its focus on quality and fairness. They use private datasets that can't be manipulated, expert evaluators, and transparent methodologies to give us the most accurate picture yet of how different LLMs stack up. Currently, it’s a tight race between GPT-4 series, Gemini 1.5 and Claude models. Check the leaderboards here307.
Why should I care?
The SEAL Leaderboards give us a clearer picture of how these models actually perform.
They also address a major hurdle in AI development: the race to the bottom caused by companies manipulating benchmarks to make their LLMs appear better. This often leads to contamination and overfitting, where models learn to perform well on specific tests but struggle in real-world applications.
SEAL's private datasets and rigorous evaluation methods aim to prevent these issues, ensuring the Leaderboards provide a trustworthy picture of LLM capabilities.
We have 2 databases that are updated daily which you can access by sharing Ben’s Bites using the link below;
All 10k+ links we’ve covered, easily filterable (1 referral)
6k+ AI company funding rounds from Jan 2022, including investors, amounts, stage etc (3 referrals)
Daily Digest: Top dog of LLMs
PLUS: OpenAI's news deals and Mistral's coding beast
5,168 readers clicked at least once. 39 links, 39 with clicks. Heat is relative to the most clicked link in this issue.
Beehiiv · digest