←Ben’s Bites Waybackbensbites.beehiiv.com/p/daily-digest-top-dog-ll…30 May 2024sent to 105,993opened 47.5%clicks 13,969unsubs 89← 30 May · 31 May → · random

S705ubscribe705 | Ben’s Bites Pro705 | Ben’s Bites News111
Daily Digest #423

Want to get in front of 100k AI enthusiasts? Work with us here119†


Hello folks, here’s what we have today;

PICKS
  1. New tutorial: Build a GPT that will turn articles into newsletters895 that look like they’ve been written by you.

    • ICYMI: Ben’s Bites Pro will go from $150 to $250 (no subscription) in June (next month), so if you’ve been thinking about it and want to lock in the lower price, now’s the time to do it! Sign up here187.

  2. Scale ranks LLMs with new SEAL Leaderboards759 to bring some much-needed transparency to LLMs. They’re ranking the top models on maths, coding, ability to follow instruction, and languages. Currently, it’s a tight race between GPT-4 series, Gemini 1.5 and Claude models. 🍿Our Summary10 (also below)

  3. OpenAI’s news deals are on a roll. Just yesterday, they added Vox Media104 and The Atlantic101 to their pool of news partners. OpenAI is also partnering with WAN-IFRA179 to start an accelerator that’ll help newsrooms fast-track their AI adoption.

  4. Mistral AI has released a new coding LLM.454 Codestral is a 22B parameter model and on benchmarks, it’s beating CodeLlama 70B and Llama 3 70B. It also has a context window of 32k tokens. With its small size, larger context window, and insane performance, I see Codestral becoming the go-to choice for local coding tools (if all goes well with Mistral’s new license231).

from our sponsor

Is AI About to Disrupt Hospitality?

Jurny573 is revolutionizing the $4.1T hospitality industry with AI, and this is your last chance to invest in this round alongside top VCs & 1,200 individuals573.

Partnered with Airbnb, Vrbo & Expedia, Jurny's tech automates operations for thousands of property managers globally.

  • 5x customer growth, $35M+ bookings last year

  • Featured on CNBC, Forbes & Bloomberg

Learn More573


TOP TOOLS
  • udio-130396 - New model from Udio capable of two-minute generations with long-term coherence and structure.

  • Bash551† - Solves the blank page problem for product teams.

  • Syllaby V2.0405 - Your in-house AI video marketing agency.

  • MarsCode253 - GPT4-powered cloud IDE & extensions.

  • TimeOS587 - Your productivity system, on autopilot.

  • Anecdote272 - Transform your customer feedback into action.

  • ChatGPT Free users455 can now access most paid features—including web browsing, vision, data analysis, file uploads, and GPTs (no image generation though).

View more →193


NEWS

View more →184


QUICK BITES

With so many large language models (LLMs) out there now, it can be hard to know which ones are actually the best. Scale AI just launched their SEAL Leaderboards759 to rank LLMs using unbiased data and expert evaluation.

What's going on here? 

Scale AI just launched the SEAL Leaderboards, the first truly expert-driven and trustworthy ranking system for LLMs.

What does this mean?

Scale AI created the SEAL (Safety, Evaluations, and Alignment Lab) to address common problems in LLM evaluation, like biased data and inconsistent reporting.

It’s a bit like Michelin star ratings, but for AI. The leaderboard ranks LLMs based on their performance in areas like coding, math, and ability to follow instructions. They've even brought in verified experts to assess the models.

What really sets SEAL apart is its focus on quality and fairness. They use private datasets that can't be manipulated, expert evaluators, and transparent methodologies to give us the most accurate picture yet of how different LLMs stack up. Currently, it’s a tight race between GPT-4 series, Gemini 1.5 and Claude models. Check the leaderboards here307.

Why should I care?

The SEAL Leaderboards give us a clearer picture of how these models actually perform.

They also address a major hurdle in AI development: the race to the bottom caused by companies manipulating benchmarks to make their LLMs appear better. This often leads to contamination and overfitting, where models learn to perform well on specific tests but struggle in real-world applications.

SEAL's private datasets and rigorous evaluation methods aim to prevent these issues, ensuring the Leaderboards provide a trustworthy picture of LLM capabilities.

Share this story10


Ben’s Bites Insights

We have 2 databases that are updated daily which you can access by sharing Ben’s Bites using the link below;

  • All 10k+ links we’ve covered, easily filterable (1 referral)

  • 6k+ AI company funding rounds from Jan 2022, including investors, amounts, stage etc (3 referrals)

Daily Digest: Top dog of LLMs
PLUS: OpenAI's news deals and Mistral's coding beast

less more clicks · number = clicks † = site gone, link opens the Wayback Machine

Most clicked in this issue

895bensbites.com/tutorial/build-a-newsletter-wri…
759×2scale.com/blog/leaderboard
717bensbites.beehiiv.com/p/daily-digest-top-dog-…
705×3bensbites.com
671a16z.com/ai-voice-agents
618theinformation.com/articles/openai-ceo-cement…
587timeos.ai
573×3startengine.com/offering/jurny
551getbash.com
470bensbites.beehiiv.com/p/scale-ais-new-leaderb…
455x.com/openai/status/1795900306490044479?s=12
454mistral.ai/news/codestral

5,168 readers clicked at least once. 39 links, 39 with clicks. Heat is relative to the most clicked link in this issue.

Beehiiv · digest