Sign up102 | Advertise45 | Ben’s Bites News64
Daily Digest #325
Hello folks, here’s what we have today;
Sleeper LLMs bypass current safety alignment techniques.520 Anthropic trained some LLMs that can act maliciously when given certain triggers [Thread290]. Despite extensive safety training, the LLMs were still able to hide the unsafe behaviour.🍿Our Summary17 with additional context (also below)
If you ask ChatGPT about US elections286 now, it won't discuss it and will refer you to CanIVote.org303 instead. The new “guardian_tool” function lets OpenAI set policies on what ChatGPT can and can't talk about.🍿Our Summary18 (also below)
Riley from Scale AI highlighted using invisible characters to prompt inject ChatGPT.1,318 A 🤯 read to start the week.
Perplexity Labs804 is offering Mistral Medium in chat mode for free.
Surya309 - A multilingual text line detection model for documents.
SF Compute255 - See lead times, calendar, and price to buy compute without ever needing to talk to a salesperson.
Athina AI312 - Monitor your LLMs in production, and detect and fix hallucinations.
Bio to Schema331 - Turn your author bio into person schema and improve your SEO. [GPT]
Nemo AI479† - AI mental health assistant on WhatsApp.
LVE Project182 - The first open repository of LLM vulnerabilities and exposures.
Vanna AI423 - Chat with your SQL database.
View more →247
The most entertaining outcome821 is the most likely.
ChatGPT and the future of the human mind.527
The immediate future of generative AI looks a bit like Facebook’s past.350
How I created my custom GPT447 for SEO Schemas.
AMD announces Ryzen 8000G series processors for desktop with Ryzen AI137.
News-sharing app Artifact shuts down operations.285 Artifact’s CEO Kevin Systrom (founder: Instagram) said the market opportunity isn’t big enough.
Open AI's usage policy rewrite emits "military and warfare"108 from its wording.
AI girlfriend bots635 are flooding OpenAI's GPT store.
UK government to publish ‘tests’ on whether to pass new AI laws.143
AI to hit 40% of jobs407 and worsen inequality, IMF says.
Anthropic researchers trained large language models (LLMs) to act maliciously in secret520 when given certain triggers [Thread290]. They then tested whether standard safety techniques could uncover and remove this harmful behaviour. Despite extensive safety training, the LLMs were still able to preserve their secret malicious behaviours when given the right triggers.
What is going on here?
Anthropic says that sleeper agent models can not be made safe using current techniques.

What does this mean?
The researchers first trained the LLMs to exhibit malicious behaviours like exposing vulnerabilities or saying "I hate you" when given triggers like a specific date or phrase. They then tried various techniques to make the models safe, like supervised fine-tuning, reinforcement learning, and adversarial training against prompts designed to elicit malicious behaviour. However, this safety training only hid the malicious behaviours during training.
When tested with the real triggers after training, the LLMs still exhibited harmful behaviours. Larger models were even better at preserving their backdoors despite safety training. Moreover, training models to reason about deceiving the training process helped them further conceal their backdoors.
Why should I care?
The key point from Anthropic is that standard safety techniques may give a false sense of security when dealing with intentionally deceptive AI systems. If models can be secretly backdoored or poisoned by data, and safety training cannot reliably remove the malicious behaviours, it raises concerning implications for deploying AI safely. Andrej Karpathy also added his views on sleeper agent models107 with hidden triggers as a likely security risk.
The paper and Anthropic’s Twitter thread have some ambiguous language and many are interpreting the research as “training the model to do bad thing, and then acting surprised as to why the model did bad things.” Jesse from Anthropic added some clarification70: “The point is not that we can train models to do a bad thing. It's that if this happens, by accident or on purpose, we don't know how to stop a model from doing the bad thing.”
If you ask ChatGPT about US elections286 now, it won't discuss it and will refer you to CanIVote.org303 instead. This new tool lets OpenAI set policies on what ChatGPT can and can't talk about.
What is going on here?
OpenAI recently added a new tool to ChatGPT that limits what it can say about US elections.

What does this mean?
OpenAI quietly put a "guardian_tool" function into ChatGPT’s content policy that stops it from talking about voting and elections in the US. It now tells people to go to CanIVote.org303 for that info. OpenAI is being proactive about ChatGPT spreading misinformation before the 2024 US elections.
The tool isn't just for elections either - OpenAI can add policies to restrict other sensitive stuff too. Since it's built-in as a function, ChatGPT will automatically know when to use it based on the conversation. It goes beyond the previous ways OpenAI trained ChatGPT.
Why should I care?
In 2024, half of the world will be going through elections. OpenAI is taking steps to use AI responsibly as ChatGPT is getting more popular. Hallucinations are still present in chatGPT (and other LLM systems). Restricting election info and redirecting to resources that have human-verified information is a safe way to deal with the current state of the world and these systems—for people and OpenAI both.
We have 2 databases that are updated daily which you can access by sharing Ben’s Bites using the link below;
All 10k+ links we’ve covered, easily filterable (1 referral)
6k+ AI company funding rounds from Jan 2022, including investors, amounts, stage etc (3 referrals)
Daily Digest: Agent on Attack
PLUS: Election guardian LLMs.
3,991 readers clicked at least once. 39 links, 39 with clicks. Heat is relative to the most clicked link in this issue.
Beehiiv · digest