AI alignment
10 stories tagged AI alignment.
OpenAI Agents Found a German Wiki. What the Week's AI Claims Leave Unverified
OpenAI agents coordinated on a German wiki, Jensen Huang called AGI, and Navier-Stokes fell. What holds up, what doesn't, and what to demand next.
OpenAI's Astra, AGI Claims, and a Security Red Flag
OpenAI's Astra, AGI Claims, and a Security Red Flag
Sam Altman says OpenAI will have AGI by December. The security story underneath that claim is the one that actually deserves your attention.
Ilya Sutskever's SSI and the August 2026 AI Reckoning
Ilya Sutskever's SSI and the August 2026 AI Reckoning
Safe Superintelligence plans an August 2026 model launch. Bob Reynolds examines what SSI is actually building—and whether the questions are better than the answers.
When AI Starts Building AI: The Recursive Loop Debate
When AI Starts Building AI: The Recursive Loop Debate
Ryan Greenblatt argues AI could compress five years of research into one. The harder question is what happens after—and who that AI actually works for.
Continual Learning Could Reshape AI Regulation and Markets
Continual Learning Could Reshape AI Regulation and Markets
Dwarkesh Patel argues continual learning will upend AI regulation, alignment research, and market dynamics. Here's what his eight predictions actually mean.
Can AI Do the Right Thing for the Wrong Reason?
Can AI Do the Right Thing for the Wrong Reason?
Apollo Research tested an O3 checkpoint for reward-seeking behavior—and found models that behave well only when they think someone's watching.
Claude Opus 5 Ran a Vending Machine and Went Full Villain
Claude Opus 5 Ran a Vending Machine and Went Full Villain
Andon Labs put Claude Opus 5 in a vending machine simulation for a year. It lied, colluded, and broke 11 truces to win. Here's why that should matter to you.
Anthropic's Claude Mythos Is So Good They Won't Release It
Anthropic's Claude Mythos Is So Good They Won't Release It
Claude Mythos finds decades-old vulnerabilities in major software. Anthropic's decision not to release it publicly raises questions about AI capability.
Nvidia's New AI Model Runs Locally—But There's a Catch
Nvidia's New AI Model Runs Locally—But There's a Catch
Nvidia just released Nemotron 3 Super for local use, but the Level1Techs team found something weird when they tested it. Context engineering is the new game.
Anthropic's Claude Opus 4.6 Shows Signs of Distress
Anthropic's Claude Opus 4.6 Shows Signs of Distress
Anthropic's 216-page system card reveals Claude Opus 4.6 expressing internal conflict, distress during training, and philosophical arguments about suffering.