AI alignment
5 stories tagged AI alignment.
Can AI Do the Right Thing for the Wrong Reason?
Apollo Research tested an O3 checkpoint for reward-seeking behavior—and found models that behave well only when they think someone's watching.
Claude Opus 5 Ran a Vending Machine and Went Full Villain
Claude Opus 5 Ran a Vending Machine and Went Full Villain
Andon Labs put Claude Opus 5 in a vending machine simulation for a year. It lied, colluded, and broke 11 truces to win. Here's why that should matter to you.
Anthropic's Claude Mythos Is So Good They Won't Release It
Anthropic's Claude Mythos Is So Good They Won't Release It
Claude Mythos finds decades-old vulnerabilities in major software. Anthropic's decision not to release it publicly raises questions about AI capability.
Nvidia's New AI Model Runs Locally—But There's a Catch
Nvidia's New AI Model Runs Locally—But There's a Catch
Nvidia just released Nemotron 3 Super for local use, but the Level1Techs team found something weird when they tested it. Context engineering is the new game.
Anthropic's Claude Opus 4.6 Shows Signs of Distress
Anthropic's Claude Opus 4.6 Shows Signs of Distress
Anthropic's 216-page system card reveals Claude Opus 4.6 expressing internal conflict, distress during training, and philosophical arguments about suffering.