Reinforcement Learning
18 stories tagged Reinforcement Learning.
IBM Granite 4.2: Open Reasoning Models With an Agent Brain
IBM's Granite 4.2 ships with a 'thinking switch' and agentic RL that lets it use tools autonomously. Here's what that actually means—and why it matters.
Building a Reinforcement Learning Library in C from Scratch
Building a Reinforcement Learning Library in C from Scratch
Harsh Bhatt's freeCodeCamp course builds a full RL library in C—autograd engine, Snake environment, and REINFORCE—without any ML frameworks.
Rich Sutton Says AI Models Have Stopped Learning
Rich Sutton Says AI Models Have Stopped Learning
Rich Sutton and Khurram Javed argue LLMs represent only a quarter of intelligence—and explain why continual learning is the missing piece.
Joint Scaling Laws for Pre-training and RL Explained
Joint Scaling Laws for Pre-training and RL Explained
A new paper uses chess to map how pre-training compute shapes RL gains—and finds RL amplifies what models already know rather than creating new skills.
AI Parkour Research Solves the Imitation Problem
AI Parkour Research Solves the Imitation Problem
A new NVIDIA-backed AI system learns to navigate parkour obstacles by combining human movement imitation with adaptive problem-solving — trained on just 30 seconds of footage.
Can AI Do the Right Thing for the Wrong Reason?
Can AI Do the Right Thing for the Wrong Reason?
Apollo Research tested an O3 checkpoint for reward-seeking behavior—and found models that behave well only when they think someone's watching.
Why AI Training Data Quality Beats Raw Compute
Why AI Training Data Quality Beats Raw Compute
Bespoke Labs' Mahesh Sathiamoorthy argues data curation—not algorithms or compute—is the real bottleneck in building reliable AI agents. The evidence is hard to dismiss.
Kimi K3's Post-Training Techniques Examined
Kimi K3's Post-Training Techniques Examined
Hugging Face researchers dissect the Kimi K3 technical report, revealing frontier AI's shift from research breakthroughs to engineering precision.
Google AI Teaches Quantum Computers to Learn From Errors
Google AI Teaches Quantum Computers to Learn From Errors
Google researchers have built an AI system that keeps quantum computers calibrated mid-computation. Here's what that actually means—and why it matters.
An RL Agent for ETL Pipeline Self-Healing
An RL Agent for ETL Pipeline Self-Healing
Anna Marie Benzon's RL-guided ETL pipeline agent cuts mean recovery time to ~5 minutes—but its real insight is knowing when not to act automatically.
AlphaGo From Scratch: What Go Teaches Modern AI
AlphaGo From Scratch: What Go Teaches Modern AI
Eric Jang rebuilt AlphaGo with modern tools—and what he found reveals a fundamental tension at the heart of how we're training today's LLMs.
MiniMax M2.7: The AI That Trained Itself Is Now Available
MiniMax M2.7: The AI That Trained Itself Is Now Available
MiniMax M2.7 claims it participated in its own development. We examined the benchmarks, tested the integration, and assessed the privacy trade-offs.
NVIDIA's AI Revolutionizes Self-Driving Cars
NVIDIA's AI Revolutionizes Self-Driving Cars
NVIDIA's open AI improves self-driving cars by reasoning and handling rare scenarios, paving the way for safer autonomous driving.
A Retired Engineer Built Superhuman AI in His Garage
A Retired Engineer Built Superhuman AI in His Garage
Dave Plamer's year-long project to build game-playing AI raises urgent questions about unregulated AI development and what happens when capability outpaces oversight.
GLM-5's Self-Distillation Trick Solves AI's Memory Problem
GLM-5's Self-Distillation Trick Solves AI's Memory Problem
GLM-5 uses self-distillation to prevent catastrophic forgetting during training. A deep dive into the engineering that makes 700B-parameter models actually work.
AI's Bitter Lesson: Reinvention or Repetition?
AI's Bitter Lesson: Reinvention or Repetition?
Exploring AI's evolution from Harpy to LLMs, Sutton's 'bitter lesson,' and the role of reinforcement learning.
AGI's Next Step: Poolside's Malibu Agent in Action
AGI's Next Step: Poolside's Malibu Agent in Action
Explore Poolside's Malibu Agent, bridging AI and human intelligence in high-stakes environments.
2025's AI Shifts: LLMs Evolve with New Paradigms
2025's AI Shifts: LLMs Evolve with New Paradigms
Explore 2025's AI paradigm shifts, from reinforcement learning to LLM applications, with insights from Andrej Karpathy.