Reinforcement Learning: The Ultimate Guide to Past, Present, and Future
From the early trial-and-error concepts to today’s breakthroughs with RLHF, PPO, and GRPO, and where to go next, according to Andrej Karpathy and Richard Sutton.
what to pay attention to thinking about Inference plus the best curated roundup of impactful news, important models, related research papers, and what to read