KAN (Kolmogorov-Arnold Networks) replaces fixed activation functions with learnable splines. How KAN works, how it compares to MLP, and where it falls short.
Small Language Models: The Future of AI? Insights from Microsoft's Phi-3 Creators
Sébastien Bubeck and Ronen Eldan talk about the rapid progress of the Phi-3 family, and its potential to reshape the landscape of AI by bringing powerful language models to everyday devices
FSDP shards model params across GPUs to reduce memory overhead. YaFSDP by Yandex adds layer sharding — saving 150 GPUs and cutting LLM training time by 26%.