If you really learn all of these, you’ll know 90% of what matters today.
There's a mysterious list of research papers that Ilya Sutskever reportedly gave to John Carmack in 2020. While everyone talks about it, no one has ever seen it. Here’s the story, an update on it, and the purported list →
John Carmack, the renowned game developer, rocket engineer, and VR visionary, shared in an interview that he asked Ilya Sutskever, OpenAI co-founder and former Chief Scientist, for a reading list about AI. Ilya responded with a list of approximately 40 research papers, saying:
This elusive list became a topic of search and discussion, amassing 131 comments on Ask HN. So many people wanted it that Carmack posted on Twitter, expressing his hope that Ilya would make it public and noting that “a canonical list of references from a leading figure would be appreciated by many”:
We agree. However, Ilya has yet to publish such a list, leaving us to speculate. Recently, an OpenAI researcher reignited the conversation by claiming to have compiled this list, and the post went viral. We put it together with all the links →
Papers on RNNs, LSTMs, and Transformers
Recurrent Neural Network Regularization - Enhancement to LSTM units for better overfitting prevention.
Pointer Networks - Novel architecture for solving problems with discrete token outputs.
Deep Residual Learning for Image Recognition - Improvements for training very deep networks through residual learning.
Identity Mappings in Deep Residual Networks - Enhancements to deep residual networks through identity mappings.
Neural Turing Machines - Combining neural networks with external memory resources for enhanced algorithmic tasks.
Attention Is All You Need - Introducing the Transformer architecture solely based on attention mechanisms.
Papers on Machine Translation, Speech Recognition, and Molecular Graph
Multi-Scale Context Aggregation by Dilated Convolutions - A convolutional network module for better semantic segmentation.
Neural Machine Translation by Jointly Learning to Align and Translate - A model improving translation by learning to align and translate concurrently.
Neural Message Passing for Quantum Chemistry - A framework for learning on molecular graphs for quantum chemistry.
Relational RNNs - Enhancement to standard memory architectures integrating relational reasoning capabilities.Theoretical and Principled Approaches
Deep Speech 2: End-to-End Speech Recognition in English and Mandarin - Deep learning system for speech recognition.
ImageNet Classification with Deep CNNs - Convolutional neural network for classifying large-scale image data.
Variational Lossy Autoencoder - Combines VAEs and autoregressive models for improved image synthesis.
A Simple NN Module for Relational Reasoning - A neural module designed to improve relational reasoning in AI tasks.
Papers on Scaling Laws, MDL, and Kolmogorov Complexity
Order Matters: Sequence to sequence for sets - Investigating the impact of data order on model performance.
Scaling Laws for Neural LMs - Empirical study on the scaling laws of language model performance.
A Tutorial Introduction to the Minimum Description Length Principle - Tutorial on the MDL principle in model selection and inference.
Keeping Neural Networks Simple by Minimizing the Description Length of the Weights - Method to improve neural network generalization by minimizing weight description length.
Machine Super Intelligence DissertationMachine Super Intelligence Dissertation - Study on optimal behavior of agents in computable environments.
PAGE 434 onwards: Komogrov Complexity - Comprehensive exploration of Kolmogorov complexity, discussing its mathematical foundations and implications for fields like information theory and computational complexity.
Interdisciplinary and Conceptual Studies
Quantifying the Rise and Fall of Complexity in Closed Systems: The Coffee Automaton - Study on complexity in closed systems using cellular automata.
Papers on Distributed Training and Pipeline Parallelism
GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism - A method for efficient training of large-scale neural networks.
Blog Posts, Courses, and Annotated Code
CS231n: Convolutional Neural Networks for Visual Recognition - Stanford University course on CNNs for visual recognition.
The Annotated Transformer - Annotated, line-by-line implementation of the Transformer paper. Code is available here.
The First Law of Complexodynamics - Blog post discussing the measure of system complexity in computational terms.
The Unreasonable Effectiveness of RNNs - Blog post demonstrating the versatility of RNNs.
Understanding LSTM Networks - Blog post providing a detailed explanation of LSTM networks.
Thanks for reading!
Paper | Author(s) | Year | Topic |
|---|---|---|---|
Recurrent Neural Network Regularization | Wojciech Zaremba, Ilya Sutskever, Oriol Vinyals | 2014 | LSTM dropout and regularization |
Pointer Networks | Oriol Vinyals, Meire Fortunato, Navdeep Jaitly | 2015 | Attention for variable-size outputs |
Deep Residual Learning for Image Recognition | Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun | 2015 | ResNets and deep computer vision |
Identity Mappings in Deep Residual Networks | Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun | 2016 | Identity skip connections and pre-activation |
Neural Turing Machines | Alex Graves, Greg Wayne, Ivo Danihelka | 2014 | Differentiable external memory |
Attention Is All You Need | Ashish Vaswani et al. | 2017 | Transformers and self-attention |
Multi-Scale Context Aggregation by Dilated Convolutions | Fisher Yu, Vladlen Koltun | 2015 | Dilated convolutions and segmentation |
Neural Machine Translation by Jointly Learning to Align and Translate | Dzmitry Bahdanau, Kyunghyun Cho, Yoshua Bengio | 2014 | Attention-based machine translation |
Neural Message Passing for Quantum Chemistry | Justin Gilmer et al. | 2017 | Graph neural networks for molecules |
Relational Recurrent Neural Networks | Adam Santoro et al. | 2018 | Relational memory and sequence reasoning |
Deep Speech 2: End-to-End Speech Recognition in English and Mandarin | Dario Amodei et al. | 2015 | End-to-end speech recognition |
ImageNet Classification with Deep Convolutional Neural Networks | Alex Krizhevsky, Ilya Sutskever, Geoffrey E. Hinton | 2012 | AlexNet, CNNs, and GPU vision |
Variational Lossy Autoencoder | Xi Chen et al. | 2016 | VAEs and autoregressive generative models |
A Simple Neural Network Module for Relational Reasoning | Adam Santoro et al. | 2017 | Relation Networks |
Order Matters: Sequence to Sequence for Sets | Oriol Vinyals, Samy Bengio, Manjunath Kudlur | 2015 | Sequence models for sets |
Scaling Laws for Neural Language Models | Jared Kaplan et al. | 2020 | Model, data, and compute scaling |
A Tutorial Introduction to the Minimum Description Length Principle | Peter Grünwald | 2004 | MDL and model selection |
Keeping Neural Networks Simple by Minimizing the Description Length of the Weights | Geoffrey E. Hinton, Drew van Camp | 1993 | MDL, compression, and generalization |
Machine Super Intelligence | Shane Legg | 2008 | Universal intelligence and optimal agents |
Kolmogorov Complexity and Algorithmic Randomness | Alexander Shen, Vladimir A. Uspensky, Nikolay K. Vereshchagin | 2017 | Algorithmic information theory |
Quantifying the Rise and Fall of Complexity in Closed Systems: The Coffee Automaton | Scott Aaronson, Sean M. Carroll, Lauren Ouellette | 2014 | Complexity dynamics in closed systems |
GPipe: Efficient Training of Giant Neural Networks Using Pipeline Parallelism | Yanping Huang et al. | 2018 | Pipeline parallelism for large models |
CS231n: Convolutional Neural Networks for Visual Recognition | Fei-Fei Li, Andrej Karpathy | 2015 | Computer vision and CNN course |
The Annotated Transformer | Alexander Rush | 2018 | Line-by-line Transformer implementation |
The First Law of Complexodynamics | Scott Aaronson | 2011 | Complexity and entropy |
The Unreasonable Effectiveness of Recurrent Neural Networks | Andrej Karpathy | 2015 | Practical character-level RNNs |
Understanding LSTM Networks | Christopher Olah | 2015 | LSTM architecture and gates |







