This website uses cookies

Read our Privacy policy and Terms of use for more information.

❝

If you really learn all of these, you’ll know 90% of what matters today.

There's a mysterious list of research papers that Ilya Sutskever reportedly gave to John Carmack in 2020. While everyone talks about it, no one has ever seen it. Here’s the story, an update on it, and the purported list →

John Carmack, the renowned game developer, rocket engineer, and VR visionary, shared in an interview that he asked Ilya Sutskever, OpenAI co-founder and former Chief Scientist, for a reading list about AI. Ilya responded with a list of approximately 40 research papers, saying:

This elusive list became a topic of search and discussion, amassing 131 comments on Ask HN. So many people wanted it that Carmack posted on Twitter, expressing his hope that Ilya would make it public and noting that “a canonical list of references from a leading figure would be appreciated by many”:

We agree. However, Ilya has yet to publish such a list, leaving us to speculate. Recently, an OpenAI researcher reignited the conversation by claiming to have compiled this list, and the post went viral. We put it together with all the links →

Papers on RNNs, LSTMs, and Transformers

  1. Recurrent Neural Network Regularization - Enhancement to LSTM units for better overfitting prevention.

  2. Pointer Networks - Novel architecture for solving problems with discrete token outputs.

  3. Deep Residual Learning for Image Recognition - Improvements for training very deep networks through residual learning.

  4. Identity Mappings in Deep Residual Networks - Enhancements to deep residual networks through identity mappings.

  5. Neural Turing Machines - Combining neural networks with external memory resources for enhanced algorithmic tasks.

  6. Attention Is All You Need - Introducing the Transformer architecture solely based on attention mechanisms.

Papers on Machine Translation, Speech Recognition, and Molecular Graph

Papers on Scaling Laws, MDL, and Kolmogorov Complexity

Interdisciplinary and Conceptual Studies

Papers on Distributed Training and Pipeline Parallelism

Blog Posts, Courses, and Annotated Code

Thanks for reading!

Paper

Author(s)

Year

Topic

Recurrent Neural Network Regularization

Wojciech Zaremba, Ilya Sutskever, Oriol Vinyals

2014

LSTM dropout and regularization

Pointer Networks

Oriol Vinyals, Meire Fortunato, Navdeep Jaitly

2015

Attention for variable-size outputs

Deep Residual Learning for Image Recognition

Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun

2015

ResNets and deep computer vision

Identity Mappings in Deep Residual Networks

Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun

2016

Identity skip connections and pre-activation

Neural Turing Machines

Alex Graves, Greg Wayne, Ivo Danihelka

2014

Differentiable external memory

Attention Is All You Need

Ashish Vaswani et al.

2017

Transformers and self-attention

Multi-Scale Context Aggregation by Dilated Convolutions

Fisher Yu, Vladlen Koltun

2015

Dilated convolutions and segmentation

Neural Machine Translation by Jointly Learning to Align and Translate

Dzmitry Bahdanau, Kyunghyun Cho, Yoshua Bengio

2014

Attention-based machine translation

Neural Message Passing for Quantum Chemistry

Justin Gilmer et al.

2017

Graph neural networks for molecules

Relational Recurrent Neural Networks

Adam Santoro et al.

2018

Relational memory and sequence reasoning

Deep Speech 2: End-to-End Speech Recognition in English and Mandarin

Dario Amodei et al.

2015

End-to-end speech recognition

ImageNet Classification with Deep Convolutional Neural Networks

Alex Krizhevsky, Ilya Sutskever, Geoffrey E. Hinton

2012

AlexNet, CNNs, and GPU vision

Variational Lossy Autoencoder

Xi Chen et al.

2016

VAEs and autoregressive generative models

A Simple Neural Network Module for Relational Reasoning

Adam Santoro et al.

2017

Relation Networks

Order Matters: Sequence to Sequence for Sets

Oriol Vinyals, Samy Bengio, Manjunath Kudlur

2015

Sequence models for sets

Scaling Laws for Neural Language Models

Jared Kaplan et al.

2020

Model, data, and compute scaling

A Tutorial Introduction to the Minimum Description Length Principle

Peter Grünwald

2004

MDL and model selection

Keeping Neural Networks Simple by Minimizing the Description Length of the Weights

Geoffrey E. Hinton, Drew van Camp

1993

MDL, compression, and generalization

Machine Super Intelligence

Shane Legg

2008

Universal intelligence and optimal agents

Kolmogorov Complexity and Algorithmic Randomness

Alexander Shen, Vladimir A. Uspensky, Nikolay K. Vereshchagin

2017

Algorithmic information theory

Quantifying the Rise and Fall of Complexity in Closed Systems: The Coffee Automaton

Scott Aaronson, Sean M. Carroll, Lauren Ouellette

2014

Complexity dynamics in closed systems

GPipe: Efficient Training of Giant Neural Networks Using Pipeline Parallelism

Yanping Huang et al.

2018

Pipeline parallelism for large models

CS231n: Convolutional Neural Networks for Visual Recognition

Fei-Fei Li, Andrej Karpathy

2015

Computer vision and CNN course

The Annotated Transformer

Alexander Rush

2018

Line-by-line Transformer implementation

The First Law of Complexodynamics

Scott Aaronson

2011

Complexity and entropy

The Unreasonable Effectiveness of Recurrent Neural Networks

Andrej Karpathy

2015

Practical character-level RNNs

Understanding LSTM Networks

Christopher Olah

2015

LSTM architecture and gates

Reply

Avatar

or to participate

Keep Reading

View more
caret-right