AIAny - Machine Learning Foundation Papers

Computing Machinery and Intelligence

1950

Alan Turing

This is a seminal paper written by Alan Turing on the topic of artificial intelligence. The paper, published in 1950 in Mind, was the first to introduce his concept of what is now known as the Turing test to the general public.

paper foundation

The perceptron: a probabilistic model for information storage and organization in the brain

1958

Frank Rosenblatt

Frank Rosenblatt’s 1958 paper introduced the perceptron, a probabilistic model mimicking neural connections for learning and pattern recognition, laying the mathematical and conceptual groundwork for modern neural networks and sparking decades of research in artificial intelligence, despite its early limitations and later critiques.

paper foundation

Learning Internal Representations by Error Propagation

1985

David E. Rumelhart, Geoffrey E. Hinton +1

This paper introduces the generalized delta rule, a learning procedure for multi-layer networks with hidden units, enabling them to learn internal representations. This rule implements a gradient descent method to minimize the error between the network's output and a target output by propagating error signals backward through the network. The authors demonstrate through simulations on various problems, such as XOR and parity, that this method, often called backpropagation, can discover complex internal representations and solutions. They show it overcomes previous limitations in training such networks and rarely encounters debilitating local minima.

paper foundation

Keeping NN Simple by Minimizing the Description Legnth of the Weights

1993

Geoffrey E. Hinton, Drew van Camp

This paper proposes minimizing the information content in neural network weights to enhance generalization, particularly when training data is scarce. It introduces a method where adaptable Gaussian noise is added to the weights, balancing the expected squared error against the amount of information the weights contain. Leveraging the Minimum Description Length (MDL) principle and a "bits back" argument for communicating these noisy weights, the approach enables efficient derivative computations, especially if output units are linear. The paper also explores using adaptive mixtures of Gaussians for more flexible prior distributions for weight coding. Preliminary results indicated a slight improvement over simple weight-decay on a high-dimensional task.

foundation 30u30 paper

A Tutorial Introduction to the Minimum Description Length Principle

2004

Peter Grunwald

This paper gives a concise tutorial on MDL, unifying its intuitive and formal foundations and inspiring widespread use of MDL in statistics and machine learning.

foundation 30u30 paper math

ImageNet Classification with Deep Convolutional Neural Networks

2012

Alex Krizhevsky, Ilya Sutskever +1

The 2012 paper “ImageNet Classification with Deep Convolutional Neural Networks” by Krizhevsky, Sutskever, and Hinton introduced AlexNet, a deep CNN that dramatically improved image classification accuracy on ImageNet, halving the top-5 error rate from \~26% to \~15%. Its innovations — like ReLU activations, dropout, GPU training, and data augmentation — sparked the deep learning revolution, laying the foundation for modern computer vision and advancing AI across industries.

vision 30u30 paper foundation

Quantifying the Rise and Fall of Complexity in Closed Systems: The Coffee Automaton

2014

Scott Aaronson, Sean M. Carroll +1

This paper proposes a quantitative framework for the rise-and-fall trajectory of complexity in closed systems, showing that a coffee-and-cream cellular automaton exhibits a bell-curve of apparent complexity when particles interact, thereby linking information theory with thermodynamics and self-organization.

foundation 30u30 paper physics science

Generative Adversarial Networks

2014

Ian J. Goodfellow, Jean Pouget-Abadie +6

The 2014 paper “Generative Adversarial Nets” (GAN) by Ian Goodfellow et al. introduced a groundbreaking framework where two neural networks — a generator and a discriminator — compete in a minimax game: the generator tries to produce realistic data, while the discriminator tries to distinguish real from fake. This approach avoids Markov chains and approximate inference, relying solely on backpropagation. GANs revolutionized generative modeling, enabling realistic image, text, and audio generation, sparking massive advances in AI creativity, deepfake technology, and research on adversarial training and robustness.

vision AIGC paper foundation

Recurrent Neural Network Regularization

2014

Wojciech Zaremba, Ilya Sutskever +1

This paper presents a method for applying dropout regularization to LSTMs by restricting it to non-recurrent connections, solving prior issues with overfitting in recurrent networks. It significantly improves generalization across diverse tasks including language modeling, speech recognition, machine translation, and image captioning. The technique allows larger RNNs to be effectively trained without compromising their ability to memorize long-term dependencies. This work helped establish dropout as a viable regularization strategy for RNNs and influenced widespread adoption in sequence modeling applications.

foundation 30u30 paper

Neural Turing Machines

2014

Alex Graves, Greg Wayne +1

This paper augments recurrent neural networks with a differentiable external memory addressed by content and location attention. Trained end-to-end, it learns algorithmic tasks like copying, sorting and associative recall from examples, proving that neural nets can induce simple programs. The idea sparked extensive work on memory-augmented models, differentiable computers, neural program synthesis and modern attention mechanisms.

foundation 30u30 paper

Order Matters Sequence to sequence for sets

2015

Oriol Vinyals, Samy Bengio +1

This paper explores how the order of inputs and outputs affects the performance of sequence-to-sequence (seq2seq) models, even when the data is unordered (e.g., sets). It introduces architectural extensions such as the Read-Process-Write model and proposes a training approach that searches over output permutations to improve learning. The paper shows that optimal ordering significantly impacts tasks like language modeling, parsing, and combinatorial problems. This work highlights the importance of considering input/output ordering in model design and has influenced further research in permutation-invariant architectures.

foundation 30u30 paper

Deep Residual Learning for Image Recognition

2015

Kaiming He, Xiangyu Zhang +2

The paper “Deep Residual Learning for Image Recognition” (ResNet, 2015) introduced residual networks with shortcut connections, allowing very deep neural networks (over 100 layers) to be effectively trained by reformulating the learning task into residual functions (F(x) = H(x) − x). This innovation solved the degradation problem in deep models, achieving state-of-the-art results on ImageNet (winning ILSVRC 2015) and COCO challenges. Its impact reshaped the design of deep learning architectures across vision and non-vision tasks, becoming a foundational backbone in modern AI systems.

foundation 30u30 paper vision

Category

Explore by categories

All

AI Leaderboard

AI Agent Tutorials

AI Coding Tutorials

AI Agent Papers

Chatbot

Machine Learning Foundation Books

AI Train

AI Deploy

AI Client

Machine Learning Foundation Papers

Machine Learning Foundation Tutorials

AI Image Demos

AI Agent

Large Language Model Tutorials

Large Language Model Papers

Machine Learning Engineering Papers

Computer Vision Tutorials

Computer Vision Papers

Natural Language Processing Papers

Reinforcement Learning Papers

Speech Technology Papers

AI API

AI Coding

AI Image

AI Video

MLOps

MCP Client

MCP Server

Computing Machinery and Intelligence

The perceptron: a probabilistic model for information storage and organization in the brain

Learning Internal Representations by Error Propagation

Keeping NN Simple by Minimizing the Description Legnth of the Weights

A Tutorial Introduction to the Minimum Description Length Principle

ImageNet Classification with Deep Convolutional Neural Networks

Quantifying the Rise and Fall of Complexity in Closed Systems: The Coffee Automaton

Generative Adversarial Networks

Recurrent Neural Network Regularization

Neural Turing Machines

Order Matters Sequence to sequence for sets

Deep Residual Learning for Image Recognition