Browse abstracts and use AI to summarize the ones you don't have time to read.
The foundational paper of information theory, defining entropy and channel capacity.
The dominant sequence transduction models are based on complex recurrent or convolutional neural networks. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.
We introduce BERT, a new language representation model which stands for Bidirectional Encoder Representations from Transformers.
A purely peer-to-peer version of electronic cash would allow online payments to be sent directly from one party to another without going through a financial institution.
We present a residual learning framework to ease the training of networks substantially deeper than those used previously. Deep residual nets are easier to optimize and gain accuracy from considerably increased depth.
A formal model for privacy guarantees in statistical databases.
We propose a new framework for estimating generative models via an adversarial process, in which we simultaneously train two models: a generative model G and a discriminative model D.
Proteins are essential to life, and understanding their structure can facilitate a mechanistic understanding of their function. We present a computational method, AlphaFold, that predicts protein structures with atomic accuracy.
We trained a large, deep convolutional neural network to classify the 1.2 million high-resolution images in the ImageNet contest. On the test data, we achieved top-1 and top-5 error rates substantially better than the previous state-of-the-art.
MapReduce is a programming model and an associated implementation for processing and generating large data sets.
Here we introduce an algorithm based solely on reinforcement learning, without human data, guidance or domain knowledge beyond game rules.
We report quantum supremacy using a programmable superconducting processor.