The Attention Is All You Need paper proposed an architecture built around attention rather than recurrent or convolutional layers. It was initially demonstrated on machine translation.

Why it mattered: processing relationships between words through attention opened a different path for building and scaling language models.

What came next: Transformer-based models became a major foundation for later language systems.