paper

Language Models are Few-Shot Learners

Landmark paper demonstrating that large language models can perform diverse tasks with minimal task-specific training through in-context learning.

Language Models are Few-Shot Learners
Type
paper
Year
2020
By
Tom Brown et al. (OpenAI)
Publisher
OpenAI
DOI
10.48550/arXiv.2005.14165

This paper introduces GPT-3, a 175-billion parameter autoregressive language model trained on diverse internet text. The key innovation is demonstrating that large-scale language models can perform a wide range of natural language tasks—including translation, question-answering, and creative writing—with only a few examples provided in the prompt (few-shot learning), without requiring task-specific fine-tuning.

The paper systematically evaluates GPT-3 across 42 different tasks, showing that performance generally improves with model scale and the number of in-context examples. It reveals emergent abilities in few-shot settings that were not present in smaller models, suggesting that scale itself is a path to more general language understanding.

Published in May 2020, this work became foundational to modern large language model research and applications. It demonstrated the viability of scaling as a primary approach to AI capability and established the paradigm of prompt-based interaction with language models that underpins contemporary AI assistants.

Last updated 31 August 2026