OpenAI’s researchers tested a 175-billion-parameter language model on tasks specified through text, without updating the model’s weights for each task.
Why it mattered: examples inside a prompt became a practical way to steer a general language model.
The perspective: the paper also documented weaknesses and evaluation challenges. Broader capability did not remove the need to check answers.
Book a consultation