- Reuses knowledge from a pre-trained model, cutting data and compute needs
- Achieves strong performance on tasks with limited labelled data
- Speeds development by starting from a capable base rather than scratch
- Makes advanced models practical for narrow, data-scarce problems