Interactive · Scaling laws

The Scaling Curve

Model capability rises predictably as dataDataThe examples a model learns from — text, images, code. More, higher-quality data generally yields a more capable model., computeComputeThe processing power used to train a model (GPU/TPU time, measured in FLOPs). More compute lets a model train on more data for longer., and parametersParametersThe internal weights a model tunes during training. More parameters let it capture more complex patterns. grow together. Drag the scale to move along the curve and watch the model generations light up.
Capability → Scale: data × compute × parameters (log) →
1e20
Estimated capability:  ·  nearest generation:

Illustrative — real scaling laws (Kaplan 2020, Chinchilla 2022) follow power laws in loss.