Scaling Laws¶ The training loss \(L\) decreases predictably as we scale up model size \(N\) , dataset size \(D\) , and compute \(C\) , following a power-law curve, which appears as a straight line on a log-log plot.