Aug 11, 2026 | Journal
A persistent mystery in scaling language models is that certain abilities seem to appear abruptly. Multi-step reasoning, code generation, and some forms of instruction following improve slowly or not at all across small models, then jump at a particular scale....