Language Model Shape
Really interesting framing from Alex Zhang on the “shape” of language models: We’ve spent the last few years building increasingly sophisticated harnesses around a mostly fixed primitive: models that take in and put out arbitrary text. As a result, we've mostly treated agents, tool use, memory, compaction, context management, and planning as problems for the harness to solve. Alex explores what it might look like to invert that: instead of always designing the harness around the model, what if we designed models around the harness / the tasks to be completed? Jev is an example of this — rather than generating arbitrary text, its output is constrained to a value between 0 and 1, making it dramatically more specialized, and also enabling a different computational tradeoff (very fast fuzzy decisions that can become primitives inside larger agent systems). The implication is that the future may not be one increasingly intelligent general-purpose model sitting inside increasingly complicated agent harnesses. It may instead be systems composed of models with different “shapes,” each optimized for a particular role. Once frontier capabilities can be distilled into smaller/open models, this design space becomes much more practical to explore. So maybe the question worth asking now is: what should models look like if we designed them specifically for agents?
Read original source ↗