Will transformers scale to AGI or do we need a different architecture?
I’m curious if people think transformers will be the architecture AGI is born out of. People seem to fall into 3 buckets: 1. Scaling the existing recipe is enough and more parameters, data and compute will get us there. Scaling laws predict improvements, but are they enough for general intelligence? 2. The architecture is sufficient, but the learning recipe needs to evolve (more RL, interaction with environments, more inference time compute, multimodal learning cross text, images, video and audio). 3. We need a totally different architecture or representation of the world (e.g. changes to memory, learning mechanisms, or tokenizing in 4D/3D vs 2D might be needed for spatial intelligence). What do folks here think?