Patchwork AGI: a new mental model for thinking about AI

Abhay

When AI was getting started, everyone predicted that the first to get automated would be the people with soft skills, the poetry writers, or the "wordcels". The opposite has happened instead. AI has produced a proposed solution to a Millennium Prize Problem, yet we still cannot rely on it as an effective executive assistant. We predict that AI, on the current paradigm, may not progress in the way in which everyone expects (or wants) i.e. fast RSI takeoff, but rather in a "patchwork" manner: patching AI capabilities one by one (which can still be fast, but not super general).

Read original source ↗

Discussion

why expect new work created by AI to require new capabilities? what if a core set of skills lets AI handle an indefinitely expanding range of future work?

one thing I’m interested in: which gaps are intrinsic to how models learn, and which persist because nobody has found a profitable way to train them yet?

In reply to a comment

That would be a pretty strong form of generalization/portability where you have a "core" that carries over to work it wasn't trained on. Right now we argue that we're not seeing that. For example, a model that knows how to write an algorithm in Python does much worse in a language it hasn't been trained on, even if it knows the rules of that language.

In reply to a comment

Profitability is a part of it, but less so than the nature of these models' learning. In the essay we argue that models learn by piling up skills one field at a time, while people start with a general way of learning that comes before any particular skill

In reply to a comment

It might get close for a lot of work, and it'll be enormously valuable. Model providers can now spread the cost of teaching one skill across billions of requests. But without true generalization, these models won't be able to make the kinds of leaps that require a really strong world model. In the essay we give the example of general relativity. It's also not clear that, even when there is task coverage, models will handle tasks that need combinations of skills very well. Even if you're just coding, when you take a model out of its distribution and try to make something somewhat novel, it gets less useful.

That's what we've been working on since 2023. Deep Mind wrote a paper on it last December. The requirement is to be in control of it at the database and schema layer from the inside out not outside in. All labs have gotten it wrong to date.

I agree that current AI capabilities remain somewhat jagged and vary greatly between fields (and even within fields, among tasks). Especially for tasks that are easy to verify or simulate, progress has been very impressive, and it doesn't seem like we've reached a point where we have exhausted the creation of new RL environments yet. If the idea of RSI and AI accelerating AI research rapidly (through superhuman capabilities in math and software engineering) proves accurate, there may be a fast takeoff regardless. If that's not the case, data may become an even greater bottleneck. I can definitely see an industry of collecting data/workflows from businesses (for the purpose of automation of their tasks) becoming a major thing.

Sign in to join the discussion.