Division of labor between AI models

The meaningful unit of AI work is shifting from a single model to a division of labor where different specialized models route tasks across different systems.


Nathan Labenz, who hosts the Cognitive Revolution podcast and has spent years interviewing the people building AI systems, has landed on a concise way of describing where AI work is heading:

"The interesting unit is no longer one model. It's the division of labor between models."

The claim behind that line is worth unpacking, because it changes how you should think about getting good results out of AI.

For most of the time AI assistants have been widely available, the implicit question has been: which single model is best? People argue over rankings, switch subscriptions when a new release tops a leaderboard, and generally treat the choice of one model as the decision that determines the quality of the output. Labenz's point is that this framing is becoming outdated. The meaningful unit of AI work is shifting away from the individual model and toward how tasks are split up and routed between different specialized models and systems.

In plain terms: instead of one generalist trying to do everything, a well-designed setup hands each part of a job to whatever handles it best. A complex task — say, researching a question, drafting an answer, checking it for errors, and formatting it for a particular audience — can be broken into steps, and each step routed to a model or tool suited to it. Smaller, cheaper, more specialized components can outperform a single expensive general-purpose model, because each piece is doing the narrow thing it is good at rather than one thing doing everything passably.

Who is this for, honestly? Mostly people designing workflows and pipelines — the systems that move work through an organization. In practice that skews toward developers and technical teams, because today the routing between models is something you build: you decide which model sees which task, you wire up the steps, you handle the handoffs. If you are a non-developer using a single chatbot on your phone, this idea does not yet hand you much you can act on directly — and it is worth saying so plainly rather than pretending otherwise. What it does offer you is a more accurate mental model. When a product you use quietly improves, it is increasingly likely to be because of this kind of behind-the-scenes orchestration rather than because one model got smarter.

Is this usable now, or just talk? It is shipping. Multi-model systems, routing layers, and agentic pipelines — where one model's output becomes another's input — are running in production today, not just being discussed on podcasts. Labenz's remark is a description of what is already happening, not a prediction.

The limits deserve equal billing. Splitting work across models adds complexity: every handoff is a place where context gets lost, errors compound, and costs become harder to predict. A pipeline of five cheap models is not automatically better than one good one — it is better only if someone designed the division of labor well. And the evaluation problem gets harder, not easier: when the final output is wrong, figuring out which step in the chain failed is its own job. The claim tells you where the leverage is moving. It does not promise the leverage is easy to use.

automationfinanceproductsvideodeveloper
Source: youtube.com