Writing · Architecture

When the bottleneck moves: notes on consulting after LLMs

2026 · 5 min read

I have spent enough time watching small teams outbuild large ones, in 2026, that I have stopped treating it as the exception. It is the new shape of the work.

A pattern I keep seeing: a five-person team with strong domain experts and a thoughtful architect delivers, in eight weeks, what a 500-person firm had bid six months and a small fortune for. The five-person team did not invent a new methodology. They did not have a secret toolkit. They had two ingredients the big firm had structurally optimized away: deep domain understanding, and the willingness to let LLMs do the work that used to require fifteen engineers.

I do not think most consulting buyers, or most consulting firms, have fully internalized what just happened to their business model.

What changed

The cost of producing software, in time, has collapsed.

A senior engineer paired with a capable model produces, in a day, what a small team used to produce in a week. Not for every kind of work. Not at every quality bar. But the surface area is large enough that the economics have flipped.

The cost of producing software, in money, has become a different question entirely. The old question was: how many engineers do we need, and which tech stack. The new question is: which model do we use, how many tokens will it cost, and who is sitting in the seat directing it.

A practical consequence: instead of pitch decks and MVPs, it is now possible to demonstrate a minimum working module in the same time it used to take to write the proposal. The artefact has changed. The economics behind the artefact have changed with it.

Per-resource-per-hour billing made sense when hours were the bottleneck. Hours are not the bottleneck anymore.

What did not change

What did not get cheaper is knowing what to build.

LLMs are tools, and like every tool that has ever existed, they reward the wielder who knows what they are doing. A domain expert with an LLM moves ten times faster than they did a year ago. A generalist with an LLM produces confident-looking output that solves the wrong problem at impressive speed.

The mistake the industry is making, in my reading, is concluding that LLMs reduce the importance of expertise. The opposite has happened. Expertise has become more valuable, not less, because every other input in the equation got cheaper. The leverage of being the right person in the seat has multiplied.

In economic terms: when execution costs collapse, judgement costs rise. The bottleneck moved.

Where the bottleneck moved to

The new bottleneck is not engineering capacity. It is the chain of:

  • Did we understand the actual problem the customer has?
  • Did we frame it correctly?
  • Did we pick the right model, prompt, architecture, and abstraction for it?
  • Do we know what good looks like once the model produces a draft?
  • Do we know what to throw away?

Each of those is a judgement task. None of them is solved by adding headcount. Several of them are made worse by adding headcount, because judgement does not parallelize cleanly.

This is why a five-person team of domain experts and a senior architect can outperform a 500-person firm on a specific engagement. The five-person team is unblocked at the bottleneck. The 500-person firm is paying overhead on the wrong constraint.

The two questions that have switched places

Two questions used to be central to scoping a software engagement:

  1. How many engineers do we need?
  2. Which technology stack?

Both assumed that engineers were the scarce input, and that the stack determined how productively those engineers could work.

In 2026, both questions are downstream of two different ones:

  1. Which LLM do we route which class of work to?
  2. Whose domain knowledge and whose architectural judgement are in the seat?

If you cannot answer those two, picking a stack and counting heads is premature. It is solving for the wrong constraint.

What this means for buyers

If you are buying consulting in 2026, the question to ask is no longer "how many people will you put on this." That question used to correlate with seriousness; it now correlates with overhead.

The questions that correlate with outcomes are different:

  • Who specifically will be doing the thinking, and what is their depth in our domain?
  • What does your team's working module look like in week two, not your presentation in week six?
  • How do you decide which work goes to a model and which to a human?
  • How do your fees relate to outcomes, not hours?

A small firm that answers those four well is worth more to you than a large firm with a hundred-page proposal and a billing sheet.

What this means for builders

If you run a small firm, the structural advantage you have is real. You should be explicit about it.

Stop competing on headcount. You will lose. Compete on the things that have actually become scarce: domain depth, architectural judgement, the willingness to ship a working module in weeks instead of a proposal in months, and the discipline to use LLMs as force multipliers rather than as the entire product.

The risk for a small firm is the opposite one: assuming that because LLMs are doing more of the typing, you can be lighter on expertise. You cannot. The expertise is what made the LLM useful in the first place. Take it out and you become every other LLM-wrapper, which is a market with no moat and a falling floor.

A short, fast window

I do not think this inversion will stay quiet for long. Big firms will adjust, in their own way. Buyer expectations will catch up. Pricing models will get reset.

In the meantime, a window is open that I have not seen open in software services in twenty years. A small team of people who know what they are doing, paired with capable models, can produce work that used to require scale. The customer benefits. The team benefits. The middlemen do not.

If your industry has not noticed yet, that is your window.

Closing

The sentence I keep coming back to as the cleanest frame for the present moment: when execution costs collapse, judgement costs rise.

The consultancies that will compound through the next five years are the ones that hold judgement and domain expertise at the centre and treat LLMs as the cheapest part of the stack, rather than the most expensive part of the pitch.

The ones that will not are the ones still selling hours.

Drafted in 2026
Updated for site in October 2026