outmanage.

In the room · step 5

Build, buy, or extend — and when to do none of them

Microsoft's own blueprint makes this a testable decision, which tells you it's a real one. The framing that helps is what you'd be sad to have someone else own.

8 min readpublished and checked 2026-07-29

"Should we build this or buy it" is the question that arrives once someone has decided the capability is worth having. It is usually posed as a binary, and it has at least four answers.

Microsoft's AI Transformation Leader blueprint includes "identify when to build, buy, or extend" as an explicit objective — it sits in the domain about mapping business processes to AI tooling, which carries 35–40% of that exam. Worth noting because it means the three-way distinction is considered part of a manager's job rather than an architect's.

The four options, in order of how often they're right

Buy a product that does the whole job. Someone else owns the model choice, the evaluation, the maintenance, and the bill. You own configuration and adoption.

Extend something you already pay for. This is the option most often missed. If your organisation already runs a productivity suite with AI features, an existing search platform, or a CRM with assistive tooling, some meaningful proportion of proposals arriving as new projects are things that platform does already, or does adequately, or will do next quarter. Microsoft names extensibility frameworks specifically. Checking costs an afternoon.

Build on a managed model API. You write the application; the provider runs the model. AWS frames the deployment choice as managed API service versus self-hosting, and for almost every organisation that is not itself an AI company, the managed API is the answer — you inherit scaling, patching, and availability rather than staffing them.

Wait. Not a failure of nerve. Sometimes the honest position is that the capability is six months from being reliable enough, or that you cannot yet write an evaluation for the thing you want, and starting without one means you will not be able to tell whether you succeeded.

Self-hosting your own model deserves a mention mainly so it can be dismissed. It makes sense at sustained high volume, with strict data-residency constraints that rule out the alternatives, and with people who want to own that operational burden. If your team does not currently run inference infrastructure, this is a much larger commitment than it appears in the slide where it costs less per request.

The question that actually decides it

Not cost. Cost differences are usually smaller than the confidence intervals on your volume forecast.

The useful question is: which part of this would you be unhappy for a competitor to have an identical copy of?

If the answer is "none of it" — meeting transcription, document search, first-draft replies, translation — buy or extend. This is undifferentiated work. Every organisation in your sector needs it, nobody wins on it, and building it yourself means maintaining forever something you could have rented.

If the answer names something specific — your pricing logic, your triage rules, the twelve years of resolved cases nobody else has, the particular way your service works — that part is worth building, and only that part. The pattern that tends to hold up: buy the commodity layer, build the thin piece that is actually yours, and resist the urge to build the plumbing underneath it because the plumbing is more fun.

Six questions before you sign anything

Applies to buy and extend equally.

1. What happens to our data? Is it used for training, retained, and where does it sit? For a regulated organisation this can eliminate options outright, so ask before you get attached.

2. Can we get our work back out? Prompts, configuration, evaluation sets, conversation history. If the evaluation set you spent three weeks building lives only inside their product, you cannot ever run a fair comparison against an alternative. This is the lock-in that bites hardest and the one nobody asks about.

3. What model is underneath, and what happens when it changes? Providers update models. Behaviour shifts. Ask whether you get notice, whether you can pin a version, and whether they re-run their evaluations when the underlying model moves.

4. How does the price scale — with seats, usage, or conversation length? These behave very differently at scale, and the third one surprises people, because every turn of a conversation resends everything before it.

5. What do we own if we leave in two years? Not a hostile question. A planning one.

6. Who is accountable when it produces something wrong? The answer, under every shared responsibility model you will encounter, is you — for the use case, the data, and the configuration. Any supplier suggesting otherwise is either confused or selling.

If you build: start cheaper than you think you need

One piece of engineering guidance translates straight into a management decision. Anthropic's advice on choosing a model offers two routes: start with a fast, inexpensive model and upgrade only where you find a real capability gap, or start with the most capable model and optimise downward once the workflow is understood. The first suits high-volume, latency-sensitive, straightforward work; the second suits genuinely hard reasoning where accuracy dominates cost.

The management version: do not let the first build use the most expensive option by default. The usual pattern is to reach for maximum capability to de-risk the pilot, then discover the unit economics don't work at volume and rebuild. Starting cheap and upgrading where the evaluation demands it gets you the same place with the cost curve already understood.

There is a further lever worth knowing exists, because it changes the shape of this decision: recent models expose an effort setting that trades reasoning depth against latency and cost within a single model. Tuning that is often a better move than switching models entirely — so "it's too slow" and "it's too expensive" are not automatically arguments for a different supplier.

The recommendation to write down

Whatever you decide, record the reasoning in three lines: what we judged undifferentiated, what we judged ours, and what we'd need to see to revisit this. In eighteen months somebody will ask why the architecture looks the way it does, and the alternative to those three lines is an archaeology exercise.

Next in In the room. Deciding not to proceed 8 min read.

Also worth reading

In the roomWhat a demo can't show youA demo is a performance of the best case. Three requests turn it into evidence, and all of them take under a minute.In the roomHow to read an accuracy claim"94% accurate" is not a fact about a system. It's a fact about a test somebody designed, and the design is where the interesting part lives.

Get the next one.

One email when something new lands. Nothing else.

Get an email when a new guide or article is published. Read how we use your email address in our privacy.