An engineer on your team spends a weekend with an API key and comes back Monday with something that works. It reads a support ticket, recognizes the customer is furious, drafts a response, flags the account for escalation, and suggests a next step. Total cost: about forty dollars in tokens.

The room goes quiet in the good way. Then someone says the sentence that launches a thousand internal platform teams: “Why are we paying a vendor for this?”

It is a fair question. It is also the wrong one.

In 2026, the answer to “can we build it?” is almost always yes. Capability has been democratized. The interesting question is narrower and much harder: can you run it? Not on Monday in the demo room, but on a Tuesday eighteen months from now, at 3 a.m., under production load, under audit, with the original engineer at a different company and the underlying model deprecated.

That is a different question. It has a different answer.

The industry is confusing time to wow with time to value

The evidence on this is not subtle anymore.

95%of organizations reported no measurable P&L impact from their generative AI initiativesMIT Project NANDA, 2025
40%+of agentic AI projects forecast to be scrapped by end of 2027 on cost, unclear value, or weak risk controlsGartner, 2025
2×deployment success rate for vendor and partnership approaches versus internal buildsMIT Project NANDA, 2025

That last figure is the one least often quoted and most relevant here: purchasing from specialized vendors succeeded roughly 67% of the time, while internal builds succeeded about one third as often.

Two caveats, because the argument is stronger when it is honest. There is a reasonable selection-effect critique: vendors often get the projects that have already survived budget scrutiny and defined their data scope, while internal teams inherit the messy ones. And the 95% headline has been challenged for lacking underlying data in the report itself. Take both numbers as directional rather than dispositive.

But the direction is unambiguous, and it matches what anyone who has shipped AI in production already knows.

The demo is the cheapest five percent of the work.

A demo runs against twenty tickets you selected. Production runs against forty thousand you did not. Between those two states sits everything that actually determines whether the project returns money:

  • The long tail. The 3% of cases that are ambiguous, adversarial, multilingual, or wrong in ways your prompt never anticipated.
  • Evaluation. How do you know the system got better after a change? Vibes do not survive a QBR. Labeled golden sets, regression suites, and drift detection are a real engineering program.
  • Regression. You fix one prompt behavior and silently break three others. Without evals you will not find out until a customer does.
  • Human in the loop. Review queues, override paths, confidence thresholds, audit trails, and the workflow surgery required to make any of it fit how people actually work.
  • Unit economics at volume. Forty dollars becomes forty thousand. Caching, routing, distillation, and batching are the difference between a margin and a write-off.
  • Failure behavior. What the system does when the model is down, slow, rate-limited, or confidently wrong.

None of this appears in the demo. All of it appears in the invoice.

Everyone has the same models. Nobody has the same plumbing.

Here is the part that quietly undermines most build cases.

The frontier models are available to you and to every vendor you might buy from, on the same terms, through the same APIs, priced within noise of each other. Your engineer’s weekend prototype and a funded product team’s roadmap start from an identical substrate. Capability is not scarce. It is a utility.

Which means the thing your build case is implicitly betting on, that you can match a vendor’s features, is not the hard part. It is the easy part. Features are downstream of a model everyone shares.

We have written before about why “just use Claude” is the new “just use Oracle”: the single-model assumption breaks the moment you hit production, because reliability, cost, latency, and regional data residency all push you toward a portfolio of models rather than one. Build teams tend to discover this the hard way, usually in quarter three.

The differentiator was never the model. It is the plumbing underneath: the substrate that makes intelligence dependable, affordable, auditable, and durable across model generations. That layer is unglamorous, invisible when it works, and existential when it does not. It is also where essentially all of the cost lives.

The enterprise AI iceberg A small visible tip labelled “what you demo” sits above the waterline; a much larger submerged mass labelled “what you operate forever” lists architecture, model strategy, reliability, scale, cost control, data foundation, security, compliance, and observability. What you demo Prompts, features, model calls, a happy path, 20 chosen tickets ~5% of the work ~1 weekend PRODUCTION WATERLINE What you operate forever Architecture / orchestration framework Model choice, routing, fallback, re-qualification Reliability, failover, graceful degradation Scale, queueing, backpressure, concurrency Token economics, caching, tiering, batching Data ingestion, normalization, lineage, retention Prompt injection, tool permissioning, exfiltration GDPR, CCPA, EU AI Act, SOC 2, data residency Evals, tracing, cost attribution, decision logs On-call, incident response, migrations ~95% of the cost none of it moves the needle
Fig 1. Every item below the waterline is mandatory, permanent, and commercially invisible. This is the textbook definition of undifferentiated heavy lifting.

The part of the iceberg that nobody demos

Ask what an internal AI platform actually has to own, permanently, and the list is sobering.

Architecture

Orchestration model, state management, retrieval strategy, tool interfaces, memory, and the agent framework itself. Frameworks in this category have a shelf life measured in quarters.

Model strategy

Selection, qualification, routing, fallback, and re-qualification with every release and every deprecation notice. End-of-life for a model version is not a vendor’s problem when you are the vendor.

Reliability

Multi-region failover, graceful degradation, timeouts, idempotent retries, and circuit breakers around a dependency that is probabilistic by design.

Scale

Queueing, backpressure, and concurrency limits that absorb the traffic spike arriving exactly when your customers are angriest.

Cost vs. performance

Token accounting, prompt caching, context compression, model tiering, batch versus real time. A continuous optimization problem, not a one-time decision.

Data foundation

Ingestion from every channel, normalization, entity resolution, lineage, retention, and a warehouse architecture that stays queryable and governed.

Security

Prompt injection, tool permissioning, exfiltration through context, and supply chain risk across the agent stack. The OWASP Top 10 for LLM Applications is a floor, not a ceiling, and it moves.

Privacy and regulation

GDPR, CCPA/CPRA, EU AI Act obligations for high-risk and transparency, Colorado’s AI Act, India’s DPDP Act, sector rules such as HIPAA, plus the residency commitments in your own customer contracts. Then SOC 2, the annual audit, and every enterprise security questionnaire.

Observability

Tracing, cost attribution, quality monitoring, model cards, and decision logs good enough to explain to a regulator or a customer why the system did what it did.

Read that list again and notice the pattern: not one line of it moves the business needle, and every line of it is mandatory. It is work you cannot skip, cannot monetize, and cannot stop doing.

Architecture decisions now have a half-life

Enterprise IT used to make an architecture decision and live with it for seven to ten years. That assumption is gone.

Look at the last three years alone. Naive RAG gave way to agentic retrieval. Fine-tuning gave way to long context, which gave way to deliberate context engineering. Orchestration frameworks rose, forked, and were abandoned. Tool-calling conventions fragmented, then began converging on the Model Context Protocol. Model generations turned over faster than most enterprises complete a procurement cycle.

An internal build does not cost you one architecture. It commits you to re-deciding architecture roughly every eighteen months, forever, with full regression and migration cost each time, funded out of a budget that competes against revenue-generating work in every planning cycle.

Cumulative cost of building versus buying over five years A line chart showing buy cost rising steadily and predictably, while build cost starts lower but steps sharply upward at each re-architecture cycle, crossing above the buy line early in year two and diverging afterwards. Year 0 Year 1 Year 2 Year 3 Year 4 Year 5 Cumulative cost Buy: predictable, absorbed Build: stepped by every rebuild Re-architecture Re-architecture Crossover, typically before the first renewal would have come due the only stretch where build looks cheap
Fig 2. Illustrative. The build case is almost always argued on the first few months, which is the only period in which it wins.

That budget competition has a predictable winner. It is why so many internal AI platforms plateau at 60% of their intended scope and then quietly stop being upgraded. Not because the team was weak, but because the second and third rebuilds never got funded.

Vendors have no such option. Rebuilding the substrate is the business. It gets funded because it must.

The vantage point problem: what one deployment cannot teach you

An internal team learns from one deployment. Sample size: N = 1.

A vendor learns from every customer. Which prompt patterns hold up after nine months and which look brilliant in week one and collapse under drift. Which human-in-the-loop designs get adopted and which get routed around. Which taxonomies survive contact with a real support organization. Which agent autonomy levels are safe in a regulated context and which produce a postmortem.

The hardest thing to acquire alone is negative knowledge: the catalog of approaches that already failed somewhere else.

That is the difference between a software license and a trusted advisor. The license is the artifact. The pattern library, the benchmarks across comparable organizations, the answer to “what usually goes wrong at month four,” and the ability to say do not do that, we have watched three companies try is the part you cannot replicate by hiring well.

Software is not the deliverable. Accountability is.

Strip away the technology and this is the real question: who stands behind it?

When you buy, someone else carries the pager. Someone else owns the SOC 2 report your customer’s security team demands, signs the DPA, maintains the subprocessor list, funds the penetration test, indemnifies you, and takes the call when the system is wrong in a way that reaches a customer. Someone else is contractually obligated to keep the thing current through the next three model generations.

When you build, you do not merely become the engineer. You become the vendor. You inherit the roadmap, the on-call rotation, the compliance calendar, the security questionnaire, the documentation debt, and the bus factor. Permanently, for a system that is not the product your customers buy from you.

There is one more asymmetry worth naming. A vendor faces a renewal every year. If the product stops earning its keep, it gets fired. An internal platform never faces that test. No renewal moment, no competitive pressure, no mechanism that forces an honest reckoning with whether it is still the right answer. That is not a small governance detail. It is why underperforming internal systems persist for years after the market has passed them.

So when should you build?

Sometimes you should. The answer is not “always buy,” and any vendor telling you otherwise is selling rather than advising. Three questions decide it.

  1. Is this capability something your customers pay you for? Core versus context. If the AI is the product, build it. If it supports the product, buying it back as infrastructure is almost always the better trade.
  2. Do you hold proprietary data or process that no vendor can replicate? Not “we are unique” in the way every company believes it is unique. Genuinely proprietary: a dataset, domain model, or workflow structurally unavailable to anyone else.
  3. Can you fund the second, third, and fourth rebuild? Anyone can fund the first build; it is exciting and it demos well. The question is whether you can protect that budget in year three, when the excitement is gone and the work is a migration.

Three yeses: build, and build with conviction. Two or fewer: you are not choosing to build, you are choosing to find out.

Where the money actually goes, over a realistic horizon
Cost lineBuildBuy
Initial capabilityWeeks. Cheap. Genuinely impressive.Weeks. Cheap. Also impressive.
Data integration and normalizationYours, per source, foreverIncluded and maintained
Eval harness and golden setsBuild it, staff it, keep it currentIncluded
Security posture, OWASP LLM coverageYours to design, test, and re-testIncluded and independently audited
SOC 2, GDPR, CCPA, EU AI ActYours to attest and renew annuallyContractual, with indemnity
Model migration and re-qualificationFull project, every generationAbsorbed by the vendor
Reliability, HA, on-callYour SRE headcountYour vendor’s SLA
RoadmapCompetes with revenue work each quarterFunded as the vendor’s core business
Cross-customer pattern learningN = 1N = many
Failure accountabilityInternalContractual

The third answer: build on infrastructure, not from scratch

“Build versus buy” is a false binary, and it is the reason so many of these decisions go badly. The real choice is about where in the stack you spend your engineering.

Nobody builds their own database anymore. Nobody builds their own identity provider, message queue, or observability pipeline. All of those were once things smart teams built themselves, and the industry eventually concluded that the plumbing should be bought and the differentiation built on top.

Enterprise AI is at exactly that inflection. It is the thesis behind SupportLogic’s AI Infrastructure for CX positioning: buy the substrate, build the differentiation.

  • A Snowflake-backed Data Cloud that unifies signals across every support channel, with your data inside your governance perimeter.
  • CRM-Less Architecture, so your intelligence layer is not hostage to a system of engagement’s data model or release cycle.
  • An MCP Server that exposes that intelligence to your own agents, copilots, and internal tools through an open standard rather than a proprietary one.
  • A suite of autonomous AI agents for escalation, sentiment, prioritization, routing, and voice, maintained and upgraded as our full-time job rather than your side project.

Your engineers keep the weekend-prototype energy. They just spend it on the five percent that is actually yours, instead of rebuilding the ninety-five percent that is not.

The honest case against buying

I run a vendor, so treat what follows as the strongest argument against my own position.

Buying carries real risk. Roadmaps drift from your priorities. Vendors get acquired or go away. Proprietary data models create lock-in that gets expensive precisely when you have the least leverage. And some vendors treat “AI infrastructure” as a label rather than an architecture, the practice Gartner has called agent washing.

The mitigations are structural, and you should demand them in writing:

  • Open standards at the interfaces. MCP for tool and context interop, standard warehouse formats for data. A proprietary integration surface is a lock-in decision, not a technical one.
  • Your data in your perimeter. If the architecture requires a full copy of your data in an opaque vendor store, ask why.
  • A documented exit. What you can extract, in what format, on what timeline.
  • Evidence of the plumbing. Ask for the eval methodology, the model migration history, the uptime record, and the compliance attestations. Vendors doing the hard part are happy to show you. The ones doing agent washing change the subject to features.

A good vendor reduces your lock-in and takes the undifferentiated work off your plate. Hold us to that standard.

Five questions for your next build-vs-buy meeting

  1. What is our evaluation harness, who maintains it, and how do we prove the system got better after a change?
  2. What is our plan for the day our chosen model is deprecated, and how many prompts and workflows need re-qualification?
  3. Who is on call for this at 3 a.m., and what internal SLA are we committing to?
  4. What do we tell a customer’s security team when they ask for a SOC 2 report and a DPA covering this system?
  5. If the two engineers who built it leave, what happens in the following quarter?

Crisp answers to all five? Build it. You have thought about the right things and you will probably succeed. Anything less, and you have not decided to build. You have decided to find out, on a schedule you do not control, in front of customers.

Just because you can build it does not mean you should. That is not a knock on your engineers. It is a recognition that their time is the scarcest asset you have, and that spending it on plumbing everyone else has already built is the most expensive thing you can do with it.

Closing thoughts

When you build software internally, you don’t just become the engineer. You become the vendor.

You inherit the roadmap. The bug fixes. The on-call rotation. Secure SDLC practice for SOC 2. The DPA. GDPR, CCPA, and EU AI Act obligations. Running the pen test and answering the same security questionnaire three times a year. Model drift. The FinOps discipline to keep cloud costs from quietly eating the business case.

And you repeat all of it for every generation of LLM.

Time to wow is easy. TCO is not.

All for a system that is not the product your customers buy from you.

An external vendor faces a renewal every year. If the product stops earning its keep, it gets fired. An internal platform never faces that test. No renewal moment, no competitive pressure, no mechanism that forces an honest reckoning about whether it’s still the right answer.

It’s why underperforming internal systems live on as zombies for years after the market has passed them.

Software isn’t the deliverable. Accountability is. That’s what you’re actually paying for.

Take eggs.

You can raise chickens in your backyard. Fence the run, feed them, vaccinate them, clean the mess, and yes, you will be rewarded with eggs. None of that is out of reach for anyone.

The question was never whether you can.

Frequently asked questions

Should we build or buy enterprise AI?

Build when the AI capability is the product your customers pay for, when you hold genuinely proprietary data or process, and when you can fund the second and third rebuilds as well as the first. Buy when the capability is infrastructure supporting your product rather than being your product. MIT’s 2025 NANDA research found vendor-purchased and partnership approaches reached successful deployment roughly 67% of the time versus about one third as often for internal builds.

Isn’t building cheaper now that the models are so cheap?

Model inference is the smallest line item in a production AI system. The dominant costs are data integration, evaluation infrastructure, reliability engineering, security, compliance attestation, and model migration work repeated with every generation. Those costs recur and scale with ambition, not with token spend.

What is the real TCO of an internal enterprise AI build?

Budget for the initial build, then assume a substantial re-architecture roughly every eighteen months, plus continuous cost lines for evaluations, on-call, security review, annual compliance audits, and model re-qualification. The initial build is typically the cheapest year you will have.

Doesn’t buying create vendor lock-in?

It can, which is why interfaces matter more than features. Require open standards such as the Model Context Protocol at the integration surface, insist your data stays in your governance perimeter in a standard warehouse format, and get a documented exit path in the contract. Lock-in is an architecture choice, not an inevitability of buying.

What should we build ourselves?

Your differentiation: domain-specific workflows, proprietary taxonomies, business logic reflecting how your company actually operates, and experiences unique to your customers. Build those on top of purchased infrastructure rather than beneath them.

How often does an enterprise AI architecture need to be revisited?

Meaningful architectural change is arriving on a roughly twelve to eighteen month cycle. Retrieval strategies, orchestration frameworks, tool-calling conventions, and model generations have all turned over multiple times in the last three years. Any TCO model assuming a stable multi-year architecture understates the cost of building.

KR

Krishna Raj Raja is the Founder and CEO of SupportLogic and the author of The Support Experience. He has spent roughly three decades in enterprise computing, most of it close to the support function.

See the substrate before you rebuild it

If you are weighing a build, it is worth seeing exactly what you would be reconstructing. We will walk your architects through the Data Cloud, the MCP Server, the eval methodology, and the compliance posture.

Request a technical walkthrough Read the architecture overview

Don’t miss out

Want the latest B2B Support, AI and ML blogs delivered straight to your inbox?

Subscription Form