How to Build a Multi-Model AI Strategy Without Running Everything Twice
Fable 5 vanished for 19 days. GPT-5.6 launched gated. The past month is the strongest case yet for multi-model, and for not over-hedging.

Ryan Drake
Founder, Ential · Jul 13, 2026 · 8 min read
Key takeaways
- Single-vendor dependency now means capability can disappear mid-quarter for regulatory reasons; outages are no longer the main threat.
- Your evals, data layer, skills and task definitions travel between vendors; prompt quirks and vendor-specific features do not.
- Judge models with your own evals, because leaderboards flip by index: Sol leads one coding index by 2.8 points while Mythos 5 leads SWE-Bench Pro by 15.7, and both are true.
- Go deep with one primary vendor for the compounding benefits, keep a tested live secondary per critical workflow, and review the pairing quarterly.
An executive asked us last month: "How do we adopt AI without ending up locked to one frontier lab?" A year ago I would have given a slightly theoretical answer about API abstraction layers. Now I just point at the last 30 days, because we have never had a cleaner demonstration of why multi-model matters, or a cleaner demonstration of why over-hedging is its own mistake.
Both lessons arrived in the same month. If you only take the first one, you will end up running every workload twice and paying for the privilege.
What Actually Happened in the Last 30 Days?
The short version: the two leading labs each had their flagship model taken off the table for weeks, for regulatory reasons, within days of each other.
Anthropic launched Claude Fable 5 on 9 June, a model above the Opus tier that works autonomously for extended periods and holds persistent memory across millions of tokens. Three days later, on 12 June, the US government imposed export controls after Amazon researchers found a safeguard bypass, and Anthropic suspended access. The controls lifted on 30 June and Fable 5 was redeployed globally on 1 July with an improved classifier. Total time off the board: 19 days. Mid-quarter. For every customer on the planet.
OpenAI's month rhymed. On 26 June the company previewed GPT-5.6 but restricted it, at the US government's request, to some 20 approved partner organisations. It was the first non-public major launch since GPT-4, and it stayed gated for two weeks before general availability arrived on 9 July with the new Sol, Terra and Luna tier names.
Google, meanwhile, quietly shipped computer-use agents in Gemini 3.5 Flash on 24 June (agents that see, click, type and navigate real interfaces) and crossed 900 million monthly Gemini users. Three labs, three genuinely frontier products, one month.
The lesson executives keep drawing from outages ("keep a backup vendor in case the API goes down") is now too narrow. Fable 5 did not go down. It was suspended by a government. Your vendor's capability can vanish mid-quarter for reasons that have nothing to do with their infrastructure and nothing to do with you. Pre-release government review, export controls and gated rollouts are the new normal, which means single-vendor dependency is no longer a resilience question. It is a regulatory exposure question.
How Do You Avoid Vendor Lock-In With One Frontier Lab?
You avoid lock-in by being precise about which of your AI assets are portable and which are not, then investing deliberately in the portable ones. Most companies get this backwards. They obsess over prompt libraries (barely portable) and neglect evals and data plumbing (entirely portable).
Here is the split as we see it after running rollouts across both major vendors.
What travels between vendors:
- Your evals. A test suite of 30 real tasks from your business, with pass criteria, runs against any model with an API. This is the single most portable and most valuable asset you can build.
- Your data layer. Clean, well-structured, retrievable company data serves whichever model reads it. Nobody's model licence covers your CRM hygiene.
- Your skills and SOPs. A written procedure for "how we produce a client proposal" or "how we reconcile the month" is model-agnostic by nature. The model executes it; it does not own it.
- Your harness patterns. Approval gates, permission scopes, review checkpoints, audit trails. The supervision architecture around agents transfers almost untouched.
- Your task definitions. Knowing exactly which 40 workflows you have delegated to AI, with owners and success criteria, is worth more than any individual integration.
What does not travel:
- Prompt quirks. The phrasing that coaxes one model into the right format is folklore, not infrastructure. Budget to rewrite it on migration and stop treating it as an asset.
- Vendor-specific features. OpenAI's new caching rules (cache writes at 1.25x uncached input, explicit breakpoints, 30-minute minimum cache life) are real money if you use them well, and they are meaningless the moment you switch. Same for any vendor's memory, connectors or proprietary tool-calling modes. Use them, enjoy them, but never let a critical workflow depend on one without a documented fallback.
If 80% of your AI investment sits in the first list, switching vendors is a fortnight of work. If it sits in the second list, you are locked in regardless of what your procurement contract says.
How Do You Know Which Models Are Actually Good?
You know by testing them on your own work, because the public leaderboards now disagree with each other on the most basic questions. This month's benchmark results are a beautiful illustration. On the Artificial Analysis Coding Agent Index, GPT-5.6 Sol scores 80, which is 2.8 points above Claude Fable 5, while using under half the output tokens at a third of the cost (index-reported). On SWE-Bench Pro, Claude Mythos 5 leads Sol by 15.7 points, 80.3% to 64.6%. Both results are true. They simply measure different things, and neither measures your business.
The practical move is a private eval: 20 to 30 real tasks pulled from your actual operations (the proposal you sent last Tuesday, the reconciliation that takes your finance lead half a day, the contract summary your ops manager writes weekly), each with a clear pass or fail definition. Run every candidate model through it quarterly. The whole exercise costs a few hundred dollars in tokens and settles arguments that would otherwise run on vibes and vendor keynotes. I have written a fuller method in how to evaluate AI agents before you trust them, and it applies verbatim to model selection.
One more reason to own your evals: they are how you route. Which brings us to architecture.
What Does a Model-Agnostic Setup Actually Look Like?
It looks like routing by task type behind one thin interface, not running every workload on two models simultaneously. Model-agnostic does not mean model-indifferent, and it certainly does not mean duplicated. Here is the shape we build for clients:
- 1One gateway. All AI calls go through a single internal service or router, so a model swap is a configuration change, not a six-week engineering project.
- 2Routing by task type. High-stakes, long-horizon work goes to your frontier model. High-volume, well-defined work goes to a mid-tier model like Terra at $2.50/$15 per million tokens or Opus 4.8, which is where most mid-market workloads belong anyway. Cheap classification and extraction go to the budget tier. The pricing spread is now wide enough that routing well can change your bill materially, which is why token spend is the new cloud bill.
- 3A tested fallback per critical workflow. Not "we could probably move it". For each workflow you genuinely depend on, you have run it end-to-end on the secondary model in the last quarter and recorded the result. When Fable 5 went dark for 19 days, the companies that shrugged were the ones with this in place. Untested fallbacks are hope wearing a lanyard.
- 4Evals as the referee. When a new model ships, you run the suite, compare, and reroute if the numbers justify it. No migration by press release.
Notice what is absent: no dual-running of production workloads, no lowest-common-denominator prompts that use nobody's strengths, no six-vendor procurement spread that leaves you shallow everywhere. Over-hedging has a real cost, and it is usually paid in mediocrity.
How Should You Structure Vendor Partnerships?
Go deep with one primary vendor, keep one secondary genuinely live, and review the pairing quarterly. Depth compounds in ways a hedged posture never captures: your team builds fluency with one set of tools, your skills library matures against one model's behaviour, your admin and governance settings get properly configured rather than half-configured twice. The upside of commitment is visible in what deep adopters achieve; Anthropic reports Stripe migrated a 50-million-line codebase in a day with Fable 5, against an estimated two months manually. Nobody produces that result while keeping three vendors at polite arm's length.
The secondary vendor is not a museum piece. Give it real work: a workflow or two in production, enough volume that your team stays current and your fallback paths stay tested. Then, once a quarter, review with fresh eval numbers: is the primary still earning its position, has pricing moved, has a capability gap opened or closed. This is a two-hour meeting, not a re-platforming. Most quarters the answer is "no change", and that is fine. The point is that the option is warm.
Executives also ask me where the big bets belong, given all this hedging. My answer: make your big bets in the workflow layer, not the model layer. Committing hard to "we will automate our entire proposal pipeline this year" is a bet that pays off under any vendor, because the task definitions, data and evals travel. Committing hard to "we are a single-lab shop forever" is a bet a regulator can void on a Thursday. Bet big on what you automate. Stay flexible on what you automate it with.
What Does the Bleeding Edge of Applied AI Look Like Right Now?
It looks like supervised delegation across a portfolio of models, not loyalty to a single chat window. The frontier companies we work with kick off long-horizon agent work on a frontier model, route the routine volume to mid-tier models, and use computer-use agents (like the ones Google shipped in Gemini 3.5 Flash) to reach the legacy, API-less software that every mid-market business still runs. The human bottleneck has moved from doing the work to specifying and reviewing it, and the vendor question has become an operations question: routing, fallbacks, evals, quarterly reviews.
The past month settled the argument. One lab's flagship was suspended for 19 days; the other's launched gated to a partner list. Multi-model is no longer a philosophy, it is insurance you can price. And the over-hedging trap is just as real: portability comes from evals, data and workflow definitions, not from running everything twice.
If you want help designing the routing, the fallback paths, or the eval suite that makes all of this honest, our AI consulting engagements do exactly this work with mid-market teams. Either way, run your critical workflows on your secondary model this quarter. You want to learn what breaks while it is a drill.
More from the Blog

The 14 Levels of an AI Rollout, Translated for a Mid-Market Budget
A well-known rollout map runs fourteen levels from first audit to self-guided agent. Here is how each level actually plays out when you have $1M to $50M in revenue, not a data department.
Read article
How to Run an AI Audit That Actually Leads Somewhere
Most "AI audits" are a sixty-page deck and an invoice. Here is the one-week version that ends in a ranked queue of work, not a strategy binder nobody opens.
Read article
The Internal AI Hackathon Playbook
The best AI use case in your business is almost never the one leadership guessed. A two-day internal hackathon drags it into the open, if you run it so the prototypes survive Monday.
Read article