What your model vendor's outage does to your own promises

When the model provider went down, so did the feature, and so did the uptime commitment the company had made to its customers. Their reliability was now someone else's.

What your model vendor’s outage does to your own promises

An engineering director at a company with real enterprise customers told me about a day that reframed how the team thought about their model provider. The provider had an outage. During it, the AI feature stopped working, which the team expected. What they had not fully internalized was that their own availability commitments to customers did not have an asterisk for it. The company had promised a level of uptime, and for the duration of the provider’s problem, they could not meet it, because a core part of the product depended on a service they did not run and could not fix. Their reliability, in that moment, was entirely in someone else’s hands, and there was nothing to do but wait.

When you depend on a model provider, you inherit their availability as a ceiling on your own. Their outage is your outage. Their degraded performance is your degraded performance. And the commitments you have made to your customers do not automatically bend to accommodate a vendor you chose. The provider’s reliability becomes a load-bearing part of your promises, and unless you planned for it to fail, its failure is simply your failure, passed through.

Their uptime is now your ceiling

The uncomfortable math is that your availability can be no better than the availability of everything you critically depend on. Add a model provider to the critical path, and your feature is up only when the provider is up. If you promised your customers a certain reliability, you have implicitly promised that the provider will meet it too, even though you have no control over whether it does.

This gets sharper in the moments that matter. Outages have a way of arriving during high-stakes periods, and a provider outage during a customer’s important moment, or during your own audit, does not care about your commitments. You are left explaining, to a customer who is measuring you against a promise, that the reason you are down is a third party. That explanation rarely satisfies, because from the customer’s side they bought a product from you, not from your vendor.

The reason teams get caught is that in normal operation the provider is reliable enough that the dependency is invisible. The feature works, day after day, so the fact that its uptime is borrowed never surfaces. Then the borrowed uptime is called in, at the worst time, and there is no plan.

Plan for the provider to fail

The fix is to decide, before it happens, what your feature does when the provider is unavailable, and to make sure your promises account for it.

What that involves:

  • Design graceful failure. Decide how the feature behaves when the model is unreachable, so it degrades in a controlled way rather than simply breaking.
  • Consider a fallback path. For critical features, a secondary provider or a reduced-capability mode can turn a total outage into a partial one.
  • Align your commitments with reality. Your availability promises to customers should reflect the dependencies you actually have, so you are not promising an uptime your provider can take away.
  • Have an incident plan for it. Know in advance who does what when the provider goes down, so the response is a runbook rather than an improvisation.

The teams that get this right treat provider availability as a risk they own, because their customers hold them responsible regardless of whose fault it is. They cannot control the outage. They can control what their product does during one, and what they promised in the first place.

How we approach it at Density Labs

In the AI Opportunity Assessment, our fixed two-week, $2,500 engagement, we look at what your feature does when the model provider is unavailable, and whether your customer commitments account for that dependency. Borrowed uptime is invisible until it is called in, and by then the plan needed to already exist. We would rather design the failure behavior in a diagnosis than watch it improvised during an incident.

Your model vendor’s outage becomes your outage, and your customers still hold you to your promises. Plan for the provider to fail, because its reliability is quietly standing in for yours.