provider array. The first route is the primary and the only one that can run a Flex Race. The rest are backups for the same model. We try them in order when a route fails.
gpt-*ando*run onopenai, withfoundryon Azure AI Foundry as the backup.gemini-*runs ongoogle, withvertexon Google Vertex AI as the backup.claude-*runs onanthropic, withbedrockon Amazon Bedrock as the backup.- The open-source families run on
cloudflareordeepinfra. Most run on one of them. Two run on both.
429, a 5xx, a connection error, or a timeout, and only before the response commits. A chain that used up its routes on faults returns 502 all_routes_failed. When the last failure was a 429, you get that 429 back instead.
Every route you name needs its own key. We refuse a missing key before any upstream call, and we never fall through to the next route, because that would put you on a provider you didn’t choose.
Claude on Bedrock
We address every Claude model on Bedrock through a global inference profile, so it can run in any AWS region. A request is not guaranteed to run in the region your Bedrock key names. Bedrock can carry a model and still refuse it for your AWS account, which returns403 bedrock_model_access_denied. Grant the model in the Amazon Bedrock console under Model access.
The five newest Claude models need two more one-time steps on the AWS account: claude-opus-5, claude-sonnet-5, claude-fable-5, claude-opus-4-8, and claude-opus-4-7. Run both with an administrator’s credentials. The first command opts the whole AWS account into provider data sharing, Amazon requires it for these models, and it stays on until you turn it off. The Bedrock inference key you added in the dashboard can’t run them.
403 bedrock_model_access_denied. Amazon requires this, not us.
GPT on Foundry
Thefoundry route runs GPT through Azure AI Foundry instant access, with no Azure deployment step. Add your resource name and one of its API keys under API, then Provider keys. Create the resource in West US 3, the one instant-access region during the preview.
Azure decides which part of the GPT catalog instant access carries. A model it lacks returns 404 foundry_model_unavailable.
Open source on DeepInfra
Thedeepinfra route runs open-source models. Add your DeepInfra token under API, then Provider keys.
DeepInfra sells a flex tier. We don’t race it.
The route never tells us when it has read your prompt but hasn’t started writing yet. So a race we lose has already paid to read your prompt, then pays again in full. That costs more than a standard call, so we send you standard.
gemma-4-26b-a4b-it and llama-4-scout-17b-16e-instruct run on both cloudflare and deepinfra. They go to Workers AI unless you name deepinfra first.