Skip to main content
FlexInference is a deadline aware LLM router. You set a time window (5 seconds to 10 minutes for most models) and we find you cheaper inference. We can hunt for cheaper inference across a bunch of different providers. This saves about 48 percent for roughly 20 percent more latency. We also work on default and auto requests. We provide BYOK support for free. We also have trace storage, PII masking, and moderation all behind a single parameter. We are currently working on building an A/B testing platform so you can test out different models. We also support cheap inference search for tools like Codex, OpenClaw, Cursor, and more. Our goal is to make working with AI as easy and affordable as possible.

How does it work?

When a provider offers Flex API support we first check if your request can start in that time using Flex. If it can’t we cancel the Flex request and escalate up to a default request. This means you don’t get charged 1.5x, and always pay either 50% or 100%. We have about a 96 percent hit rate, which is why we can usually save about 48 percent. When providers don’t support Flex, they usually support Batch. With batch inference it can take up to 24 hours. However, that’s not always the case. We have built a model that forecasts compute usage for all these providers. When you send us a request we check the time window you set, model, time of day, and a few other parameters. If based on those our model says that you could get your request within time then we run your request. If not we send out a default request. If we do send that request and it doesn’t fulfill in time then we send out a default request. This offering is only behind Managed Keys. It’s because it requires eating the cost of failed attempts about 4 percent of the time. With managed keys though you never pay that yourself, we take the hit for you. This allows managed keys users to get flex race support far more often and no double charges.

How much do you charge?

For “Bring Your Own Key” we are free. We support Gemini and OpenAI model flex race support. We also have default and priority support. Support across Gemini, OpenAI, Anthropic, and Workers AI open source models. For managed keys we charge 10% and offer Anthropic flex race too. We also have a higher hit rate for flex in managed due to cross provider search. We also provide PII masking, moderation, and trace storage through managed keys. When you add cash through Managed Keys, it isn’t stored as credits. Instead you can refund all your cash back (except the fee).

Quickstart

One working request, in about a minute.

Flex Race Routing

What to set your time limit to, and what happens when it runs out.