countis required, from 1 to 5. It is the number of extra attempts after the first one fails.backoffisexponentialby default, orlinear.jitteris on by default. It spreads the waits apart, and only ever shortens them.
429, a 5xx, or a connection that never opened, and only before anything commits. We do not retry your own 4xx, a bad provider key, a timeout, or a request you cancelled, because those fail the same way twice.
A retry goes back to whichever tier your request ended on, so a lost race retries the default tier. The race itself never counts against count. When the provider sends Retry-After we honor it, up to 60 seconds.
Your client probably retries too, and the two multiply. A stock OpenAI or Anthropic client retries twice, which is three sends. Each send spends its own budget, so count: 3 makes each one up to four attempts, and the request can reach twelve. We set x-flexinference-retries to the number we ran, so you can turn your client down.
An agent key can carry a retry policy the same way it carries a duration. See setting the duration for Agent Keys.