Skip to content
ShivanshSen
Menu

Jaipur, India

All writing

The API is overloaded. What should your client do?

Decide when to retry, spread the attempts out, and put a limit on the extra work.

I had several Claude Code sessions open. They kept returning 529 overloaded_error, then counting down to another attempt. One was already at attempt nine of ten. The waiting made me think about the retry policy I would write for my own client.

Those messages tell me the API reported overload. They do not reveal Claude Code’s retry algorithm or explain what caused the overload. Anthropic distinguishes temporary overload from a rate limit. Its 429 can also mean a spend cap, which waiting will not fix. Claude API error reference.

A retry is another request. Before adding one, I want to know whether it can help, how much extra work it creates, and when it must stop.

First, decide whether to send it again

A malformed request needs a correction. An expired credential needs a new credential. A temporary failure may justify another attempt, subject to the API’s contract and the operation’s safety.

A timeout is harder: I know that my client did not receive the result. I may not know whether the server completed the operation. Retrying a payment or job submission can duplicate work. Use a supported idempotency key for the same logical operation, or check its status before sending it again. Do not assume every POST supports this. HTTP permits automatic retries of non-idempotent operations only when the client knows they are safe or knows the first request was not applied. HTTP idempotency rules.

Exponential backoff still needs some randomness

For a first retry, start with a base delay. Double the delay for each later retry, up to a cap. If the base is 500 milliseconds and the cap is 8 seconds, the delay ceilings are 0.5, 1, 2, 4, 8, then 8 seconds.

ceiling = min(cap, base × 2^retryIndex)
delay = random(0, ceiling)
// retryIndex = 0 for the first retry

This is full jitter: choose a fresh random delay inside that window. Identical delays can leave many clients waking together. Spreading them out can reduce those retry clusters; it does not create server capacity. AWS compares the approaches in Exponential Backoff and Jitter.

If a response includes Retry-After, respect its wait guidance. The value can be seconds or an HTTP date. My policy would wait at least that long, with any added jitter afterward. If the wait exceeds the operation’s remaining deadline, stop instead of shortening the server’s requested wait. HTTP Retry-After.

Compare the extra work

Give the model below the same burst of clients and change their retry policy. Compare both completed requests and extra attempts. A policy that sends fewer requests by giving up early has made a different trade-off.

Compare retry policies

Compare the work each policy sends and the clients it completes. The same client count and server capacity apply to every row.

Default run: 20 clients, 6 accepted requests per second after recovery, no shared throttle. Immediate retry: 0 completed, 140 retries, 20 gave up. Capped exponential: 18 completed, 82 retries, 2 gave up. Exponential + full jitter: 20 completed, 80 retries, 0 gave up.

Completed work and request load
PolicyCompletedAttemptsGave up
Immediate retry0 / 20160 (140 retries)20
Capped exponential18 / 20102 (82 retries)2
Exponential + full jitter20 / 20100 (80 retries)0
Model assumptions
  • Every client starts at time zero. The server rejects all requests for two seconds, then accepts up to the selected capacity in each fixed one-second window.
  • Each response takes 0.1 seconds. Immediate retry waits another 0.1 seconds; exponential backoff waits 0.5, 1, 2, 4, then at most 8 seconds. Full jitter draws a repeatable wait from zero to that same ceiling.
  • Each client gets at most eight attempts and a 20-second deadline. A shared throttle spaces initial and retry starts at six requests per second; time in its queue counts against the deadline.
  • Retries stop after success. The comparison reports both completions and clients that give up, so a policy cannot look better by silently dropping unfinished work.

This is a seeded educational simulation. It makes no API calls and does not model Anthropic’s infrastructure. Jitter spreads retries; it does not guarantee a better outcome for every run.

Throttle the rate and limit concurrency

Throttling limits how quickly the client starts requests, including initial attempts. A concurrency limit bounds requests already in flight. With slow responses, a modest start rate can still leave many requests running at once.

Share these controls across the work that uses the same upstream limit. Ten independent workers each enforcing their own allowance can still exceed the total. Bound the waiting queue too; delaying an unbounded backlog moves the problem into memory.

Give each operation a maximum attempt count and an overall deadline that includes requests and sleeps. Let the user cancel both. Across operations, a retry budget limits how much additional traffic retries may consume. During a sustained outage, an attempt limit on every request can still produce a large retry load. AWS guidance on retries and overload.

Check the SDK before adding a retry loop

Before adding a wrapper, inspect your SDK. Anthropic’s official SDK retries transient failures twice by default and honors Retry-After. That describes the SDK, not the Claude Code sessions above. SDK retry behaviour.

If an outer loop makes three attempts and each SDK call makes three attempts, one operation can send nine requests. Choose the layer that owns the policy and keep its deadline visible to the caller.

I would record attempts per operation, elapsed time, error type, server wait guidance, and the reason the client stopped. Then I can tell whether retries recovered useful work or just prolonged a failure. Keep secrets and request bodies out of those logs.

The API owner also needs admission control, bounded queues, load shedding, and graceful degradation. That side needs its own article.

All writing