# The API is overloaded. What should your client do?

Canonical: https://shivanshsen.com/blogs/api-overloaded-client-retries

Decide when to retry, spread the attempts out, and put a limit on the extra work.

![Travellers and camels queue at a narrow gateway, then continue with more space between them.](https://shivanshsen.com/illustrations/articles/api-overloaded-client-retries-960.webp)

Pacing arrivals gives a busy gateway room to recover.

AI-generated contemporary Phad-inspired illustration; not traditional artisan authorship.

I had several Claude Code sessions open. They kept returning `529 overloaded_error`, then counting down to another attempt. One was already at attempt nine of ten. The waiting made me think about the retry policy I would write for my own client.

Those messages tell me the API reported overload. They do not reveal Claude Code’s retry algorithm or explain what caused the overload. Anthropic distinguishes temporary overload from a rate limit. Its `429` can also mean a spend cap, which waiting will not fix. [Claude API error reference](https://platform.claude.com/docs/en/api/errors).

A retry is another request. Before adding one, I want to know whether it can help, how much extra work it creates, and when it must stop.

## First, decide whether to send it again

A malformed request needs a correction. An expired credential needs a new credential. A temporary failure may justify another attempt, subject to the API’s contract and the operation’s safety.

A timeout is harder: I know that my client did not receive the result. I may not know whether the server completed the operation. Retrying a payment or job submission can duplicate work. Use a supported idempotency key for the same logical operation, or check its status before sending it again. Do not assume every `POST` supports this. HTTP permits automatic retries of non-idempotent operations only when the client knows they are safe or knows the first request was not applied. [HTTP idempotency rules](https://www.rfc-editor.org/rfc/rfc9110.html#section-9.2.2).

## Exponential backoff still needs some randomness

For a first retry, start with a base delay. Double the delay for each later retry, up to a cap. If the base is 500 milliseconds and the cap is 8 seconds, the delay ceilings are 0.5, 1, 2, 4, 8, then 8 seconds.

```text
ceiling = min(cap, base × 2^retryIndex)
delay = random(0, ceiling)
// retryIndex = 0 for the first retry
```

This is full jitter: choose a fresh random delay inside that window. Identical delays can leave many clients waking together. Spreading them out can reduce those retry clusters; it does not create server capacity. AWS compares the approaches in [Exponential Backoff and Jitter](https://aws.amazon.com/blogs/architecture/exponential-backoff-and-jitter/).

If a response includes `Retry-After`, respect its wait guidance. The value can be seconds or an HTTP date. My policy would wait at least that long, with any added jitter afterward. If the wait exceeds the operation’s remaining deadline, stop instead of shortening the server’s requested wait. [HTTP Retry-After](https://www.rfc-editor.org/rfc/rfc9110.html#section-10.2.3).

## Compare the extra work

Give the model below the same burst of clients and change their retry policy. Compare both completed requests and extra attempts. A policy that sends fewer requests by giving up early has made a different trade-off.

## Throttle the rate and limit concurrency

**Throttling** limits how quickly the client starts requests, including initial attempts. **A concurrency limit** bounds requests already in flight. With slow responses, a modest start rate can still leave many requests running at once.

Share these controls across the work that uses the same upstream limit. Ten independent workers each enforcing their own allowance can still exceed the total. Bound the waiting queue too; delaying an unbounded backlog moves the problem into memory.

Give each operation a maximum attempt count and an overall deadline that includes requests and sleeps. Let the user cancel both. Across operations, a retry budget limits how much additional traffic retries may consume. During a sustained outage, an attempt limit on every request can still produce a large retry load. [AWS guidance on retries and overload](https://aws.amazon.com/builders-library/timeouts-retries-and-backoff-with-jitter/).

## Check the SDK before adding a retry loop

Before adding a wrapper, inspect your SDK. Anthropic’s official SDK retries transient failures twice by default and honors `Retry-After`. That describes the SDK, not the Claude Code sessions above. [SDK retry behaviour](https://platform.claude.com/docs/en/api/errors).

If an outer loop makes three attempts and each SDK call makes three attempts, one operation can send nine requests. Choose the layer that owns the policy and keep its deadline visible to the caller.

I would record attempts per operation, elapsed time, error type, server wait guidance, and the reason the client stopped. Then I can tell whether retries recovered useful work or just prolonged a failure. Keep secrets and request bodies out of those logs.

The API owner also needs admission control, bounded queues, load shedding, and graceful degradation. That side needs its own article.

## Sources

- https://platform.claude.com/docs/en/api/errors

- https://www.rfc-editor.org/rfc/rfc9110.html#section-9.2.2

- https://aws.amazon.com/blogs/architecture/exponential-backoff-and-jitter/

- https://www.rfc-editor.org/rfc/rfc9110.html#section-10.2.3

- https://aws.amazon.com/builders-library/timeouts-retries-and-backoff-with-jitter/
