Skip to content
Skip to main content

Handle 429 rate limit errors

Last updated:

Your batch job is humming along, then a wall of 429 Too Many Requests responses stops it cold. You’re hitting a rate limit, either one the API enforces or one the underlying provider enforces. A 429 isn’t a failure you log and forget. It’s a signal to slow down, wait the right amount of time, and retry.

This recipe is the handling playbook: how to detect a 429, how to read the Retry-After header, and how to back off with exponential delay and jitter so your retries don’t stampede. For the limit numbers themselves, see API and provider rate limits.

A 429 response means you sent too many requests in a time window, and the server is refusing more until you slow down. The API returns a consistent JSON error object: error.type names the category (for example, rate_limit_error for a throttled read), message summarizes the cause, and provider_error contains the provider’s own error, if it returned one. The same status code covers several causes: an account throttled by its provider, a Gmail quota exhausted, a Microsoft mailbox over its concurrency limit, or too many calls to the API itself.

The error body tells you which cause you hit. The provider 429 table maps each provider’s message and Retry-After behavior. The same page documents ten distinct 429 variants, such as Account throttled and Application is over its MailboxConcurrency limit. You don’t need to branch on every variant. Treat them all as “wait, then retry”, and read the Retry-After header when it’s present.

The cause matters for how long you wait, not for whether you retry. A Gmail Resource exhausted usually clears quickly. Google measures its per-user quota over a one-minute window, so a short backoff is enough. An Exchange account throttled response can last much longer. The Exchange server decides when to accept requests again, so retrying every few seconds wastes attempts. Read the body to set expectations, then let Retry-After and your backoff cap decide the actual delay.

For the full list of causes and fixes, see the client error reference for 400-499.

When a provider throttles a read request and says how long to wait, the API includes a Retry-After header with that number of seconds, rounded up so it’s never shorter than the provider asked. This applies to every Email, Calendar, and Contacts endpoint except send, for Google, Microsoft, Exchange, iCloud, and Zoom grants. Honor that value instead of guessing a delay. The Nylas SDKs don’t retry 429 responses or read Retry-After for you, so the retry logic below belongs in your application whichever SDK you use.

Not every 429 includes the header. IMAP servers don’t publish a wait time, so an IMAP rate limit hit. response never has one. When the header is missing, on a read or a send, fall back to your own exponential backoff schedule.

Retry-After is the provider telling you precisely how long it needs. Ignoring it and retrying early just earns another 429 and can extend the throttle. The Node and Python examples below check for the header first, parse it as an integer number of seconds, and only compute a backoff delay when it’s missing.

Microsoft allows up to 4 concurrent requests for each app ID and mailbox pair, so a 429 here often means too much parallelism rather than too many total calls. A single send makes several Graph calls, so 2 or 3 parallel sends to one grant can fill all 4 slots. See Microsoft Graph rate limits for Outlook mail for the MailboxConcurrency explanation and a per-grant send queue.

For EWS, an on-premises Exchange administrator sets the throttle, so the API can’t know the limit in advance. When the server throttles a read and says how long to wait, the response includes a Retry-After header. Two response headers let you act before a 429 ever fires. Nylas-Provider-Request-Count reports how many provider calls a single request consumed, for any provider, and Nylas-Gmail-Quota-Usage reports how many Gmail quota units a Drafts, Messages, Threads, Folders, or Attachments request spent. Read both on successful responses, not just failures, so you can taper your request rate while you still have headroom rather than after the provider cuts you off.

How do I retry with exponential backoff and jitter?

Section titled “How do I retry with exponential backoff and jitter?”

Exponential backoff doubles the wait after each failed attempt: roughly 1, 2, 4, 8 seconds. Jitter adds a random offset so many clients retrying at once don’t all fire on the same tick and re-trigger the limit. Cap the delay (for example, at 32 seconds) and cap retries (5 attempts is a sensible default) so a persistent 429 doesn’t loop forever.

Without jitter, a fleet of workers that all got throttled at the same moment retries in lockstep and stampedes the API again. The functions below add randomness of up to 1 second on top of the exponential base, honor Retry-After when present, and give up after 5 attempts. Both target the same GET /v3/grants/{grant_id}/messages endpoint that returns up to 200 messages per page.

The official Nylas SDKs raise exceptions on non-200 responses, so wrap SDK calls in the same retry loop. For broader error strategy beyond rate limits, see error monitoring and handling.

How do provider quotas differ from Nylas limits?

Section titled “How do provider quotas differ from Nylas limits?”

Two separate ceilings apply to every request. The API enforces its own limits, such as up to 200 requests per grant per second on the Messages endpoint. Underneath, each provider enforces its own quota, and a single API call can fan out to several provider calls. You can hit a provider limit while staying well under the API rate limit.

Google and Microsoft count usage differently, which changes how you pace requests:

ProviderQuota modelKey limitHeader to watch
GooglePer-user and per-project6,000 quota units/min per user; 1.2M/min per projectNylas-Gmail-Quota-Usage
MicrosoftPer app ID and mailbox10,000 requests/10 min; max 4 concurrentNylas-Provider-Request-Count

Because Microsoft counts per app ID and mailbox and Google counts per user and per project, a noisy single account throttles differently on each. Watch the Nylas-Provider-Request-Count header to see how many provider calls a request consumed, and space requests out accordingly.

Gmail’s quota math is the part that surprises people. The 6,000 quota units per minute per user isn’t 6,000 requests: each method spends a different number of units, so a fast loop of cheap calls and a slow loop of expensive ones can hit the same ceiling. A messages.list page costs 5 units, a messages.get body read costs 20, and a messages.send costs 100, so reading bodies costs 4x listing them and sending costs 20x. A single limit=200 list request can spend far more units than one messages.list call. Read the Nylas-Gmail-Quota-Usage header to see what each request actually cost. For the full per-method cost table and a sync-budget worksheet, see Gmail API quotas and limits.

How do I queue requests to stay under the limit?

Section titled “How do I queue requests to stay under the limit?”

Queue work through a rate limiter that caps outbound requests below the ceiling instead of firing them as fast as your code generates them. A token-bucket or fixed-window limiter set to, for example, 150 requests per second per grant keeps you under the 200-per-second Messages limit with headroom for retries. Pair the queue with the backoff loop so throttled jobs return to the queue cleanly.

Queueing turns reactive retrying into proactive pacing. For Microsoft, keep in-flight requests at or below 4 per mailbox to respect the concurrency cap. For Google, throttle per user since the 6,000-quota-units-per-minute limit is per-user, not per-application. A single shared queue per grant gives you one place to tune both the rate and the concurrency.

Combine the limiter with listMessagesWithRetry above: call acquire() before each request so you pace proactively and still recover from any 429 that slips through.

Should I throttle proactively or back off after a 429?

Section titled “Should I throttle proactively or back off after a 429?”

Use both, because they solve different problems. Proactive throttling caps your outbound rate below the ceiling so most requests never hit a limit, which keeps latency steady. Reactive backoff is the recovery net for the requests that slip through anyway: a quota you share with other traffic, a provider-side spike, or a burst your limiter didn’t predict. A limiter set to 150 requests per second under the 200 requests per second Messages limit handles the first; the backoff loop, capped at 5 attempts, handles the second.

Lean on proactive pacing for steady, predictable workloads where you control the request rate, such as a nightly sync or a polling job. Lean on reactive backoff for traffic you can’t shape, such as user-triggered actions that arrive in unpredictable bursts. The cost of getting the balance wrong is asymmetric. Throttle too aggressively and a backfill that should take an hour takes three; throttle too loosely and you trade a few saved seconds for repeated 429 responses that each cost a full Retry-After wait, often longer than the request you were trying to save. When in doubt, set the limiter conservatively and let backoff absorb the rare miss.

How do I backfill 50 accounts without tripping limits?

Section titled “How do I backfill 50 accounts without tripping limits?”

Backfill per grant, not per fleet. Each Messages limit is scoped to one grant at up to 200 requests per second, so 50 grants give you 50 independent budgets rather than one shared pool. Run a bounded number of grants in parallel, each with its own per-grant limiter set near 150 requests per second, and the fleet stays under every ceiling at once. The constraint that bites first is almost always a provider quota, not the API limit.

Concretely: spin up a worker pool, assign each grant its own token bucket and backoff loop, and cap concurrency at the slowest provider’s tolerance. For Microsoft mailboxes, hold in-flight requests at or below 4 per mailbox, since the MailboxConcurrency limit rejects a fifth simultaneous request outright. For Gmail accounts, pace per user under 6,000 quota units per minute, and remember a messages.get-heavy backfill spends units 4x faster than a list-only pass. Read Nylas-Provider-Request-Count on each response to learn the real fan-out per request, then tune each worker’s rate down until the count stops climbing. A backfill that respects per-grant budgets and honors Retry-After keeps each account well clear of provider throttling.