Credits, 202 pending and retries: a production-ready API client

How billing, rate limits and errors work, and the handful of client patterns that keep a trading product fast and within budget.

A trading product calls the API constantly, often from several workers. This article covers how billing, rate limits and errors actually behave, and the client patterns that follow from them.

How a request is billed

  1. Before the request runs, its maximum cost is reserved from your balance. For a batch of 100 handles that is 100 × 250 = 25,000 credits.
  2. When it finishes, the real cost is charged: a batch that found 12 traders costs 3,000.
  3. The rest is returned. Anything that is not a 2xx is returned in full.

Two consequences: parallel requests can never overspend your last credits, and a near-empty balance can reject a big batch with 402 even if the real cost would have fitted. Near the end of a period, send smaller batches.

Every response carries x-credits-cost and x-credits-remaining. Log them; they are your usage dashboard.

Status codes and what to do

StatusCodeDo
200Use it. Billed.
202pendingRetry after retryAfterSeconds. Free.
400invalid_*, too_manyFix the request. Do not retry.
401missing_key, invalid_keyStop and alert. Retrying will not help.
402credits_exhaustedStop non-essential work; upgrade or wait for the reset.
403account_inactiveStop. Activate the account from your dashboard, or buy a plan.
404not_found, no_snapshotCache the miss for a while. Free.
429rate_limitedSleep for Retry-After, then retry.
503busy, upstream_unavailableRetry with backoff. Free.

Pattern 1: one retry policy

import random, time

RETRY = {202, 429, 503}

def request(session, url, attempts=5, **kw):
    for i in range(attempts):
        r = session.get(url, timeout=30, **kw)
        if r.status_code not in RETRY:
            return r
        wait = r.headers.get("Retry-After")
        if r.status_code == 202:
            wait = r.json().get("retryAfterSeconds", 15)
        delay = float(wait) if wait else min(60, 2 ** i)
        time.sleep(delay + random.uniform(0, 1))  # jitter so workers do not retry in lockstep
    return r

Pattern 2: batch whenever you can

One batch of 100 is one request against your rate limit instead of 100, and you only pay for what is found. Batches answer from the wallet map and never wait on a background check, so they are also the fastest calls.

Pattern 3: cache by freshness

Wallet ownership changes rarely; PnL changes constantly. Use asOf to decide:

  • Leaderboards: re-fetch when you need a new capture; capturedAt tells you whether anything changed.
  • Wallets: cache for hours. asOf.walletsCheckedAt tells you how old the evidence is.
  • Misses (404): cache for minutes, so a hot unknown address does not hammer the API.

Pattern 4: respect the rate limit per account

Limits are per account, not per key: every key and worker shares one token bucket. Run workers through a shared limiter set a little under your plan’s rate.

PlanRateBurst
Free1/s5
Starter5/s20
Growth20/s60
Scale50/s150

Pattern 5: check your balance for free

GET /v2/me costs nothing and returns your plan, creditsRemaining and periodEnd. Poll it from a health check and alert before you hit 402, not after.