backlink api rate limits: safe retries

Handle monthly quotas, Retry-After headers and ambiguous timeouts with a bounded Python retry planner that avoids duplicate backlink jobs.

crawlgraph team
· 9 min read · 1,696 words
sharexlinkedin

Handle backlink API rate limits by identifying the exhausted allowance before deciding to retry. A monthly quota needs a scheduling or capacity decision; a temporary burst limit may allow a bounded wait. An ambiguous timeout on a cost-bearing submission needs reconciliation. Start with the backlink API guide for endpoint selection, or the free backlink research guide if a manual lookup will answer the question without building an integration.

the four-step workflow

  1. classify the endpoint Record its method, quota bucket and whether repeating it can spend quota or create another job. Default unknown endpoints to no automatic retry.
  2. inspect the response Read the HTTP status, error code, request ID and quota headers. Separate monthly exhaustion from a temporary burst limit.
  3. calculate a bounded delay Parse Retry-After as seconds or an HTTP date. Only plan another confirmed safe read when both the attempt count and total wait budget permit it.
  4. save or reconcile the outcome Persist successful results. Stop ambiguous cost-bearing submissions and investigate existing job or usage records before considering another submission.

The goal is to finish useful work within the allowance and explain any remainder. Blind retries can spend calls twice or leave accepted jobs unclear. Use each response to choose the next action; do not retry by status code alone.

Before sending a batch, write down what a unit of work means. A referring-domain lookup, a release comparison and an asynchronous gap submission are different operations even when the client groups them under one report. Keep those operations in separate queues if they use different monthly buckets. This makes a stopped queue understandable without blocking unrelated work that still has capacity.

monthly quota versus burst limits

A burst limit restricts requests over a short interval. A monthly quota restricts the total allowance for a billing or calendar period. Waiting a few seconds can help with the first; it does not create new monthly capacity for the second. RFC 6585 defines HTTP 429 as too many requests and allows a Retry-After header. It leaves the counting policy to the server, so 429 alone cannot tell your client which limit it reached.

As checked on 3 October 2026, the crawlgraph API documentation lists 15 backlink calls and zero gap calls for the free tier, and 1,000 backlink calls and 50 gap calls for paid access, per UTC calendar month. Release-change queries use the backlink bucket. A GET method therefore does not imply that an operation is free to repeat. A repeated successful query can spend another call even when the domain and release pair are unchanged.

budget for accepted attempts

The implementation reviewed on 3 October 2026 charges quota before query or job work. A later failure can consume a call. Do not assume a failed request is refunded.

Inspect X-RateLimit-Limit-Backlinks, X-RateLimit-Remaining-Backlinks, and the corresponding -Gap headers for the operation you actually sent. X-RateLimit-Reset identifies the next UTC month boundary as a Unix timestamp. A quota_exceeded error is a stop condition for that bucket. Store the reset time for scheduling; do not leave a foreground process sleeping until next month.

a status-aware decision table

This table is our conservative client decision aid, rather than a universal provider contract. Its retry rows apply only to an endpoint you have confirmed is a non-quota, safe read. For crawlgraph backlink queries, changes and gap submissions, preserve the result or investigate the failure before repeating the operation.

observed outcomeinterpretationnext client action
2xx responsethe operation returned successfullypersist the response; do not query again for the same report
429 + quota_exceededmonthly bucket exhaustedstop that queue; record reset and remaining work
429 + matching remaining = 0quota exhausted or insufficient contextstop and inspect the error contract
429 without monthly evidencepossibly a temporary request limitonly a confirmed safe non-quota read may use a bounded delay
502, 503 or 504transient infrastructure failure is possibleapply the same safe-read gate and retry budgets
400, 401 or 403input, authentication or access needs correctioninspect and fix; repeating identical input does not help
timeout with no responseserver acceptance is unknownstop automatic repetition; reconcile the original operation

Do not infer safety from the absence of headers. A proxy may return an HTML error page without the API envelope. Preserve that distinction as an unknown failure, rather than manufacturing a quota value. Likewise, a zero gap allowance on a free account is a plan constraint, not evidence that a one-second backoff will make a gap request eligible.

a bounded Python retry planner

The following standard-library example makes no HTTP requests and needs no API key. It plans the next action from supplied response data. The included assertions are offline fixtures, not live provider results. Its three-attempt and twenty-second limits are illustrative application choices; they are not crawlgraph limits or a recommended workload size.

python
from datetime import datetime, timezone
from email.utils import parsedate_to_datetime
import math

MAX_ATTEMPTS = 3
WAIT_BUDGET = 20.0

def retry_after(value, now):
    if value is None:
        return None
    value = value.strip()
    if value.isascii() and value.isdigit():
        try:
            seconds = float(int(value))
            return seconds if math.isfinite(seconds) else None
        except (ValueError, OverflowError):
            return None
    try:
        date = parsedate_to_datetime(value)
        if date.tzinfo is None:
            return None
        return max(0.0, (date - now).total_seconds())
    except (ValueError, TypeError, OverflowError):
        return None

def plan(method, status, headers, body, *, safe_read=False,
         attempt=1, waited=0.0, now=None, bucket="Backlinks"):
    # safe_read must come from the endpoint contract, never its method alone.
    now = now or datetime.now(timezone.utc)
    h = {key.lower(): str(value) for key, value in headers.items()}
    remaining = h.get("x-ratelimit-remaining-" + bucket.lower())
    if body.get("error") == "quota_exceeded":
        return ("stop: monthly quota", None)
    if status == 429 and remaining == "0":
        return ("stop: quota or unknown limit", None)
    if status is not None and 200 <= status < 300:
        return ("done: persist response", None)
    if method.upper() != "GET" or not safe_read:
        return ("stop: reconcile without resubmitting", None)
    if status not in (429, 502, 503, 504):
        return ("stop: inspect failure", None)
    if attempt >= MAX_ATTEMPTS or waited < 0 or waited > WAIT_BUDGET:
        return ("stop: retry budget", None)
    delay = retry_after(h.get("retry-after"), now)
    if delay is None:
        delay = float(2 ** (attempt - 1))
    # A past date must not create a tight retry loop.
    delay = max(1.0, delay)
    if delay > WAIT_BUDGET - waited:
        return ("stop: wait budget", None)
    return ("retry: confirmed non-quota read", delay)

now = datetime(2026, 10, 3, tzinfo=timezone.utc)
assert retry_after("3", now) == 3.0
assert retry_after("Sat, 03 Oct 2026 00:00:05 GMT", now) == 5.0
assert retry_after("Fri, 02 Oct 2026 23:59:00 GMT", now) == 0.0
assert plan("GET", 429, {}, {"error": "quota_exceeded"},
            safe_read=True, now=now)[1] is None
assert plan("GET", 429, {"X-RateLimit-Remaining-Backlinks": "0"},
            {}, safe_read=True, now=now)[1] is None
assert plan("GET", 503, {}, {}, safe_read=True, now=now)[1] == 1.0
assert plan("GET", 503, {"Retry-After": "30"}, {},
            safe_read=True, now=now)[1] is None
assert plan("GET", 503, {"Retry-After": "3"}, {},
            safe_read=True, waited=19, now=now)[1] is None
assert plan("GET", 503, {}, {}, safe_read=True,
            attempt=3, now=now)[1] is None
assert plan("POST", None, {}, {}, safe_read=True, now=now)[1] is None
assert plan("GET", None, {}, {}, safe_read=True, now=now)[1] is None
assert plan("GET", 503, {}, {}, now=now)[1] is None
assert plan("GET", 200, {}, {}, now=now)[0].startswith("done")

RFC 9110 describes both Retry-After formats: delay-seconds and an HTTP date. The parser accepts an integer delay or calculates the time until a date using an aware UTC clock. A past date becomes zero, then the planner applies a one-second minimum. Missing or malformed values use a small bounded fallback only after the endpoint passes the safe-read gate.

The caller starts at attempt one, including the initial request. If the planner permits another attempt, add its delay to the cumulative waited value and increment attempt before evaluating the next response. Never reset those counters inside an inner exception handler. The total wait budget covers retry delays; separately bound network connection and response timeouts so a slow request cannot hold the worker indefinitely.

Set safe_read=True only from a reviewed endpoint contract that confirms both safe repetition and absence of per-call quota consequences. Do not set it for crawlgraph's GET changes endpoint merely because it is a GET. This deliberately cautious sample also stops transport failures with no status. If your integration later adds a narrower recovery path, establish what happened to the original request before enabling that path.

reconcile ambiguous submissions

A timeout tells you that the client did not receive a complete response within its deadline. It does not tell you that the server rejected the request. The server may have consumed quota, started a query or accepted a job before the connection failed. RFC 9110 cautions against automatic retries of non-idempotent requests unless the client knows repetition is safe or knows the original was not applied.

crawlgraph does not document a public Idempotency-Key contract. Adding a randomly generated header does not establish deduplication. Do not automatically repeat POST backlink lookups or asynchronous gap submissions after an ambiguous timeout. If a job identifier was received, retain it and use documented status retrieval. If no identifier was received, preserve the timestamp, endpoint and available request ID for investigation rather than guessing whether a replacement is needed.

An explicit error is more information than a lost connection, but still requires interpretation. A server error after charging is not a promise that another submission will be free. A local deduplication record can prevent your own workers from sending duplicates; it cannot retroactively prove what the remote service accepted. Use a pending-unknown state until the evidence supports completion, failure or a deliberate new submission.

plan capacity before the batch

Count unique operations, reserve room for investigation, and reuse saved successful responses. For an illustrative batch of 40 distinct backlink queries, a free monthly allowance of 15 cannot cover the batch. Splitting it into faster concurrent workers changes neither the arithmetic nor the reset time. Choose a smaller research set, schedule the remaining work for a later month, or review the current paid API access plans when the larger recurring workload justifies them.

Do not confuse a locally cached result with a provider's internal query cache. Your cache can avoid sending another request. A provider cache can accelerate processing while still counting the request against its allowance. Store the domain, parameters, named release and result together, so reuse is intentional and a report does not silently mix different snapshots.

For multiple workers, coordinate admission to each quota bucket in one shared ledger. A remaining count observed by one worker may already be stale when another worker sends a request. Treat response headers as evidence from that response, not a reservation of future capacity. Keep concurrency modest enough that an exhaustion response can stop the queue before a large group of already-dispatched requests makes the final usage hard to explain.

keep an operational evidence trail

Log the request time, operation type, response status, provider request ID, matching quota bucket, remaining count, reset timestamp and chosen decision. Avoid recording keys or authorization headers. For ambiguous outcomes, keep a local operation identifier and the normalized input needed to investigate the job. These records should answer why work stopped without exposing a customer's entire report.

Separate completed, rejected, unknown and deferred operations in the batch summary. A 429 should not erase the successful work already saved. Report what remains and why: monthly exhaustion, invalid input, access restriction or an unresolved submission. The useful completion criterion is a reproducible result set plus an honest remainder, rather than a retry loop that eventually exits with no explanation.

faq

does every 429 mean the monthly quota is exhausted?

No. HTTP 429 means too many requests. A provider may use it for a burst limit or a monthly allowance. In crawlgraph, the quota_exceeded error identifies monthly exhaustion; inspect the matching quota headers and reset time.

can Retry-After be a date?

Yes. Retry-After can contain delay-seconds or an HTTP date. Calculate the date delay using UTC, clamp past dates to zero, and apply a minimum delay and a total wait budget before any permitted retry.

should a timed-out backlink POST be retried?

Not automatically. A timeout does not establish whether the server charged the call or accepted a job. crawlgraph has no documented public Idempotency-Key contract. Reconcile the outcome before submitting again.

does a cached or failed request always preserve quota?

No. Repeated successful quota-bearing lookups can consume another call. The implementation charges before query or job work, so a failure after charging can also consume quota. Save results locally and inspect usage rather than assuming failed attempts were free.

ahrefs · backlinkslocked
upgrade required · $129/mo
crawlgraph · live $99 once
G
github.io92
C
css-tricks.com88
L
lobste.rs86
A
algolia.com84
W
web.dev80
same data · one-time
$99oncevs $129/mo
unlock the data →
stripe checkout · instant access
guides#backlink api#rate limits#developer workflow
sharexlinkedin
crawlgraph team
author

documents crawlgraph backlink workflows, open-data methods, and the limits of each report.

the dispatch
one email a month.

plus one when a new common crawl release lands. that is all.