Handle backlink API rate limits by identifying the exhausted allowance before deciding to retry. A monthly quota needs a scheduling or capacity decision; a temporary burst limit may allow a bounded wait. An ambiguous timeout on a cost-bearing submission needs reconciliation. Start with the backlink API guide for endpoint selection, or the free backlink research guide if a manual lookup will answer the question without building an integration.
the four-step workflow
- classify the endpoint Record its method, quota bucket and whether repeating it can spend quota or create another job. Default unknown endpoints to no automatic retry.
- inspect the response Read the HTTP status, error code, request ID and quota headers. Separate monthly exhaustion from a temporary burst limit.
- calculate a bounded delay Parse Retry-After as seconds or an HTTP date. Only plan another confirmed safe read when both the attempt count and total wait budget permit it.
- save or reconcile the outcome Persist successful results. Stop ambiguous cost-bearing submissions and investigate existing job or usage records before considering another submission.
The goal is to finish useful work within the allowance and explain any remainder. Blind retries can spend calls twice or leave accepted jobs unclear. Use each response to choose the next action; do not retry by status code alone.
Before sending a batch, write down what a unit of work means. A referring-domain lookup, a release comparison and an asynchronous gap submission are different operations even when the client groups them under one report. Keep those operations in separate queues if they use different monthly buckets. This makes a stopped queue understandable without blocking unrelated work that still has capacity.
monthly quota versus burst limits
A burst limit restricts requests over a short interval. A monthly quota restricts the total allowance for a billing or calendar period. Waiting a few seconds can help with the first; it does not create new monthly capacity for the second. RFC 6585 defines HTTP 429 as too many requests and allows a Retry-After header. It leaves the counting policy to the server, so 429 alone cannot tell your client which limit it reached.
As checked on 3 October 2026, the crawlgraph API documentation lists 15 backlink calls and zero gap calls for the free tier, and 1,000 backlink calls and 50 gap calls for paid access, per UTC calendar month. Release-change queries use the backlink bucket. A GET method therefore does not imply that an operation is free to repeat. A repeated successful query can spend another call even when the domain and release pair are unchanged.
The implementation reviewed on 3 October 2026 charges quota before query or job work. A later failure can consume a call. Do not assume a failed request is refunded.
Inspect X-RateLimit-Limit-Backlinks, X-RateLimit-Remaining-Backlinks, and the corresponding -Gap headers for the operation you actually sent. X-RateLimit-Reset identifies the next UTC month boundary as a Unix timestamp. A quota_exceeded error is a stop condition for that bucket. Store the reset time for scheduling; do not leave a foreground process sleeping until next month.
a status-aware decision table
This table is our conservative client decision aid, rather than a universal provider contract. Its retry rows apply only to an endpoint you have confirmed is a non-quota, safe read. For crawlgraph backlink queries, changes and gap submissions, preserve the result or investigate the failure before repeating the operation.
| observed outcome | interpretation | next client action |
|---|---|---|
| 2xx response | the operation returned successfully | persist the response; do not query again for the same report |
| 429 + quota_exceeded | monthly bucket exhausted | stop that queue; record reset and remaining work |
| 429 + matching remaining = 0 | quota exhausted or insufficient context | stop and inspect the error contract |
| 429 without monthly evidence | possibly a temporary request limit | only a confirmed safe non-quota read may use a bounded delay |
| 502, 503 or 504 | transient infrastructure failure is possible | apply the same safe-read gate and retry budgets |
| 400, 401 or 403 | input, authentication or access needs correction | inspect and fix; repeating identical input does not help |
| timeout with no response | server acceptance is unknown | stop automatic repetition; reconcile the original operation |
Do not infer safety from the absence of headers. A proxy may return an HTML error page without the API envelope. Preserve that distinction as an unknown failure, rather than manufacturing a quota value. Likewise, a zero gap allowance on a free account is a plan constraint, not evidence that a one-second backoff will make a gap request eligible.
a bounded Python retry planner
The following standard-library example makes no HTTP requests and needs no API key. It plans the next action from supplied response data. The included assertions are offline fixtures, not live provider results. Its three-attempt and twenty-second limits are illustrative application choices; they are not crawlgraph limits or a recommended workload size.
from datetime import datetime, timezone
from email.utils import parsedate_to_datetime
import math
MAX_ATTEMPTS = 3
WAIT_BUDGET = 20.0
def retry_after(value, now):
if value is None:
return None
value = value.strip()
if value.isascii() and value.isdigit():
try:
seconds = float(int(value))
return seconds if math.isfinite(seconds) else None
except (ValueError, OverflowError):
return None
try:
date = parsedate_to_datetime(value)
if date.tzinfo is None:
return None
return max(0.0, (date - now).total_seconds())
except (ValueError, TypeError, OverflowError):
return None
def plan(method, status, headers, body, *, safe_read=False,
attempt=1, waited=0.0, now=None, bucket="Backlinks"):
# safe_read must come from the endpoint contract, never its method alone.
now = now or datetime.now(timezone.utc)
h = {key.lower(): str(value) for key, value in headers.items()}
remaining = h.get("x-ratelimit-remaining-" + bucket.lower())
if body.get("error") == "quota_exceeded":
return ("stop: monthly quota", None)
if status == 429 and remaining == "0":
return ("stop: quota or unknown limit", None)
if status is not None and 200 <= status < 300:
return ("done: persist response", None)
if method.upper() != "GET" or not safe_read:
return ("stop: reconcile without resubmitting", None)
if status not in (429, 502, 503, 504):
return ("stop: inspect failure", None)
if attempt >= MAX_ATTEMPTS or waited < 0 or waited > WAIT_BUDGET:
return ("stop: retry budget", None)
delay = retry_after(h.get("retry-after"), now)
if delay is None:
delay = float(2 ** (attempt - 1))
# A past date must not create a tight retry loop.
delay = max(1.0, delay)
if delay > WAIT_BUDGET - waited:
return ("stop: wait budget", None)
return ("retry: confirmed non-quota read", delay)
now = datetime(2026, 10, 3, tzinfo=timezone.utc)
assert retry_after("3", now) == 3.0
assert retry_after("Sat, 03 Oct 2026 00:00:05 GMT", now) == 5.0
assert retry_after("Fri, 02 Oct 2026 23:59:00 GMT", now) == 0.0
assert plan("GET", 429, {}, {"error": "quota_exceeded"},
safe_read=True, now=now)[1] is None
assert plan("GET", 429, {"X-RateLimit-Remaining-Backlinks": "0"},
{}, safe_read=True, now=now)[1] is None
assert plan("GET", 503, {}, {}, safe_read=True, now=now)[1] == 1.0
assert plan("GET", 503, {"Retry-After": "30"}, {},
safe_read=True, now=now)[1] is None
assert plan("GET", 503, {"Retry-After": "3"}, {},
safe_read=True, waited=19, now=now)[1] is None
assert plan("GET", 503, {}, {}, safe_read=True,
attempt=3, now=now)[1] is None
assert plan("POST", None, {}, {}, safe_read=True, now=now)[1] is None
assert plan("GET", None, {}, {}, safe_read=True, now=now)[1] is None
assert plan("GET", 503, {}, {}, now=now)[1] is None
assert plan("GET", 200, {}, {}, now=now)[0].startswith("done")
RFC 9110 describes both Retry-After formats: delay-seconds and an HTTP date. The parser accepts an integer delay or calculates the time until a date using an aware UTC clock. A past date becomes zero, then the planner applies a one-second minimum. Missing or malformed values use a small bounded fallback only after the endpoint passes the safe-read gate.
The caller starts at attempt one, including the initial request. If the planner permits another attempt, add its delay to the cumulative waited value and increment attempt before evaluating the next response. Never reset those counters inside an inner exception handler. The total wait budget covers retry delays; separately bound network connection and response timeouts so a slow request cannot hold the worker indefinitely.
Set safe_read=True only from a reviewed endpoint contract that confirms both safe repetition and absence of per-call quota consequences. Do not set it for crawlgraph's GET changes endpoint merely because it is a GET. This deliberately cautious sample also stops transport failures with no status. If your integration later adds a narrower recovery path, establish what happened to the original request before enabling that path.
reconcile ambiguous submissions
A timeout tells you that the client did not receive a complete response within its deadline. It does not tell you that the server rejected the request. The server may have consumed quota, started a query or accepted a job before the connection failed. RFC 9110 cautions against automatic retries of non-idempotent requests unless the client knows repetition is safe or knows the original was not applied.
crawlgraph does not document a public Idempotency-Key contract. Adding a randomly generated header does not establish deduplication. Do not automatically repeat POST backlink lookups or asynchronous gap submissions after an ambiguous timeout. If a job identifier was received, retain it and use documented status retrieval. If no identifier was received, preserve the timestamp, endpoint and available request ID for investigation rather than guessing whether a replacement is needed.
An explicit error is more information than a lost connection, but still requires interpretation. A server error after charging is not a promise that another submission will be free. A local deduplication record can prevent your own workers from sending duplicates; it cannot retroactively prove what the remote service accepted. Use a pending-unknown state until the evidence supports completion, failure or a deliberate new submission.
plan capacity before the batch
Count unique operations, reserve room for investigation, and reuse saved successful responses. For an illustrative batch of 40 distinct backlink queries, a free monthly allowance of 15 cannot cover the batch. Splitting it into faster concurrent workers changes neither the arithmetic nor the reset time. Choose a smaller research set, schedule the remaining work for a later month, or review the current paid API access plans when the larger recurring workload justifies them.
Do not confuse a locally cached result with a provider's internal query cache. Your cache can avoid sending another request. A provider cache can accelerate processing while still counting the request against its allowance. Store the domain, parameters, named release and result together, so reuse is intentional and a report does not silently mix different snapshots.
For multiple workers, coordinate admission to each quota bucket in one shared ledger. A remaining count observed by one worker may already be stale when another worker sends a request. Treat response headers as evidence from that response, not a reservation of future capacity. Keep concurrency modest enough that an exhaustion response can stop the queue before a large group of already-dispatched requests makes the final usage hard to explain.
keep an operational evidence trail
Log the request time, operation type, response status, provider request ID, matching quota bucket, remaining count, reset timestamp and chosen decision. Avoid recording keys or authorization headers. For ambiguous outcomes, keep a local operation identifier and the normalized input needed to investigate the job. These records should answer why work stopped without exposing a customer's entire report.
Separate completed, rejected, unknown and deferred operations in the batch summary. A 429 should not erase the successful work already saved. Report what remains and why: monthly exhaustion, invalid input, access restriction or an unresolved submission. The useful completion criterion is a reproducible result set plus an honest remainder, rather than a retry loop that eventually exits with no explanation.
faq
does every 429 mean the monthly quota is exhausted?
No. HTTP 429 means too many requests. A provider may use it for a burst limit or a monthly allowance. In crawlgraph, the quota_exceeded error identifies monthly exhaustion; inspect the matching quota headers and reset time.
can Retry-After be a date?
Yes. Retry-After can contain delay-seconds or an HTTP date. Calculate the date delay using UTC, clamp past dates to zero, and apply a minimum delay and a total wait budget before any permitted retry.
should a timed-out backlink POST be retried?
Not automatically. A timeout does not establish whether the server charged the call or accepted a job. crawlgraph has no documented public Idempotency-Key contract. Reconcile the outcome before submitting again.
does a cached or failed request always preserve quota?
No. Repeated successful quota-bearing lookups can consume another call. The implementation charges before query or job work, so a failure after charging can also consume quota. Save results locally and inspect usage rather than assuming failed attempts were free.
documents crawlgraph backlink workflows, open-data methods, and the limits of each report.
plus one when a new common crawl release lands. that is all.