export backlinks to csv by choosing a report with the right row unit, downloading its available slice, and saving the source, date and limits alongside it. if your source returns json, convert the saved response with a csv writer that handles quoting correctly. the resulting file is a portable dataset for review or comparison, not proof that you have every backlink on the web. preserve its context before opening it in a spreadsheet.
start with the backlink data guide if you need to choose a source. our guide to finding backlinks for free explains the available research routes. here, the task is narrower: turn a bounded observation into a file another person can understand, reproduce and compare without guessing what the rows represent.
the portable export workflow
- choose the row unit. Decide whether your task needs referring domains or individual linking pages. Record the target, source, filters and selected release before exporting.
- save the bounded dataset. Download the available report or save one API JSON response. Keep the original file and record the requested limit, returned rows and any reported cap.
- convert with explicit columns. Use a CSV writer rather than joining strings. Keep source field names, preserve unknown values and store provenance in a separate metadata file.
- validate before comparing. Read the CSV back, check its headers and row count, and compare only datasets with compatible units, releases and filters.
keep an untouched original and work on a copy. give the working file a descriptive name containing the target and observation date. a filename is convenient for people, but it cannot carry every filter or limitation. the companion metadata file is where those details belong. this small separation lets you clean or annotate a spreadsheet without erasing the evidence behind it.
choose what each row means
a page-link report can describe a source url pointing to a target url. a referring-domain report groups relationships at a higher level. one domain may link from many pages, so a hundred rows in the first file can describe fewer than a hundred independent referring domains. do not compare those totals as though they measure the same thing. see the worked distinction in referring domains versus backlinks.
common crawl publishes host and pay-level-domain graphs. its domain aggregation uses the public suffix list, and graph edges can arise from assets and other link types as well as ordinary page hyperlinks. those aggregated relationships do not provide the source-page anchor text you would need for a placement audit. this is a property of the data layer, not something csv conversion can repair. the official graph documentation describes those units and link types.
source coverage also matters. common crawl describes its collection as a sample, rather than an archive of every page on each website. a missing domain in an export may reflect collection or release scope. it does not, on its own, establish that somebody removed a live link.
| dataset | row or field | interpretation |
|---|---|---|
| page-link report | source and target urls, when supplied | individual observed page relationships |
| crawlgraph web csv | rank, domain, tld, hosts, percentage | referring-domain rows in the retrieved slice |
| public api json | linking_domain, num_hosts, tld, cg_authority, cg_rank | domain results with optional graph metrics |
choose web export or api json
in crawlgraph, signed-in paid web users can export the selected domain's available csv or json slice. choose the intended release and format, save the download, then record that selection. the web csv contains exactly the five columns shown above. it does not include source page urls, anchors, link attributes or authority columns. its rank is the position in that export's ordering, not the api's graph rank.
the web percentage divides a row's host count by the sum of host counts in the retrieved slice. it is not a percentage of all links on the web. for example, in a hypothetical two-row slice with host counts of three and one, the percentages are 75 and 25. adding another retrieved row changes that denominator even when the original relationships have not changed. do not use this column as a universal market-share measure.
the public api has separate access and limits. its backlink endpoint returns json; a request defaults to 1,000 rows and accepts a maximum of 10,000. the free api allowance is 15 backlink calls per utc month, while paid api access allows 1,000 backlink calls and 50 gap calls. these api allowances do not confer paid web-export eligibility. check the currentapi documentation and plan options before choosing a route.
repeated requests can consume additional quota, and a failure after the quota charge can still count. save a response once and experiment with conversion locally. there is no documented pagination cursor that lets you recover omitted rows by repeating the same request. the underlying retrieved slice is globally capped at 100,000 rows; the api's total_linking_domains describes loaded results before its smaller request limit, and can differ from the source count above that global cap. neither number establishes whole-web coverage.
convert saved api json locally
the following python recipe uses only the standard library and performs no network requests. place a saved public api response in response.json. create context.json yourself with observed_at, source_scope and requested_limit: for example, an iso observation timestamp, a description of the selected graph source, and the actual limit used in your request. preserve those original inputs with the output.
import csv
import json
from pathlib import Path
# Use a saved public API response, not a web-export JSON file.
payload = json.loads(Path("response.json").read_text(encoding="utf-8"))
context = json.loads(Path("context.json").read_text(encoding="utf-8"))
rows = payload["results"]
if payload["returned"] != len(rows):
raise ValueError("returned does not match results")
fields = ["linking_domain", "num_hosts", "tld", "cg_authority", "cg_rank"]
with Path("backlinks.csv").open("w", encoding="utf-8", newline="") as handle:
writer = csv.DictWriter(handle, fieldnames=fields, extrasaction="ignore")
writer.writeheader()
for row in rows:
writer.writerow({key: row.get(key) for key in fields})
metadata = {
"target": payload["domain"],
"observed_at": context["observed_at"],
"release_id": payload["release_id"],
"source_scope": context["source_scope"],
"requested_limit": context["requested_limit"],
"retrieved_rows": len(rows),
"reported_loaded_domains": payload["total_linking_domains"],
"request_slice_omitted_rows": payload["total_linking_domains"] > len(rows),
"upstream_truncation": "unknown from these fields alone",
}
Path("backlinks.metadata.json").write_text(
json.dumps(metadata, ensure_ascii=False, indent=2), encoding="utf-8"
)run the script in a working directory dedicated to this export. it writes backlinks.csv and backlinks.metadata.json, so choose a fresh directory or retain older versions before running it again. the output headers deliberately use public api field names. the script does not impersonate the web export or manufacture a percentage column from a limited api response. extra response fields are left in the original json for later inspection.
null graph metrics become empty csv cells. keep them unknown rather than replacing them with zero, which would turn missing information into a numeric claim. a local round-trip check with hypothetical example.com and example.org rows verifies the five-column layout and empty-value handling. a separate csv-writer fixture checks a review note containing a comma, a newline and unicode text. these tests establish conversion behavior on constructed inputs, not the accuracy of a live backlink dataset or the validity of arbitrary strings as domains.
keep a provenance sidecar
the sidecar answers the questions a spreadsheet alone cannot: what was queried, when it was observed, which release supplied the rows and what was requested. observation time means when you retrieved the response; it is not necessarily the time the source crawled a link. retain both concepts if your source provides them. avoid labeling a historical graph release as a live crawl.
distinguish request slicing from upstream truncation. if the loaded domain count exceeds the returned rows, the api request omitted some loaded results. when the two counts match, you can say all reported loaded rows were returned, but you cannot infer that the global cap or collection coverage had no effect. the recipe therefore preserves upstream truncation as unknown. when another source supplies an explicit truncation flag, record it separately with that source's definition.
matching the csv row count to the json results verifies that conversion retained the saved records. it does not verify that the provider found every backlink. keep row-count validation and coverage claims separate in your handoff.
validate and compare your files
read the output with a csv parser, rather than counting physical lines: a quoted cell can contain a newline. check the header sequence, parsed record count and several values against the saved json. then import the file as utf-8 and review the spreadsheet's type choices. urls and domains should remain text, and empty metrics should remain empty. csv quoting protects delimiters; it is not a spreadsheet formula-safety policy for arbitrary user-entered notes. keep such notes in a separate review worksheet and treat imported text with appropriate caution.
for a comparison, align target normalization, grouping unit, filters, sort and release. retain each original export before deduplicating. if you build a distinct-domain list from a page-level file, label that transformed list and document the grouping rule. a fall in exported rows could be a smaller limit, different scope or different observation; investigate those explanations before reporting lost links.
google's links report documentation explicitly says its report is not comprehensive. use ourgoogle search console backlinks guide for that report's export process andbing webmaster backlinks guide for bing's interface. preserve each report's scope when joining those files with a domain graph. the useful result is a documented set of observations you can act on, with missing information visible to the next reader.
faq
can i export every backlink to a csv?
You can export the rows a source makes available, subject to its coverage, report scope and limits. A downloaded CSV does not prove complete web coverage. Record whether rows represent source pages, hosts or referring domains, and preserve the selected release or observation date with the file.
is crawlgraph web export the same as its public api?
No. Signed-in paid web users can download the selected domain's CSV or JSON slice. Public API access uses an API key and separate quotas. Its backlink response is JSON, with different field names and a request limit of up to 10,000 rows; converting that response locally does not unlock paid web exports.
why do two backlink csv files have different totals?
They may use different crawl periods, filters, grouping rules and caps. One file may list page links while another lists referring domains. Compare those definitions and the retrieved row counts before treating a difference as a gained or lost backlink. Neither file is automatically a complete inventory.
documents crawlgraph backlink workflows, open-data methods, and the limits of each report.
plus one when a new common crawl release lands. that is all.