why backlink counts differ

Reconcile two backlink reports with a worksheet for scope, dates, link definitions, filters, and export limits, plus a worked example.

crawlgraph team
· 9 min read · 1,734 words
sharexlinkedin

Backlink counts differ because reports can measure different targets, units, dates, and kinds of links, then expose different slices of those observations. Two hypothetical reports, one with 80 source pages and another with 12 referring domains, may describe the same relationships. Before choosing a winner, align the report settings and separate the reported total from the rows you retrieved. The worksheet and illustrative example below show how to reconcile the numbers without assuming any tool sees the entire web.

If you are assembling your first comparison, start with our guide to finding backlinks for free. For the wider map of datasets and their limitations, use the backlink data hub. This article focuses on reconciling two exports you already have, rather than estimating which vendor has the largest index.

make both reports answer the same question

Write the question above your spreadsheet before comparing totals. For example: how many distinct referring domains were observed linking to this whole domain in this selected dataset? That differs from asking how many source pages currently contain a link to one product URL. A target typed into a search box does not tell you which interpretation the report uses. Check its scope control and documentation.

Separate the target scope from the counting unit. An exact URL, a host such as shop.example.com, and a pay-level domain such as example.com describe possible target scopes. A link occurrence, source URL, referring host, or referring pay-level domain describes what you count. A whole-domain target can still return page-level rows. Conversely, one target URL can receive links from several hosts belonging to one referring domain.

Common Crawl publishes both host and pay-level-domain graphs. Its domain graph aggregates hosts using the Public Suffix List, and graph edges can include assets and other link types. A graph relationship therefore should not automatically be relabeled as an editorial page backlink. Common Crawl graph documentation explains those levels. Record the aggregation rule rather than stripping every hostname to its final two labels, which would mishandle suffixes such as co.uk.

one dataset, two different totals

The following is an invented teaching dataset, not a vendor benchmark or a crawl of a customer site. Every source links to garden.example. Assume six distinct source URLs, with one recorded relationship per URL, in one frozen dataset. Four are ordinary page hyperlinks, one is a stylesheet reference, and one is a page hyperlink retained from an earlier observation but marked historical.

Source URLRelationshipRecorded status
news.example/story-aPage hyperlinkActive in example
news.example/story-bPage hyperlinkActive in example
blog.news.example/postPage hyperlinkActive in example
club.example/resourcesPage hyperlinkActive in example
assets.example/style-pageStylesheet referenceActive in example
archive.example/old-postPage hyperlinkHistorical in example

Report A counts distinct source URLs across all recorded relationship types and statuses. Its total is six. Report B includes only active page hyperlinks and counts referring pay-level domains. It keeps the first four rows, then merges news.example and blog.news.example into news.example. Its total is two: news.example and club.example. Both totals follow their declared definitions; the difference does not require either report to have missed a page.

Reconcile in stages so the explanation remains auditable. Restrict A to active page hyperlinks: six becomes four source URLs. Aggregate those four URLs by host: four becomes three hosts. Aggregate those hosts by pay-level domain: three becomes two domains. Each reduction has a named cause. The two story URLs explain the first aggregation; the blog subdomain explains the second. Avoid attributing the entire six-versus-two difference to freshness.

Now imagine A displays a total of six but its downloaded file contains only the first three rows. That file describes two hosts within one pay-level domain. You cannot reconstruct the six-row domain total from those three rows, because the excluded rows might introduce additional domains. Write “three retrieved source URLs from six reported” instead of presenting one domain as the complete result. This small example illustrates why changing the aggregation of a capped export can create a second, separate discrepancy.

the larger number answers a different questionIn this example, six source URLs and two active referring domains are compatible. Preserve both original reports, then create a derived comparison with matching definitions. Do not overwrite the source totals to make the spreadsheet agree.

the reconciliation worksheet

Add one column for each report and complete every row below. An unknown value is a useful finding: it tells you which part of the comparison remains unsupported. Leave it marked unknown until the report or documentation supplies an answer. Do not infer a complete export merely because the download succeeded.

FieldWhat to record for each reportWhy it changes the comparison
Target scopeEntered target; exact URL, host, or whole domainA subdomain or one page can exclude relationships to other targets.
Counting unitOccurrences, source URLs, referring hosts, or pay-level domainsRepeated pages and subdomains collapse at different levels.
Dataset and dateNamed release or index; observation dates where available; retrieval timeDownloading today does not make every observation current.
Link definitionPage hyperlinks, assets, redirects, and attributes included or excludedDifferent relationship types produce different sets.
Status and filtersHistorical or active; target paths; any score or attribute filtersA filtered count is not the unfiltered dataset total.
NormalizationURL deduplication, host grouping, canonical target treatmentEquivalent-looking addresses may merge in one report and remain separate in another.
Reported and retrieved countsReported total, actual row count, cap, and truncation noticeAn export can expose only part of the observed set.
Query outcomeSuccessful observations, no observations, unavailable dataset, or errorA failed query cannot establish that the target has zero links.

Match the fields you can actually control, then keep the unresolved fields next to your conclusion. If one export has URL rows and the other only has domains, aggregate the URLs upward to the documented domain unit. You cannot reverse that process to recover missing page URLs from a domain graph. If one report lacks timestamps or link-type metadata, state that the comparison remains partial.

Copy this blank metadata checklist alongside each untouched export. It is a recordkeeping aid, not an API request or a claim about fields every tool supplies. Fill unavailable fields with “unknown” and explain what evidence would resolve them. Saving this small note prevents the same disagreement from returning when someone opens the files a month later.

text
report_name:
target_entered:
target_scope: # exact URL / host / pay-level domain
counting_unit: # link occurrence / source URL / host / pay-level domain
source_dataset:
release_or_index_date:
retrieved_at:
query_status:
link_definition:
status_filter:
other_filters:
normalization_rules:
reported_total:
retrieved_rows:
export_or_response_cap:
truncation_or_sample_notice:
missing_metadata:
comparison_conclusion:

separate observation dates from absence

Once the units and filters match, different observations can still produce different results. Common Crawl states that it samples the web and does not generally archive every page of a website. A source missing from that sample is not proof that its link never existed or was removed. Common Crawl FAQ. Different datasets can discover different pages even when their releases cover similar periods. An overlap calculation measures agreement between those sets, not either dataset's share of every link on the web.

Distinguish the crawl observation, index or release publication, and your retrieval date. Save the selected release identifier when one exists. Two exports downloaded on the same afternoon can contain observations from different periods. Our guide to how often backlink tools crawl explains the timing stages; the article on where backlink tools get their data explains source differences. The worksheet adds the settings needed to compare the resulting reports reproducibly.

Treat Search Console as another observation source with its own definition. Google says its Links report is not comprehensive, groups target pages by canonical, and can include historical links that no longer exist. These rules can explain disagreement with a current-page export or with a report that keeps separate target URLs. Google's Links report documentation. Its total is useful context, not a universal reconciliation target.

For an important disputed relationship, inspect the source page and record the date, target address, link type, and whether the check succeeded. A present link establishes that observation at check time. An inaccessible page leaves its current content unresolved. Keep “not in this export,” “page unavailable,” and “checked page no longer contains the link” as separate outcomes rather than reducing them all to lost.

account for limits before making a decision

crawlgraph exposes host-based referring relationships from Common Crawl rather than a page-level inventory of every link occurrence. Its query result storage has a 100,000-row cap; the public backlinks API accepts a maximum of 10,000 rows per response. Read the reported total and truncation information separately from the rows returned, using the current API documentation. A paid plan does not make a capped response evidence of whole-web completeness.

If your goal is competitor research, use the same dataset, target scope, and unit for every competitor before interpreting a gap. A relationship present only in one retrieved set is a research candidate, especially when either set is capped. It does not establish that the other site has no such relationship. See the gap-analysis dashboard for comparing profiles, and current plans for paid access and export options.

Finish with a conclusion another person can reproduce: the totals differed because one report counted source URLs and the other domains; both became two domains after matching filters. Or: the normalized exports still differ, but one is capped and its observation dates are unknown, so coverage cannot be ranked. Those are actionable findings. Choosing the largest headline number is not a substitute for identifying what that number measures.

faq

Not by itself. A larger report may count individual source pages while a smaller report counts referring domains, or include historical links and additional link types. Align the target, counting unit, dates, filters, and export limits before comparing coverage. Any remaining difference concerns observations under those settings, not proven whole-web completeness.

Search Console adds another source of observations, but its Links report is not comprehensive. Google groups target pages by canonical and may retain historical links that have since disappeared. Use it to investigate particular source and target relationships, rather than treat its total as the definitive count against which every other report must agree.

No. An empty result can mean no observations for the selected target and dataset, an active filter, or an incomplete retrieval. Record the source, selected release or date, query status, and filters. Check known source pages separately before concluding that a specific link was removed.

ahrefs · backlinkslocked
upgrade required · $129/mo
crawlgraph · live $99 once
G
github.io92
C
css-tricks.com88
L
lobste.rs86
A
algolia.com84
W
web.dev80
same data · one-time
$99oncevs $129/mo
unlock the data →
stripe checkout · instant access
methodology#backlink data#methodology#referring domains
sharexlinkedin
crawlgraph team
author

documents crawlgraph backlink workflows, open-data methods, and the limits of each report.

the dispatch
one email a month.

plus one when a new common crawl release lands. that is all.