where backlink data comes from.
every backlink number you have ever seen is a count over somebody's crawl, and no crawler sees the whole web. that one fact explains why two tools report different totals for the same site, why a new link takes weeks to show up, and why the free sources show you so little.
the posts here go source by source: the commercial indexes and how often they recrawl, the first-party data google and bing will give you about your own site for free, and the open common crawl hyperlink graph this site is built on.
commercial indexes
how the paid tools crawl, and why they lag behind the live web.
where ahrefs and moz get their backlink data
ahrefs, moz, semrush, and majestic each run their own proprietary crawler and ship a curated index - which is why no two backlink tools ever agree. here's the methodology, the size claims, and why common crawl is the only open alternative.
11 min readhow often do backlink tools crawl? why your new link is not showing up yet
no tool sees the web live. a link must be crawled, indexed, then published before it appears. how to read your tool's date and confirm a link yourself.
6 min readmoz's crawler explained: dotbot, rogerbot, and the crawled flag
moz runs two crawlers: dotbot builds the link index behind DA, rogerbot runs site crawl in a moz pro campaign. what crawled means and why DA moves.
14 min read
free first-party sources
what google and bing will show you about links to your own site.
google search console backlinks: what it shows
gsc is the only free backlink data source straight from google - but it caps the report at ~1,000 referring domains, only shows data for sites you've verified, and surfaces a filtered subset of what google actually indexed. here's the full walkthrough plus how to get around the limits.
9 min readbing webmaster tools backlinks (free, no signup)
bing webmaster tools shows backlink data for free - and unlike google search console, it surfaces backlinks for sites you don't own through the similar sites report. here's the full walkthrough, the gsc comparison, and the export workflow.
9 min read
the open graph: common crawl
the public dataset behind this site, and when it refreshes.
common crawl, explained for SEOs
the 3.90B-edge open dataset that quietly powers half the AI training runs - and how we turn it into rank intelligence.
6 min readcommon crawl release schedule (next crawl)
common crawl publishes new web crawls every 1-2 months. here's the recent release cadence, crawlgraph's current composite, and why the lag matters for backlink freshness.
5 min readCommon Crawl index API 2026: the current CC-MAIN index IDs
The current index is cc-main-2026-apr-may-jun, built from CC-MAIN-2026-17, CC-MAIN-2026-21 and CC-MAIN-2026-25. Copy-paste API URLs for both the Common Crawl CDX index and crawlgraph.
4 min read
- query the common crawl web graph yourself. the streaming pipeline that pulls one domain's referring domains out of the raw files on a laptop.