Common Crawl's host graph keeps hostnames separate; its domain graph groups them into pay-level domains using Public Suffix List rules. Choose hosts when subdomain relationships matter and domains when your question concerns aggregated referring domains. Neither level gives you the exact linking page. This guide explains that choice; for the wider research workflow, start with how to find backlinks for free.
This is a worked resolution example, using invented hostnames and relationships throughout. Common Crawl's host and pay-level-domain definitions were checked on October 3, 2026 against its graph documentation. Dataset availability is a separate, changing fact. For the underlying terminology, see our backlink graph definition; the backlink data learning hub covers the broader source questions.
map hosts with public suffix rules
The following invented names form a toy mapping, not observations from a release. A registered domain comprises the public suffix plus the label immediately before it. With example.co.uk, the public suffix is co.uk. Taking the last two labels would produce the suffix itself and merge unrelated domains beneath it. A fixed last-two-label rule is therefore unsuitable for this task.
| host node | public suffix | domain node |
|---|---|---|
shop.example.com | com | example.com |
blog.example.com | com | example.com |
news.example.co.uk | co.uk | example.co.uk |
shop.example.co.uk | co.uk | example.co.uk |
news.other.co.uk | co.uk | other.co.uk |
archive.news.example.co.uk | co.uk | example.co.uk |
Notice that example.com and example.co.ukremain different domain nodes. Similar names do not merge them. The nested archive.news.example.co.uk also maps toexample.co.uk: counting labels from the hostname's left end would not identify the boundary reliably. Andother.co.uk stays separate from example.co.uk, despite sharing the same two-label suffix. Conversely, the shop and blog under example.com do merge. When reproducing a published dataset, use its documented extraction rules and record the relevant suffix-list version rather than assuming your own grouping will reproduce every node exactly.
what happens to the edges
Now map both ends of each directed host relationship. This table shows a conceptual transformation of five invented host edges. It is not a specification of Common Crawl's release-specific treatment of duplicate edges, self-loops or weights.
| host edge before mapping | mapped domain pair | information lost |
|---|---|---|
blog.example.com → shop.example.com | example.com → example.com | which subdomain linked to which |
news.example.co.uk → shop.example.com | example.co.uk → example.com | the news and shop hostnames |
shop.example.co.uk → blog.example.com | example.co.uk → example.com | the shop and blog hostnames |
news.other.co.uk → shop.example.com | other.co.uk → example.com | the news and shop hostnames |
shop.example.com → news.example.co.uk | example.com → example.co.uk | the shop and news hostnames |
Read the example as a sequence of explicit operations. First, replace each source hostname using the mapping table. Next, replace each destination hostname. Finally, compare the resulting ordered pairs. Rows two and three now have identical endpoints; row four stays separate because other.co.uk is a different registered domain. Row five preserves the reverse direction. Aggregation does not make example.com → example.co.uk equivalent toexample.co.uk → example.com.
For the three incoming cross-domain host edges in rows two through four, there are three distinct source hosts but two distinct source domains. Those are exact counts for this invented example only. If the question is which external registered domains refer toexample.com, the answer here isexample.co.uk and other.co.uk. If the question is whether the UK shop or UK news hostname supplied a relationship, that grouped answer is insufficient.
Row one maps both endpoints to example.com, creating a potential self-loop at the mapping stage. This does not establish whether a published graph retains that loop. Similarly, identical mapped pairs do not establish whether its artifact stores duplicate rows, a unique edge or a weight. Consult the selected release's construction documentation before interpreting those policies. A weight must have a documented meaning before it can be treated as a count of anything.
why the mapping cannot be reversed
Consider two alternative host-edge sets. Set A contains onlynews.example.co.uk → shop.example.com. Set B contains only shop.example.co.uk → blog.example.com. Each maps to the same domain pair:example.co.uk → example.com. Even knowing that one host edge existed would not tell you which set produced it. This is a consequence of the many-to-one mapping, not a measured coverage result or an undocumented claim about Common Crawl's weights.
The domain pair cannot identify the original source hostname, the original destination hostname or their pairing. A separate list of all hosts belonging to each domain would still leave the pairing unknown. Nor can a set of aggregated pairs reconstruct the order of an original input list. That order is not evidence of when links were created; neither a host edge nor a domain edge alone establishes a chronological page-link history.
Some details were absent even before domain aggregation. A host edge does not identify the source article URL, destination path, anchor text or placement. Switching back to hosts can preserve subdomain identity in a new analysis, but cannot recreate page details that were never present in the host-level input. Keep the source host edges separately if you will need to explain a grouped result later.
choose the resolution for your question
- Choose the host graph for subdomain separation. Ask whether observed relationships involve a blog, shop or another hostname. Keep the hostname in your comparison key.
- Choose the domain graph for grouped discovery. Ask which registered domains have observed relationships with a target, combining its hosts under the documented domain rule.
- Choose page evidence for a specific link. Ask which URL linked, what anchor it used, where it appeared and whether it is still present. Neither graph level alone supplies those answers.
Label an export with its release, node level and counting rule. For example, distinguish distinct referring hosts from distinct referring domains. Compare like units across targets and releases; otherwise a change in grouping can look like a change in links. This is a reporting recommendation, not a ranking formula.
Once you have chosen the domain dataset, the raw Common Crawl graph query guide covers resolving node IDs and selecting incoming edges, rather than this resolution decision. If you use crawlgraph instead, inspect the response fields and release metadata in the backlink API documentation. Interpret the product's returned unit as documented; the existence of an upstream host graph does not mean every product endpoint exposes host-level or page-level data.
keep resolution separate from coverage
A finer resolution does not make a crawl complete. Common Crawl describes its corpus as a sample and does not generally archive every page of a website. A missing host or domain relationship can reflect missing observations rather than the absence of a live link. That limitation applies before you choose either graph level. See the Common Crawl FAQ.
Relationships also need careful labeling: the graph documentation includes link types beyond editorial page links, including assets. A domain pair is evidence of the dataset's observed relationship, not an automatic quality judgment. Validate a candidate at the page level before describing it as an editorial opportunity.
Finally, upstream publication and product availability are separate. Common Crawl's graph release manifest identifies its published datasets. Recheck it when selecting a download; the checked definitions above are not a claim about the newest release currently available. A release must separately be ingested and made queryable by crawlgraph before a product lookup can use it. Record the release actually returned with your result, then choose the resolution that preserves the distinctions your question requires.
faq
what is the difference between a host graph and a domain graph?
A host graph keeps individual hostnames as nodes, so shop.example.com and blog.example.com can remain separate. Common Crawl's pay-level-domain graph aggregates hosts using Public Suffix List rules. Both hostnames therefore map to example.com, removing the distinction between those subdomains in the aggregated graph.
is a registered domain always the last two hostname labels?
No. Public suffixes can contain multiple labels. For shop.example.co.uk, co.uk is the public suffix and example.co.uk is the registered domain. Keeping only the last two labels would return co.uk and incorrectly combine unrelated registered domains. Use Public Suffix List rules rather than a fixed label count.
can I recover the original host edges from a domain pair?
No. Different source and destination hostnames can map to the same registered-domain pair. The grouped pair does not identify which host combination produced it. Even a separate list of hosts belonging to both domains would leave the original pairing unknown. Preserve host-level input separately if you need to explain the aggregation later.
documents crawlgraph backlink workflows, open-data methods, and the limits of each report.
plus one when a new common crawl release lands. that is all.