DataPulse Methodology and Data Coverage
What is in the DataPulse corpus, how names get there, and — most importantly — what a name’s presence or absence actually tells you.
Unit of analysis: PSL+1
DataPulse operates at the PSL+1 level: the registrable domain, meaning one label below a public suffix as defined by the Public Suffix List.
| Name | Public suffix | PSL+1 (the corpus entry) |
|---|---|---|
foo.com | com | foo.com |
example.co.uk | co.uk | example.co.uk |
a.b.c.foo.com | com | foo.com |
microsoft.jp.net | jp.net (registry-run private suffix) | microsoft.jp.net |
octocat.github.io | github.io (platform private suffix) | none: not a corpus entry |
Every corpus entry is a PSL+1 name. Hostnames and subdomains are not corpus entries.
a.b.c.foo.com is not a record; foo.com is. Tools that search “labels” are searching the
PSL+1 label against its public suffix.
Two refinements at the edges of that rule:
- Delegated public suffixes are entries too.
co.uk,jp.netandgithub.ioare themselves delegated names with nameservers, and each has its own record (the history tool returnsco.uk’s nameserver timeline). They are simply not reduced any further. - Private-section suffixes count only when a registry sells delegations under them.
microsoft.jp.netis a delegated zone sold by CentralNic and is an entry. A hostname under a hosting platform’s suffix, such asoctocat.github.ioorfoo.blogspot.com, is a record inside the platform’s zone, not a delegation, and is not an entry however live the site is. The platform’s own name (github.io) is.
One concept, several names. PSL+1, eTLD+1, and registrable domain all denote the same thing and appear interchangeably across DataPulse tools and docs — the scrape tools say “registrable domain”, and
link_domainssplits page links on eTLD+1. Read them as synonyms.Second-level domain is a looser synonym that breaks on multi-label suffixes: for
example.co.ukthe registrable label isexampleand the suffix isco.uk, so the name is PSL+1 but not literally second-level. PSL+1 is the precise framing and is what the tools implement.
Registry coverage
DataPulse covers all PSL registries:
- ICANN registries — legacy gTLDs and new gTLDs
- Non-ICANN registries
- ccTLD registries
There is no tier of TLD that is out of scope. Note that these three classes describe
collection coverage; downstream analysis often groups differently — a common split is
ICANN gTLDs against everything else, with ccTLDs and private/PaaS suffixes sharing the
second bucket. A two-bucket report is a grouping choice, not a narrower corpus. Where a ccTLD appears less complete, the cause
is almost always the RDAP protocol — some ccTLD registries publish no RDAP server, which
affects registration lookups only and not search, history, similarity, or clustering. See
datapulse_help(topic="rdap").
Wildcard awareness
Some registries wildcard-resolve: every nonexistent name under them returns an answer rather than NXDOMAIN. DataPulse identifies the PSL registries that do this and accounts for it during collection.
This matters because without that handling, a wildcarding registry would appear to contain every string anyone ever queried, and the corpus would fill with names that do not exist. Wildcard responses do not manufacture corpus entries.
Discovery sources
Public sources. All known publicly available sources, including the zone files ICANN makes available.
Certificate Transparency, in realtime. DataPulse reviews CT logs as they publish and
extracts the PSL+1 of every name observed. If a.b.c.foo.com appears in a CT entry, then
foo.com is added to the corpus under .com if it is not already present.
Note carefully what this does and does not add: CT contributes the registrable domain, not the hostname. The subdomains in a CT entry are the input to that reduction, not records in their own right. CT ingestion does not give DataPulse a subdomain inventory, and there is no tool that enumerates hostnames seen in CT.
Proprietary techniques. A number of non-public methods for discovering and processing domain names, developed over a decade of operating this corpus.
The quality gate: active and functioning
This is the part that most changes how you should read a result.
Discovery is not admission. A name found by any source above enters the corpus only if it passes a liveness and function check — the label must be active in the DNS and functioning on the internet, established by a technical check. Names that fail are not added. Names that later fail are removed. Expired, dead, and certain classes of misconfigured labels are excluded.
The corpus is therefore deliberately not “every string ever registered.” It is the set of names that are live and working. DataPulse trades recall for precision on purpose, and the emphasis on filtering out false positives is a design commitment, not an artifact.
How to read presence and absence
Absence is a statement about liveness, not about discovery
With realtime CT ingestion, published zone files, and continuous proprietary discovery all running, “DataPulse never heard of this name” is a weak hypothesis. When a domain is present in WHOIS/RDAP but absent from DataPulse listings, the strong hypothesis is that it did not pass — or no longer passes — the active-and-functioning check.
That is the reasoning behind datapulse_help(topic="why_missing"), and CT ingestion makes
it more reliable rather than less: it closes the discovery gap, leaving liveness as the
explanation that actually accounts for the absence.
Do not tell a user “this domain may simply not be in our data yet” as a first explanation.
Run the why_missing workflow and establish the actual DNS and registry state.
Presence means the name was live when listed
A listed name passed the function check at listing time. That is evidence about function
only. It is not evidence of legitimacy, ownership, safety, or good standing, and it says
nothing about who controls the name — registrant identity is GDPR-redacted and is not in the
corpus at all (see datapulse_help(topic="registrant_lookup")).
Discovery source is not queryable, and not reportable
Once a name is in the corpus it is an entry like any other. There is no CT-only tier, no source filter, and no provenance field in any tool response.
Never describe a result as “found via CT” or attribute a record to a particular discovery source. The tools do not return that information, so any such claim is fabricated. Cite what the tool actually returned.
What the corpus does not contain
| Not present | Why |
|---|---|
| Hostnames and subdomains | The unit is PSL+1; hostnames reduce to their registrable domain |
| Registered but inactive names | Excluded by the active-and-functioning gate — see why_missing |
| Expired and dead names | Removed when they stop functioning |
| Certain misconfigured labels | Excluded as false positives |
| Registrant identity | GDPR-redacted at the source; no reverse-WHOIS is possible |
Related Topics
datapulse_help(topic="why_missing")— the workflow for a domain that is registered but not listeddatapulse_help(topic="searchlabels")— searching PSL+1 labels across TLDsdatapulse_help(topic="rdap")— registration data, status codes, and the ccTLD RDAP limitationdatapulse_help(topic="registrant_lookup")— why owner data is unavailabledatapulse_help(topic="history")— how listing events (version typesa/c/d) are recorded
Generated from the live server (DataPulse MCP 1.0.0) on October 1, 2026. Your AI assistant reads this page by calling datapulse_help(topic="methodology").