MCP documentation menu

DataPulse Methodology and Data Coverage

What is in the DataPulse corpus, how names get there, and — most importantly — what a name’s presence or absence actually tells you.

Unit of analysis: PSL+1

DataPulse operates at the PSL+1 level: the registrable domain, meaning one label below a public suffix as defined by the Public Suffix List.

NamePublic suffixPSL+1 (the corpus entry)
foo.comcomfoo.com
example.co.ukco.ukexample.co.uk
a.b.c.foo.comcomfoo.com
microsoft.jp.netjp.net (registry-run private suffix)microsoft.jp.net
octocat.github.iogithub.io (platform private suffix)none: not a corpus entry

Every corpus entry is a PSL+1 name. Hostnames and subdomains are not corpus entries. a.b.c.foo.com is not a record; foo.com is. Tools that search “labels” are searching the PSL+1 label against its public suffix.

Two refinements at the edges of that rule:

  • Delegated public suffixes are entries too. co.uk, jp.net and github.io are themselves delegated names with nameservers, and each has its own record (the history tool returns co.uk’s nameserver timeline). They are simply not reduced any further.
  • Private-section suffixes count only when a registry sells delegations under them. microsoft.jp.net is a delegated zone sold by CentralNic and is an entry. A hostname under a hosting platform’s suffix, such as octocat.github.io or foo.blogspot.com, is a record inside the platform’s zone, not a delegation, and is not an entry however live the site is. The platform’s own name (github.io) is.

One concept, several names. PSL+1, eTLD+1, and registrable domain all denote the same thing and appear interchangeably across DataPulse tools and docs — the scrape tools say “registrable domain”, and link_domains splits page links on eTLD+1. Read them as synonyms.

Second-level domain is a looser synonym that breaks on multi-label suffixes: for example.co.uk the registrable label is example and the suffix is co.uk, so the name is PSL+1 but not literally second-level. PSL+1 is the precise framing and is what the tools implement.

Registry coverage

DataPulse covers all PSL registries:

  • ICANN registries — legacy gTLDs and new gTLDs
  • Non-ICANN registries
  • ccTLD registries

There is no tier of TLD that is out of scope. Note that these three classes describe collection coverage; downstream analysis often groups differently — a common split is ICANN gTLDs against everything else, with ccTLDs and private/PaaS suffixes sharing the second bucket. A two-bucket report is a grouping choice, not a narrower corpus. Where a ccTLD appears less complete, the cause is almost always the RDAP protocol — some ccTLD registries publish no RDAP server, which affects registration lookups only and not search, history, similarity, or clustering. See datapulse_help(topic="rdap").

Wildcard awareness

Some registries wildcard-resolve: every nonexistent name under them returns an answer rather than NXDOMAIN. DataPulse identifies the PSL registries that do this and accounts for it during collection.

This matters because without that handling, a wildcarding registry would appear to contain every string anyone ever queried, and the corpus would fill with names that do not exist. Wildcard responses do not manufacture corpus entries.

Discovery sources

Public sources. All known publicly available sources, including the zone files ICANN makes available.

Certificate Transparency, in realtime. DataPulse reviews CT logs as they publish and extracts the PSL+1 of every name observed. If a.b.c.foo.com appears in a CT entry, then foo.com is added to the corpus under .com if it is not already present.

Note carefully what this does and does not add: CT contributes the registrable domain, not the hostname. The subdomains in a CT entry are the input to that reduction, not records in their own right. CT ingestion does not give DataPulse a subdomain inventory, and there is no tool that enumerates hostnames seen in CT.

Proprietary techniques. A number of non-public methods for discovering and processing domain names, developed over a decade of operating this corpus.

The quality gate: active and functioning

This is the part that most changes how you should read a result.

Discovery is not admission. A name found by any source above enters the corpus only if it passes a liveness and function check — the label must be active in the DNS and functioning on the internet, established by a technical check. Names that fail are not added. Names that later fail are removed. Expired, dead, and certain classes of misconfigured labels are excluded.

The corpus is therefore deliberately not “every string ever registered.” It is the set of names that are live and working. DataPulse trades recall for precision on purpose, and the emphasis on filtering out false positives is a design commitment, not an artifact.

How to read presence and absence

Absence is a statement about liveness, not about discovery

With realtime CT ingestion, published zone files, and continuous proprietary discovery all running, “DataPulse never heard of this name” is a weak hypothesis. When a domain is present in WHOIS/RDAP but absent from DataPulse listings, the strong hypothesis is that it did not pass — or no longer passes — the active-and-functioning check.

That is the reasoning behind datapulse_help(topic="why_missing"), and CT ingestion makes it more reliable rather than less: it closes the discovery gap, leaving liveness as the explanation that actually accounts for the absence.

Do not tell a user “this domain may simply not be in our data yet” as a first explanation. Run the why_missing workflow and establish the actual DNS and registry state.

Presence means the name was live when listed

A listed name passed the function check at listing time. That is evidence about function only. It is not evidence of legitimacy, ownership, safety, or good standing, and it says nothing about who controls the name — registrant identity is GDPR-redacted and is not in the corpus at all (see datapulse_help(topic="registrant_lookup")).

Discovery source is not queryable, and not reportable

Once a name is in the corpus it is an entry like any other. There is no CT-only tier, no source filter, and no provenance field in any tool response.

Never describe a result as “found via CT” or attribute a record to a particular discovery source. The tools do not return that information, so any such claim is fabricated. Cite what the tool actually returned.

What the corpus does not contain

Not presentWhy
Hostnames and subdomainsThe unit is PSL+1; hostnames reduce to their registrable domain
Registered but inactive namesExcluded by the active-and-functioning gate — see why_missing
Expired and dead namesRemoved when they stop functioning
Certain misconfigured labelsExcluded as false positives
Registrant identityGDPR-redacted at the source; no reverse-WHOIS is possible

Generated from the live server (DataPulse MCP 1.0.0) on October 1, 2026. Your AI assistant reads this page by calling datapulse_help(topic="methodology").