MCP documentation menu

How names and URLs are normalized

Every tool that takes a domain name, hostname, IP address or URL puts it into one canonical form before anything is sent. What differs between tools is how much of the name is kept, and what each tool refuses. This topic is the single statement of those rules; the per-tool topics refer back to it.


What every name goes through

These apply to every domain / query parameter on every tool:

StepExample
Surrounding whitespace removed" example.com " → example.com
ASCII case foldedExample.COM → example.com
Unicode label separators folded to . (。 . 。)example。com → example.com
Fullwidth and other compatibility characters mapped (UTS-46)example.com → example.com
IDN converted to its A-label (punycode) formbücher.de → xn--bcher-kva.de
Every trailing dot removed, with any whitespace among themexample.com., example.com.., example.com. . → example.com

The U-label and A-label spellings of a name are the same lookup. A run of trailing dots is input noise, not an empty label: example.com.. is example.com, never an error. An empty label anywhere else (a..b.com, .example.com) is refused, and so is whitespace inside a name (exa mple.com).

Refused everywhere a domain is required: a URL (https://example.com/path), a port (example.com:8080), credentials, an IP literal (except where the tool takes addresses, below), a label over 63 bytes or a name over 253, a label starting or ending with a hyphen. The refusal names the reason and nothing is sent.

Most lookup tools are lenient about one thing (datapulse_disputedb_lookup is strict): names that registries actually sold but IDNA 2008 forbids (emoji and symbol names, raw or as xn-- punycode) are accepted, because they exist and must be queryable.


Three models: how much of the name is kept

FamilyToolsWhat is looked up
Registration — reduces to the registrable domaindatapulse_live_rdap, datapulse_live_rdap_bulk, datapulse_dns_rdap_historyHostname labels above the longest public suffix are removed (full Public Suffix List, private suffixes included): WWW.Example.COM. → example.com, www.microsoft.jp.net → microsoft.jp.net. Registries only know registered names. datapulse_live_rdap and its bulk form also take IPv4/IPv6 addresses, put into canonical form.
Hostname kept — canonical spelling, no labels removeddatapulse_live_dns, datapulse_live_dns_bulk, datapulse_dns_dphistory, datapulse_dns_dptechsim, datapulse_dns_neighborhood, datapulse_live_domain_health, datapulse_domain_overview, datapulse_disputedb_lookup (domain)The canonical form of exactly the name you gave. www.example.com and example.com are different names to DNS and to the hosting corpus, and are looked up separately. A leading www. is never removed by these tools.
URLdatapulse_scrape_submit, datapulse_scrape_bulkThe input is a URL, not a name. Scheme and host are case-folded and the host canonicalized as above; path, query and a non-default port are kept byte-for-byte on the request. Only http: and https: are accepted; credentials, scheme-relative URLs and other schemes are refused.

“Hostname kept” does not mean “sent as typed”: case, whitespace, IDN spelling, Unicode separators and trailing dots are still canonicalized. It means the hostname level is preserved rather than reduced to the registrable domain.

Two tools in the middle family add their own handling downstream, without this server changing what you sent: the domain-health probe checks both the name and its www. sibling, and the DisputeDB API matches a leading www. away itself, so www.tesla.com and tesla.com find the same decisions (the response’s query still shows the name as sent).

URL identity (url_hash)

A scrape result is keyed by a hash of the URL’s identity form, which normalizes more than the request does. These are the same stored page: http:// and https://; with and without a leading www.; with and without the default port (:80, :443); with and without a #fragment; query parameters in any order; percent-encoded and literal spellings of the same character; paths with and without ./.. segments; U-label and A-label hosts. A non-default port (:8443) is part of the identity and is kept.


Single labels and underscores

InputAccepted byRefused by
A single label (localhost, com)datapulse_live_dns, datapulse_live_dns_bulk only — a TLD or a bare host is a fair DNS question (the API may still reject a reserved name such as localhost)Every other tool, before any request: nothing registrable, hosted, scraped or disputed has fewer than two labels, so the lookup could only ever say “nothing found”
An underscore in a hostname label (_dmarc.example.com)Every tool. The registration tools reduce it to example.com like any other hostname—
An underscore in the registrable domain (exa_mple.com)The hostname-kept tools except domain health: DNS permits _ in an owner namedatapulse_live_rdap, datapulse_live_rdap_bulk, datapulse_dns_rdap_history, datapulse_live_domain_health: no registry registers such a name, so the answer could only be “not found”

submitted_domain / submitted_url: what you sent, when it changed

When the name that was looked up differs from what you typed by more than ASCII case — an IDN folded to punycode, a trailing dot dropped, a hostname reduced to its registrable domain — the response carries submitted_domain (or submitted_url) with your original input beside the result. It is absent when nothing changed. This matters for homographs: pasting instagrạm.com out of a phishing sample must not silently produce a report that reads as if it were about instagram.com.

ToolEcho
datapulse_live_rdap, datapulse_dns_rdap_history, datapulse_dns_dphistory, datapulse_live_dns, datapulse_live_domain_health, datapulse_domain_overview, datapulse_disputedb_lookupsubmitted_domain field
datapulse_scrape_submitsubmitted_url field
datapulse_dns_dptechsim, datapulse_dns_neighborhoodThe output is CSV, so the echo is a hint line: submitted_domain: <input> (sent as <name>)
Bulk toolsNo echo of the raw input; submitted[].query / sent[].url is the canonical form, and index maps it back to your array

Bulk receipts: where every input is accounted for

datapulse_live_dns_bulk, datapulse_live_rdap_bulk and datapulse_scrape_bulk validate every item before the request. One bad item never costs you the batch (a batch with nothing usable is an error, since there is nothing to send).

FieldMeaning
accepted, cached, rejectedThe API’s own counts, verbatim, over the items that were sent. They are disjoint: queued, answered from cache, refused by the API.
rejected_items (rejected_urls for scrape)The items the API refused: index, query (url for scrape) and error, the API’s code. rejected counts these.
rejected_localThe items refused before the request — malformed, blank, a single label or reserved address where those are not allowed: index, query (url) and error, the reason. Not counted in rejected, because the API never saw them.
duplicatesDNS and RDAP bulk only: items that normalize to an earlier item (index, query, first_index). Submitted once. datapulse_scrape_bulk has no such list; a repeated URL is simply sent.
submitted (sent for scrape)Every item that was sent, in canonical form, with its hashes and the API’s job_id / cached when it reported them. “Sent” is not “accepted”: an item the API refused is listed here without a job_id and in rejected_items.

Every index is the position in the array you passed. To account for every input:

accepted + cached + rejected + len(rejected_local) + len(duplicates) = number of items you passed

(duplicates is absent, so zero, for datapulse_scrape_bulk.)

To find everything that failed, read both rejected_items (rejected_urls) and rejected_local. An empty rejected_items with rejected: 0 means the API refused nothing; it says nothing about what was refused before the request.

Generated from the live server (DataPulse MCP 1.0.0) on October 1, 2026. Your AI assistant reads this page by calling datapulse_help(topic="normalization").