MCP documentation menu

datapulse_domain_overview

Domain IntelligenceScope: datapulse:scrapeOutput: JSONRead-only

Start here. One call: web scrape + DNS + RDAP for a domain

Description

One-stop domain intelligence: scrape + DNS + RDAP in a single call.

Accepts a domain name, submits any missing jobs (web scrape, DNS lookup, RDAP registration data), waits for all to complete, and returns a token-efficient combined result.

This is the recommended starting point for domain research. Returns:

  • summary: key facts at a glance. Always present: resolves, dnssec (signed | unsigned | broken | unknown), registrar_name (the one derived value: looked up from the IANA registrar list by rdap ianaid). dnssec is “unknown” and resolves is false when dns is null or dns.status is “error”, i.e. the lookup itself failed and nothing was observed — that is not a finding about the name. CONDITIONAL, absent when the source had no value: registrar_iana_id (the newest RDAP observation’s ianaid), registration_date, expiration_date, entity_name, entity_type, nature, country — the last four exactly as the extractor wrote them, including “unknown” and “other” when it declined to classify.
  • scrape: the root URL’s web scrape with LLM-extracted entity data, or null when no scrape row exists yet (note says so). Link data is NOT included by default — see link_data below. other_scrapes lists any further rows the API returned.
  • dns: summarized DNS records (answer records only, no RRSIG/NSEC3), with the worker’s error text verbatim when the lookup failed; null when the API returned no DNS row
  • rdap: the API’s RDAP envelope as sent ({query, datetime, results: [observations]}), names and codes case-folded; null when there is no envelope, with rdap_status and rdap_error saying what the RDAP row reported

NOTE: summary entity_name/entity_type/nature/country come from LLM extraction of the site’s WEB CONTENT, not from RDAP registrant records (those are GDPR-redacted). country is the extractor’s own wording (“USA”, “United Kingdom”), not an ISO 3166 code like RDAP’s two-letter country; do not compare the two as codes.

NO link data and NO page content are returned unless you ask. link_data selects between link_domains (~500 bytes: the internal/external domains the page references) and the individual links, which run to hundreds of objects and tens of kilobytes on a real commercial site and dwarf everything else. include_dom adds the cleaned HTML DOM, typically 10KB-500KB.

First call for a new domain takes 15-30 seconds (jobs must complete). Subsequent calls return cached results instantly unless force=true.

For full raw DNS output (RRSIG, NSEC3, authority sections), use datapulse_live_dns instead.

For usage guidance, call datapulse_help(topic="scraping")

Parameters

ParameterTypeDescription
domain
required
string
The domain name to look up (e.g. “osu.edu”). Accepts Unicode (IDN) domains. Its spelling is canonicalized and no labels are removed (www. kept); a single label (localhost) is rejected. When canonicalizing changed more than letter case, the response carries submitted_domain with what you sent.
forceboolean
Force re-fetch even if recent results exist. Default false.
Default: false
include_domboolean
Include the cleaned HTML DOM in the scrape result. WARNING: typically 10KB-500KB. The page content is omitted by default and is not built server-side unless requested. Default false.
Default: false
link_datastring

Which link data to return. NOTHING is returned unless you ask — link data is voluminous and the caller decides whether to pay for it.

  • “none” (default) — no link data at all.
  • “domains” — link_domains: the internal hostnames and external registrable domains the page references. ~500 bytes, and the answer to most link questions.
  • “links” — the individual links: URL, source element, rel, anchor text.
  • “both” — link_domains and links together. WARNING: “links” and “both” dominate the response — 509 links and 76KB on ruggable.com against 7.5KB for everything else combined. They are NOT truncated: a partial link list is misleading rather than merely incomplete, because the link that mattered may be the one cut. An unrecognised value is treated as “none”, so a mistake costs you data rather than silently handing back a shape you did not ask for; the response says so.
Default: "none"
nonedomainslinksboth
Input schema (JSON)
{
  "additionalProperties": false,
  "properties": {
    "domain": {
      "description": "The domain name to look up (e.g. \"osu.edu\"). Accepts Unicode (IDN) domains. Its spelling is canonicalized and no labels are removed (www. kept); a single label (localhost) is rejected. When canonicalizing changed more than letter case, the response carries submitted_domain with what you sent.",
      "type": "string"
    },
    "force": {
      "default": false,
      "description": "Force re-fetch even if recent results exist. Default false.",
      "type": "boolean"
    },
    "include_dom": {
      "default": false,
      "description": "Include the cleaned HTML DOM in the scrape result. WARNING: typically 10KB-500KB. The page content is omitted by default and is not built server-side unless requested. Default false.",
      "type": "boolean"
    },
    "link_data": {
      "default": "none",
      "description": "Which link data to return. NOTHING is returned unless you ask — link data is voluminous and the caller decides whether to pay for it.\n- \"none\" (default) — no link data at all.\n- \"domains\" — `link_domains`: the internal hostnames and external registrable domains the page references. ~500 bytes, and the answer to most link questions.\n- \"links\" — the individual links: URL, source element, rel, anchor text.\n- \"both\" — link_domains and links together.\nWARNING: \"links\" and \"both\" dominate the response — 509 links and 76KB on ruggable.com against 7.5KB for everything else combined. They are NOT truncated: a partial link list is misleading rather than merely incomplete, because the link that mattered may be the one cut.\nAn unrecognised value is treated as \"none\", so a mistake costs you data rather than silently handing back a shape you did not ask for; the response says so.",
      "enum": [
        "none",
        "domains",
        "links",
        "both"
      ],
      "type": "string"
    }
  },
  "required": [
    "domain"
  ],
  "type": "object"
}

Generated from the live server (DataPulse MCP 1.0.0) on October 1, 2026.