MCP documentation menu

Web Scraping & Live Lookup Tools

DataPulse provides tools to submit URLs for scraping, perform live DNS/RDAP lookups, and retrieve structured results.

Architecture: Pages are rendered with a headless browser, then an LLM extracts structured entity data. Results are cached for 90 days.

START HERE: Domain Overview

For most domain research, use datapulse_domain_overview — it handles everything in one call.

datapulse_domain_overview(domain="osu.edu")

This single call:

  1. Runs a live DNS lookup (all record types, DNSSEC-validated)
  2. Fetches RDAP registration data (registrar, dates, status)
  3. Scrapes the website with headless browser + LLM entity extraction
  4. Returns the combined result: scrape + DNS + RDAP

First call for a new domain takes 15-30 seconds (jobs must complete). Subsequent calls return cached results instantly. Use force=true to re-fetch everything.

When to use the other tools instead:

Queue Priority: Interactive vs Bulk Scrapes

Web scrapes are processed in two classes:

Submitted viaClassWhen it runs
datapulse_scrape_submit, datapulse_domain_overviewInteractiveProcessed first — every worker keeps a scrape slot free for it
datapulse_scrape_bulkBulkLower priority. Processed only when a worker has a spare slot, behind all interactive work. During large background runs it can take hours

If someone is waiting on the answer, do not use datapulse_scrape_bulk. Call datapulse_scrape_submit(url=..., wait=true) once per URL, or datapulse_domain_overview(domain=...), instead — both are processed ahead of datapulse_scrape_bulk submissions. Use bulk for background work whose results can arrive later.

This applies to web scrapes only. datapulse_live_dns_bulk and datapulse_live_rdap_bulk share one queue with their single-lookup forms and are not deprioritised.

Bulk work can pause

Bulk scrapes also pause while the extraction model is unavailable. Bulk jobs wait in the queue, or are retried automatically, and resume on their own when it recovers. A bulk batch can therefore stall for as long as an outage lasts, on top of waiting behind interactive work. Interactive submissions are not affected: datapulse_scrape_submit and datapulse_domain_overview keep being processed throughout.

Polling a bulk batch

  • A long-pending result is not a failure. pending or processing means the URL is still queued or in progress, however long it has been. Only status: "failed" is a failure.
  • Poll with backoff. Wait at least 5 seconds, then double the interval each time (5s → 10s → 20s → 40s …), capping at a few minutes. For a large batch, check status_counts from datapulse_scrape_list(batch_id=...), or fetch many hashes in one call with datapulse_scrape_bulk_results, instead of polling every url_hash.
  • Do not resubmit the batch to unstick it. Resubmitting with datapulse_scrape_bulk uses the same lower-priority queue, so it will not run the URLs any sooner.
  • If one URL becomes urgent, resubmit it interactively with datapulse_scrape_submit(url=..., force=true, wait=true). force=true is required: without it, a URL still queued from the bulk batch returns in_flight and stays behind the bulk work.

Sync vs Async Modes

The individual scrape tools have three modes of operation. Understanding these prevents confusion — but if you’re using datapulse_domain_overview, you don’t need to worry about this.

Use wait=true on datapulse_scrape_submit. The call blocks until the scrape completes (up to 50 seconds) and returns the full result with extraction data in a single call. No polling needed.

datapulse_scrape_submit(url="example.com", wait=true)
→ Returns: complete result with status, extraction data, timing — all in one response

Use sync mode when: scraping 1 URL at a time. This is the simplest path.

Note: The sync response has the same shape as datapulse_scrape_result (including final_url, redirected, tls, quic, server_addr, and link_domains and links unless link_data narrows them). Only cleaned_dom is unavailable here — use datapulse_scrape_result(url_hash=..., include_dom=true) after the sync call completes if you need the DOM.

Async Mode (for bulk or background work)

Without wait=true, datapulse_scrape_submit returns immediately with only a job receipt — NOT the scrape result. You must then poll separately:

Step 1: datapulse_scrape_submit(url="example.com")
        → Returns ONLY: {status: "queued", job_id, domain_hash, url_hash}
        → The page has NOT been scraped yet

Step 2: datapulse_scrape_result(url_hash="...")
        → If status="pending" or "processing" → NOT DONE, poll again in 5-10 seconds
        → If status="completed" → extraction data is ready
        → If status="failed" → check error field

Use async mode when: submitting many URLs via datapulse_scrape_bulk, or when you need to submit and check results later.

Async Status Values (datapulse_scrape_submit only)

In async mode, the response status field indicates what happened:

StatusMeaningAction
queuedNew job acceptedPoll with datapulse_scrape_result(url_hash=...)
cachedPreviously cached, full result returned inlineNo polling needed — result is complete
in_flightAlready being processedPoll with datapulse_scrape_result(url_hash=...)
failedScrape failed, returns error and full result inlineNo polling needed — error detail is in the response

How to detect an inline result: When the response contains fields like extraction or completed_at, the result is complete (from cached or failed status). When the response only contains status, job_id, domain_hash, url_hash, it’s an async receipt that requires polling.

Note: This inline result behavior only applies to datapulse_scrape_submit. The datapulse_scrape_bulk tool always returns aggregate counts (accepted/rejected/cached) — it does not return per-URL results inline. To retrieve cached results for bulk-submitted URLs, use datapulse_scrape_bulk_results(url_hashes=[...]) for efficient batch retrieval, or datapulse_scrape_result(url_hash=...) for individual results.

Common Mistakes

MistakeFix
Treating the async submit response as the scrape resultWithout wait=true, submit only returns a receipt ({status, job_id, domain_hash, url_hash}). The actual content is retrieved via datapulse_scrape_result. Exception: cached and failed statuses return the full result inline.
Using async mode for a single URLUse wait=true — it returns the full result in one call, no polling needed.
Using datapulse_scrape_bulk for interactive workBulk scrapes run at lower priority, can take hours during large background runs, and pause while the extraction model is unavailable. When someone is waiting, call datapulse_scrape_submit(wait=true) per URL, or datapulse_domain_overview.
Treating a long-pending bulk result as failedA bulk-submitted URL can stay pending or processing for hours and still complete. Keep polling with backoff; only status: "failed" is a failure. If one URL is urgent, resubmit it with datapulse_scrape_submit(url=..., force=true, wait=true). See Polling a bulk batch.
Not polling after async submitAfter async submit, you MUST call datapulse_scrape_result(url_hash=...) to get actual results. Exception: cached status returns results inline.
Polling immediately without waitingScrapes take 5-30 seconds. Wait at least 5 seconds before the first poll. Use exponential backoff: 5s → 10s → 20s.
Constructing url_hash yourselfAlways use the url_hash from the submit response. URL normalization rules are non-trivial.
Using force on every requestOnly use force when you need a fresh scrape. Without it, cached results return instantly.
Requesting include_dom by defaultDOM is 10KB-1MB+ and consumes context window. Only request it for specific analysis like impersonation detection.
Expecting batch_id to include all submitted URLsCached URLs (already cached/in-flight) are excluded from the batch. Only newly accepted URLs appear in datapulse_scrape_list(batch_id=...) results. To fetch all results (including cached), pass the url_hashes array from the bulk response to datapulse_scrape_bulk_results(url_hashes=[...]) — it covers all submitted URLs regardless of accepted/cached status.
Using an invalid batch_idMalformed batch_id (not a valid UUID) returns 400 Bad Request. A valid UUID that doesn’t match any batch returns 404 Not Found. Verify your batch_id came from a datapulse_scrape_bulk response.

URL Normalization

The tool sends the URL as you gave it, with only the scheme and host case-folded and the host in its A-label form (submitted_url echoes your input when that changed it); the path, query and fragment go up byte for byte. Everything below is what the API then does on its side for consistent hashing:

  • Scheme is ignored for hashing — http:// and https:// produce the same hash
  • URLs without a scheme default to https:// for fetching; if port 443 is unreachable but port 80 is open, falls back to http:// automatically
  • www. prefix is stripped from hostname
  • Hostname is lowercased (paths are case-sensitive)
  • Default ports are removed (443 for https, 80 for http)
  • Query parameters are sorted alphabetically
  • Fragment is stripped
  • Trailing bare slash is removed
  • Unicode/IDN hostnames are converted to punycode

Important: Paths ARE significant. example.com/about and example.com/contact are different scrapes.

These all produce the same url_hash:

  • https://www.Example.com/
  • http://example.com
  • www.example.com
  • example.com

But these produce different url_hash values:

  • example.com vs example.com/about (different path)
  • example.com/page?b=2&a=1 vs example.com/page?a=1&b=2 (same — query params sorted)

Note: For root URLs (no path), url_hash equals domain_hash — both hash the bare domain string "example.com". This is expected and harmless since they are used in different contexts.

IDN domains are automatically converted per IDNA 2008 (RFC 5890-5893):

  • münchen.de → xn--mnchen-3ya.de
  • 中国.cn → xn--fiqs8s.cn

Emoji domains are not supported. Emoji are DISALLOWED by IDNA 2008 (RFC 5892, Unicode category So). Both Unicode emoji (e.g. 💩.la) and pre-encoded punycode emoji (e.g. xn--ls8h.la) are rejected per SSAC Advisory SAC095. Some ccTLD registries accept emoji registrations, but these cannot be processed consistently by IDNA-conformant applications.

Scrape Tool Reference

datapulse_domain_overview

Recommended starting point for domain research. One call to get everything: web scrape + DNS + RDAP.

Parameters

ParameterTypeRequiredDescription
domainstringYesDomain name to investigate (e.g. “osu.edu”). Accepts Unicode/IDN. Its spelling is canonicalized and no labels are removed (www. kept); a single label (localhost) is rejected before any request. When canonicalizing changed more than letter case (an IDN folded to punycode, a trailing dot dropped), the response carries submitted_domain with what you sent. See datapulse_help(topic="normalization").
forcebooleanNoRe-fetch all data even if cached (default: false)
include_dombooleanNoInclude cleaned HTML DOM in scrape results (default: false)
link_datastringNoWhich link data to return: none (default), domains, links, both

No link data is returned unless you ask. The page’s links arrive in two shapes that differ in size by two orders of magnitude:

link_dataReturnsTypical size
none (default)nothing—
domainslink_domains~500 bytes
linksthe individual links~76 KB on a commercial site
bothboth~77 KB

domains answers most link questions. links and both are returned WHOLE, never truncated — a partial link list is misleading rather than merely incomplete, because the link that mattered may be the one cut. An unrecognised value is treated as none and the response says so, so a mistake costs data rather than silently returning a shape you did not ask for.

Example

{"domain": "osu.edu"}
{"domain": "osu.edu", "link_data": "domains"}

Response

summary fields are conditional. Only resolves, dnssec and registrar_name are always present. The rest — registrar_iana_id (a number, the same value as rdap.ianaid), registration_date, expiration_date, entity_name, entity_type, nature, country — are omitted when the source had no value at all. An absent field means the source had no answer for it.

dnssec is one of signed, unsigned, broken (the validating resolver answered SERVFAIL) or unknown. unknown appears together with dns.status: "error": the DNS lookup itself failed or returned nothing usable, so nothing was observed. resolves is false in that case because no address was seen, not because the name is known to be absent — check dns.status before treating a false as a finding.

The four extraction fields are passed through exactly as the extractor wrote them, including its placeholders: entity_name: "unknown", entity_type: "other" and country: "unknown" are the extractor declining to classify, not classifications. Read them as “no answer”, never as a finding.

country is the extractor’s own wording for where the site says it operates ("USA", "United Kingdom"), not an ISO 3166 code. RDAP’s country on IP lookups is a two-letter code; the two are not comparable as codes.

The response is optimized for token efficiency with a top-level summary and compact subsections:

{
  "domain": "github.com",
  "domain_hash": "a1b2c3d4...",
  "summary": {
    "resolves": true,
    "dnssec": "unsigned",
    "registrar_iana_id": 292,
    "registrar_name": "MarkMonitor Inc.",
    "_note": "registrar_name is \"unknown\" when RDAP failed and \"not assigned (registry without IANA IDs)\" when the registry publishes no IANA registrar IDs (Nominet .uk); registrar_iana_id is then absent",
    "registration_date": "2007-10-09T18:20:50Z",
    "expiration_date": "2026-10-09T18:20:50Z",
    "entity_name": "GitHub, Inc.",
    "entity_type": "commercial",
    "nature": "Software development platform and code hosting service",
    "country": "USA"
  },
  "scrape": {
    "status": "completed",
    "url": "https://github.com",
    "final_url": "https://github.com/",
    "redirected": true,
    "tls": {
      "protocol": "TLS 1.3",
      "issuer": "DigiCert Inc",
      "subject": "github.com",
      "valid_from": "2025-02-13T00:00:00.000Z",
      "valid_to": "2026-03-14T23:59:59.000Z"
    },
    "server_addr": "140.82.121.3:443",
    "quic": {
      "supported": true,
      "alpn": "h3",
      "tls_version": "TLS 1.3",
      "server_addr": "140.82.121.3:443",
      "handshake_ms": 45
    },
    "extraction": { "entity_name": "GitHub, Inc.", "entity_type": "commercial", ... },
    "render_time_ms": 3200,
    "llm_time_ms": 4100,
    ...
  },
  "dns": {
    "status": "ok",
    "dnssec_signed": false,
    "records": {
      "A": [{"name": "github.com.", "ttl": 60, "data": {"address": "140.82.121.3"}}],
      "MX": [{"name": "github.com.", "ttl": 3600, "data": {"preference": 1, "exchange": "aspmx.l.google.com."}}],
      "NS": [{"name": "github.com.", "ttl": 900, "data": {"nameserver": "dns1.p08.nsone.net."}}],
      "TXT": [...]
    },
    "query_time_ms": 450
  },
  "rdap": {
    "name": "github.com",
    "researchdate": "2026-02-03T21:07:05Z",
    "ianaid": 292,
    "registrationdate": "2007-10-09T18:20:50Z",
    "expirationdate": "2026-10-09T18:20:50Z",
    "status": ["client delete prohibited", "client transfer prohibited"],
    "dns": ["dns1.p08.nsone.net", "dns2.p08.nsone.net"],
    "abuse_email": "abusecomplaints@markmonitor.com"
  }
}

Response sections:

  • summary: Key facts at a glance — includes resolution status, DNSSEC state, authoritative registrar ID/name, registration dates, and extracted entity classification.
  • scrape: Single scrape result object (not array). Same fields as datapulse_scrape_result. Includes final_url (actual URL rendered after redirects/protocol fallback), redirected (boolean), tls (certificate details — absent for HTTP), quic (QUIC/HTTP3 support probe — absent when not probed), server_addr (IP:port), and, when link_data asks for them, link_domains and/or links (see the parameter table above — neither is returned by default). The page content (cleaned_dom) is likewise omitted unless include_dom=true. Always present: a failed scrape carries status: "failed" and error; if the scrape has not finished yet, a placeholder with status: "processing" and an error note is returned while dns and rdap are still populated.
  • dns: Summarized DNS — answer records only, no RRSIG/NSEC3/authority sections. For NXDOMAIN domains, returns status: "nxdomain" with just the SOA. For full raw DNS output, use datapulse_live_dns.
  • rdap: The RDAP envelope as the API returned it, names and codes case-folded. results holds every observation in the API’s order; each has name, ianaid, registrationdate, expirationdate (optional), status, dns, abuse_email, abuse_phone. Null when there is no envelope, with rdap_status and rdap_error saying what the RDAP row reported. Note the expiry field is expirationdate here but expiration_date in summary and in datapulse_dns_rdap_history.

Individual job failures are non-fatal — the response includes whatever succeeded.

datapulse_scrape_submit

Submit a single URL for scraping. This is the tool for interactive use — single submissions are processed ahead of datapulse_scrape_bulk submissions. Use wait=true for single URLs (recommended).

Parameters

ParameterTypeRequiredDescription
urlstringYesURL to scrape (protocol optional)
waitbooleanNoRecommended. Sync mode: block until complete, return full result. Default: false
wait_timeoutintegerNoMax seconds to wait in sync mode (default: 50, max: 50)
forcebooleanNoBypass 90-day cache (default: false)
metadataobjectNoCustom key-value metadata stored with result (returned as metadata in all result responses)
link_datastringNowait=true only. Which link structures the sync result carries: none, domains (link_domains only), links (the links array only) or both (default). Unlike the overview, the sync result returns both by default; pass domains when you do not need every URL, since links can run to hundreds of KB. An unrecognised value is refused. Ignored without wait=true: an async receipt carries no link data, and datapulse_scrape_result always returns both.
{"url": "example.com", "wait": true}

Response — the completed result:

{
  "status": "completed",
  "domain": "example.com",
  "url_hash": "a379a6f6...",
  "extraction": {
    "entity_name": "IANA",
    "entity_type": "other",
    "nature": "Example domain",
    "language": "eng",
    "location_country": "USA",
    ...
  }
}

If the sync timeout expires before the scrape finishes, the tool returns a timeout error. To recover, resubmit without wait=true — the response will be status: "in_flight" (with url_hash), then poll with datapulse_scrape_result(url_hash=...).

Async mode example

{"url": "example.com"}

Response — just a job receipt (NOT the scrape result):

{
  "status": "queued",
  "job_id": "550e8400-e29b-41d4-a716-446655440000",
  "domain_hash": "a379a6f6eeafb9a55e378c118034e2751e682fab9f2d30ab13d2125586ce1947",
  "url_hash": "a379a6f6eeafb9a55e378c118034e2751e682fab9f2d30ab13d2125586ce1947"
}
  • status: queued (new), cached (cached result returned inline), in_flight (already being processed), or failed (failed, result inline)
  • job_id: Job identifier (UUID)
  • domain_hash: SHA-256 hash of the normalized domain, computed with the shared dpdomain convention (www. stripped, lowercase, punycode, no trailing dot) — the same value datapulse_domain_overview and datapulse_live_domain_health return
  • url_hash: SHA-256 hash of the normalized URL — use with datapulse_scrape_result to poll for actual results

Inline results: If status is cached or failed, the full result is returned inline — no polling needed. All statuses include url_hash and domain_hash; receipts (queued, in_flight) carry job_id, while inline results carry the job ID as id. For in_flight, only IDs are returned — poll with datapulse_scrape_result.

datapulse_scrape_bulk

Submit multiple URLs in a single request (up to 1000). Always async — poll results individually or use batch_id to retrieve the whole batch.

Lower priority. Bulk submissions are processed behind single submissions (datapulse_scrape_submit, datapulse_domain_overview) and can take hours during large background runs. They also pause while the extraction model is unavailable and resume automatically when it recovers, so some results can stay pending or processing much longer than usual (see Polling a bulk batch). For interactive or time-sensitive use — someone waiting on the answer — call datapulse_scrape_submit(wait=true) per URL, or datapulse_domain_overview, instead.

Parameters

ParameterTypeRequiredDescription
urlsarrayYesArray of URL objects (max 1000)

Each URL object can have: url (required), force, metadata

Example

{
  "urls": [
    {"url": "example.com"},
    {"url": "another.com", "force": true},
    {"url": "third.com", "metadata": {"campaign": "q1"}}
  ]
}

Response

Returns job receipts with a batch_id for tracking:

{
  "accepted": 2,
  "rejected": 0,
  "cached": 1,
  "batch_id": "f47ac10b-58cc-4372-a567-0e02b2c3d479",
  "job_ids": ["550e8400-...", "550e8400-..."],
  "domain_hashes": ["d1e2f3a4...", "b5c6d7e8..."],
  "url_hashes": ["a379a6f6...", "a1b2c3d4...", "c3d4e5f6..."],
  "errors": [],
  "rejected_urls": [],
  "rejected_local": [],
  "sent": [
    {"index": 0, "url": "https://example.com/"},
    {"index": 1, "url": "https://another.com/"},
    {"index": 2, "url": "https://third.com/"}
  ]
}

Every URL is validated before anything is sent, and one bad URL never costs you the batch. Every index is the position in your urls array:

  • accepted, cached and rejected are the API’s own counts over the URLs that were sent. rejected_urls (index, url, error) lists what the API refused, with its code (for example invalid_domain).
  • rejected_local (index, url, error) lists what was refused before the request — a blank item, ftp: or javascript:, credentials, a scheme-relative URL — with the reason. These are not counted in rejected, because the API never saw them. To find every failure, read both lists.
  • sent is every URL that was sent, in the form it was sent. Sent is not accepted: a URL the API refuses is in sent and in rejected_urls.
  • Every input is accounted for: accepted + cached + rejected + len(rejected_local) equals the number of URLs you passed.

Key: url_hashes contains hashes for all submitted URLs (accepted + cached), not just accepted ones. This array is what you pass to datapulse_scrape_bulk_results to fetch results. In the example above, 3 URLs were submitted → 3 hashes returned, even though only 2 were accepted.

After bulk submit:

  • Use datapulse_scrape_bulk_results(url_hashes=[...]) to fetch all results (accepted + cached) in a single call — pass the url_hashes array from this response directly.
  • Use datapulse_scrape_list(batch_id="...") to paginate through newly accepted results with status summary counts.
  • Use datapulse_scrape_result(url_hash="...") for individual results or when you need include_dom=true.

Important: cached handling. URLs that are already cached or in-flight are counted as cached and are not included in the batch_id tracking. Only newly accepted URLs appear when querying by batch_id. Cached hashes are still included in the url_hashes response array — use datapulse_scrape_bulk_results(url_hashes=[...]) to fetch all results regardless of whether they were accepted or cached.

datapulse_scrape_result

Retrieve a scrape result by url hash. Use this to poll for async results, or to retrieve results with DOM.

Parameters

ParameterTypeRequiredDescription
url_hashstringYes64-char SHA-256 hash from submit response
include_dombooleanNoInclude cleaned HTML DOM (see warning below)

Status Values

StatusMeaningAction
pendingJob waiting in queuePoll again in 5-10 seconds. A bulk-submitted URL can stay pending for hours without having failed — back off (see Polling a bulk batch)
processingCurrently being scrapedPoll again in 5-10 seconds, backing off if it persists
completedSuccessfully scraped and extractedExtraction data is ready
cachedServed from storage by a sync datapulse_scrape_submit call (the row itself stays completed)Result is inline; cached: true says the scrape did not run for this call
failedError occurredCheck error field

Warning: DOM Size and Context Windows

The include_dom=true option returns the full cleaned HTML DOM. It is now requested from the API rather than filtered on arrival, so the bytes are never built unless you ask:

  • Simple pages: 10-50KB
  • Average pages: 50-200KB
  • Complex pages: 200KB-1MB+

This can consume significant context window space. Only request DOM when needed for specific analysis tasks like impersonation detection or content comparison. For most use cases, the extraction data provides the key information without the DOM overhead.

MCP transport overflow: Results exceeding MCP context limits may be stored externally by the MCP client — you may receive a file reference rather than inline content. This is an MCP protocol-level behavior, not a DataPulse API limitation.

Response Fields

Each result includes the following fields:

FieldTypeDescription
idstringUnique job ID (UUID)
domainstringDomain extracted from URL
domain_hashstringSHA-256 hash of the normalized domain (dpdomain convention: www. stripped, lowercase, punycode)
urlstringNormalized URL
url_hashstringSHA-256 hash of normalized URL
statusstringpending, processing, completed, failed; a sync submit reports cached for a stored row
cachedbooleantrue when the row was served from storage rather than scraped for this call. Set by datapulse_scrape_submit (sync) and datapulse_scrape_list; the lifecycle status is unchanged
final_urlstringActual URL rendered after redirects and protocol fallback. Compare with url to see what changed.
redirectedbooleantrue when final_url differs from the requested URL
tlsobjectTLS certificate details: protocol (“TLS 1.3”), issuer (CA name), subject (cert CN), valid_from, valid_to (ISO 8601). Absent for HTTP connections.
quicobjectQUIC/HTTP3 support probe. supported: true with alpn (e.g. “h3”), tls_version, server_addr (IP:port) and handshake_ms; supported: false with error (a timeout means the host does not answer QUIC on UDP 443 — most hosts without HTTP/3 simply drop it) and, on rows scraped since 2026-09-15, reason (timeout, tls_rejected, resolve_failed, version_negotiation, stateless_reset, transport_error, application_error, listen_failed, invalid_args or other); a tls_rejected failure adds tls_alert (e.g. “handshake failure”) and tls_alert_code (e.g. 40), which is what a CDN with HTTP/3 disabled returns. No handshake_ms on a failure. Absent when the probe could not run.
server_addrstringIP:port of the server that responded (e.g. “140.82.121.3:443”)
errorstringError message (if failed)
created_atdatetimeWhen job was first submitted
completed_atdatetimeWhen processing finished for this submission
last_scraped_atdatetimeWhen the URL was last successfully scraped (may differ from completed_at if result was served from cache — last_scraped_at reflects the original scrape time, completed_at reflects the current job completion)
render_time_msintBrowser rendering time in ms
llm_time_msintLLM processing time in ms
total_time_msintTotal end-to-end processing time in ms
batch_idstringBatch UUID if the URL was submitted via datapulse_scrape_bulk (absent for individual submissions)
metadataobjectCustom key-value metadata from submission (always present, {} when not set). Shared, not per caller: the corpus is one store for every caller of this deployment, so metadata written by one caller is returned to any caller who retrieves the row, and a cache hit returns the metadata of whoever first scraped the URL. Never put credentials, personal data or anything confidential in it
extractionobjectStructured extracted data (see below). Omitted for non-HTML content (PDFs): a completed result with no extraction and no llm_time_ms is a PDF or other non-HTML document; pending/processing is “not yet available”.
crawler_policiesobjectRobot/AI crawler policy files discovered for the domain (robots_txt, llms_txt, ai_txt, ai_json — each a string or null). They are collected and returned verbatim, not applied: a Disallow in robots_txt does not stop the scrape. Always present (as {} when no policy files found) in datapulse_scrape_result, datapulse_scrape_bulk_results and datapulse_scrape_list; datapulse_domain_overview omits it when empty.
link_domainsobjectThe URLs referenced by the rendered page, split by ownership. Links whose host is an IP literal are in links only: an address is not a registrable domain. The two halves hold different units. internal holds full HOSTNAMES sharing the page’s registrable domain (wholesale.example.com) — you want to know which subdomains a site runs. external holds distinct third-party REGISTRABLE DOMAINS (klaviyo.com), deduplicated to the entity rather than every CDN hostname — the portfolio-discovery signal. Do not deduplicate or count across the two: they are not the same kind of thing.
linksarrayURLs referenced by the rendered page (includes JS-inserted links), up to 1500 records of {url, source, rel?, text?} where source is one of a, link, script, iframe, img, media, form, jsonld, meta. Returned by this tool, by datapulse_scrape_submit with wait=true (by default; link_data can narrow it), and by datapulse_domain_overview when link_data is links or both. List and bulk responses carry link_domains only.
cleaned_domstringCleaned HTML DOM (only with include_dom=true)

Link discovery availability: links and link_domains are populated for scrapes performed on or after 2026-07-02. Results scraped earlier lack both fields entirely — absence means “scraped before the feature existed”, not “page has no links” (a genuinely link-free page returns links: []). Re-scrape with force=true to populate them. PDF results have no rendered DOM, so both fields are always absent for PDFs.

Extraction Data

When status is completed, the extraction object contains LLM-extracted data.

Note: For non-HTML content (PDFs, binary files), the extraction field is omitted entirely and so is llm_time_ms (zero values are not sent) — LLM analysis is skipped for these content types. links and link_domains are absent too, since nothing was rendered. PDF results include markdown_content via the full API but this field is not currently surfaced through the MCP tool responses.

FieldTypeDescription
is_erroryes/noWhether page was an error page
entity_namestringOrganization/entity name
entity_typeenumcommercial, non-profit, government, personal, education, other
naturestringPage nature classification
nature_adultyes/noAdult content indicator
nature_domainstringDomain purpose classification
locationstringPhysical location
location_countrystringISO 3166-1 alpha-3 country code
phonestringContact phone number
emailstringContact email address
domain_parkedyes/noWhether domain is parked
domain_randomyes/noWhether domain appears randomly generated
ecommerceyes/noE-commerce site indicator
languagestringISO 639-3 language code
bot_detectedyes/noBot/CAPTCHA detection page identified

Note: Boolean fields return "yes" or "no" strings, not true/false.

datapulse_scrape_bulk_results

Retrieve multiple scrape results by url_hash in a single request (up to 1000). Use this any time you have multiple url_hashes and want to fetch all results efficiently in one call instead of N individual datapulse_scrape_result calls.

Common use cases:

  • After datapulse_scrape_bulk: pass the url_hashes array from the bulk response directly to get all results (accepted + cached) in one call
  • Cross-batch queries: fetch results from multiple batches or submissions at once
  • Any scenario where you have a set of known url_hashes

Parameters

ParameterTypeRequiredDescription
url_hashesarray of stringsYesSHA-256 URL hashes to retrieve (max 1000)

Each hash must be a 64-character hex string. Invalid hashes return 400 Bad Request.

Example

{
  "url_hashes": ["a379a6f6...", "a1b2c3d4...", "deadbeef..."]
}

Response

{
  "count": 2,
  "results": [
    {"url_hash": "a379a6f6...", "status": "completed", "extraction": {...}, ...},
    {"url_hash": "a1b2c3d4...", "status": "pending", ...}
  ],
  "not_found": ["deadbeef..."]
}
  • count: Number of results found
  • results: Array of scrape results with the same fields as datapulse_scrape_result (see Response Fields table above), excluding cleaned_dom, semantic_snapshot, the links array, and the per-connection fields final_url, redirected, tls, quic and server_addr — the bulk read is served from a lean projection; fetch a single result for those (link_domains is included). Includes error field on failed results, batch_id on batch-submitted results, and crawler_policies.
  • not_found: Array of url_hashes that had no matching result in storage

In-flight jobs: Hashes for URLs that are still pending or processing appear in results (with their current status), not in not_found. A hash only appears in not_found if no record exists at all — e.g., a hash that was never submitted or is simply invalid.

datapulse_scrape_list

List scrape results with pagination. Supports filtering by batch_id from bulk submissions. Each result includes an error field (present only when status is "failed") so you can identify failures without individual polling.

The listing is global, not per caller. It covers every row in this deployment’s corpus, including results and metadata submitted by other callers; nothing is scoped to the OAuth identity that made the call.

Parameters

ParameterTypeRequiredDescription
limitintegerNoResults per page (default: 100, max: 1000)
offsetintegerNoSkip N results for pagination
batch_idstringNoFilter to a specific bulk submission batch (UUID from datapulse_scrape_bulk response)

Example — list all results

{
  "limit": 50,
  "offset": 100
}

Example — list results for a batch

{
  "batch_id": "f47ac10b-58cc-4372-a567-0e02b2c3d479",
  "limit": 200
}

When batch_id is provided, the response includes additional fields:

{
  "count": 200,
  "total": 1597,
  "batch_id": "f47ac10b-58cc-4372-a567-0e02b2c3d479",
  "status_counts": {
    "pending": 50,
    "processing": 100,
    "completed": 1400,
    "failed": 47
  },
  "results": [...]
}
  • total: Total number of results in the batch (across all pages)
  • status_counts: Breakdown of all results in the batch by status (not just the current page)

Note: The total and status_counts fields are only available when filtering by batch_id. Global list responses (without batch_id) return only count (items on this page) — there is no total field for the global corpus.

DNS/RDAP Live Lookup Tools

Live DNS and RDAP lookup tools return current records as published on the internet.

Always synchronous — no polling required. Unlike datapulse_scrape_submit (which defaults to async), these tools return full results directly in the response. There is no wait parameter and no need to call datapulse_scrape_result afterward.

Bulk variants: For multi-domain submission (up to 1000 per request), use datapulse_live_dns_bulk or datapulse_live_rdap_bulk. These are async — they return job receipts, not results.

datapulse_live_dns

Live DNS lookup — returns current records as published on the internet in a single synchronous call.

Resolver: DNSSEC-validating recursive, no content filtering except that answers in private, link-local or unique-local ranges are dropped (DNS-rebinding protection; see datapulse_help(topic="dns")). The specific resolver is not exposed or configurable.

Parameters

ParameterTypeRequiredDescription
domainstringYesDomain to look up
record_typestringNoSpecific record type (e.g., SRV, PTR, TLSA). If omitted, queries 10 default types
fullbooleanNoReturn full raw DNS output (default: false — returns compact summary)
forcebooleanNoForce a fresh lookup (default: false). DNS lookups are always live — cached only at the resolver layer — so this is rarely needed
metadataobjectNoCustom key-value metadata

Example

{"domain": "example.com"}

By default returns a compact summary with status, dnssec_signed, and records grouped by type (A, AAAA, MX, etc.), with DNSSEC metadata stripped. Queries 10 record types (A, AAAA, MX, NS, TXT, SOA, CNAME, CAA, DNSKEY, DS) plus a _dmarc.<domain> TXT lookup (returned in the dmarc field) when record_type is omitted. Use record_type for specific types (SRV, PTR, TLSA, etc.).

Set full=true for the raw DNS output including per-query flags, RRSIG records, TTL, authorities, and additionals.

Caching: DNS lookups are always live — results are cached only at the DNS resolver layer, not the application layer — so force=true is rarely needed (datapulse_live_dns_bulk always reports cached: 0 for the same reason).

datapulse_live_rdap

Live RDAP lookup — returns current registration data as published on the internet in a single synchronous call. Supports domain names and IP addresses (IPv4 and IPv6). The query parameter accepts any of these (e.g., example.com, 8.8.8.8, 2601:249:8080:5296::35).

For full documentation — parameters, the {query, type, datetime, results} envelope and field reference, lookup errors, caching behavior, and how to interpret RDAP output — see datapulse_help(topic="rdap").

Workflow Patterns

datapulse_domain_overview(domain="example.com")
→ Done. Web scrape + DNS + RDAP in one call.

Single URL — use sync mode

datapulse_scrape_submit(url="example.com", wait=true)
→ Done. Full result with extraction in one call.

Single URL with DOM analysis

1. datapulse_scrape_submit(url="example.com", wait=true)
   → Get url_hash from the response

2. datapulse_scrape_result(url_hash="...", include_dom=true)
   → Get full result with cleaned HTML DOM

Bulk scraping — use batch_id to track progress

Background work only: bulk scrapes run at lower priority, can take hours during large background runs, and pause while the extraction model is unavailable (see Queue Priority).

1. datapulse_scrape_bulk(urls=[...up to 1000...])
   → Returns batch_id + url_hashes array

2. Wait at least 5 seconds

3. Check batch progress:
   datapulse_scrape_list(batch_id="...")
   → Returns status_counts showing how many are pending/processing/completed/failed
   → Paginate through results with limit/offset
   → If pending/processing remain, repeat with backoff (5s → 10s → 20s …,
     capping at a few minutes). A batch that stops moving is paused or
     queued behind interactive work, not failed; do not resubmit it

4. For individual results with DOM:
   datapulse_scrape_result(url_hash="...", include_dom=true)

5. If one URL becomes urgent while the batch is still pending:
   datapulse_scrape_submit(url="...", force=true, wait=true)
   → force=true is required: without it, wait=true waits for the queued bulk job and usually times out

Efficient multi-hash retrieval

Use datapulse_scrape_bulk_results to fetch many results in a single call. Works after any bulk submission — pass the url_hashes array directly:

1. datapulse_scrape_bulk(urls=[...])
   → Returns url_hashes=[...] covering all submitted URLs (accepted + cached)

2. datapulse_scrape_bulk_results(url_hashes=[...up to 1000...])
   → Returns all results + not_found list in one call

This is especially useful when most or all URLs are cached (accepted=0, cached=N), since batch_id only covers accepted URLs. But it works equally well for mixed batches or any set of known hashes.

Force Fresh Scrape

Use force: true to bypass the 90-day cache:

{"url": "example.com", "force": true, "wait": true}

Live DNS + RDAP Lookup

datapulse_live_dns(domain="example.com")
→ Returns comprehensive DNS records

datapulse_live_rdap(query="example.com")
→ Returns RDAP registration data (domains)

datapulse_live_rdap(query="8.8.8.8")
→ Returns RDAP IP network data (IPv4)

datapulse_live_rdap(query="2601:249:8080:5296::35")
→ Returns RDAP IP network data (IPv6)

DOM Comparison for Impersonation Detection

When investigating whether a suspicious domain is impersonating a legitimate brand, scrape BOTH domains, then retrieve with include_dom=true. Compare DOM structures: similar page layouts, copied text, replicated forms, matching CSS/JS patterns, and reused images or branding elements are strong evidence of impersonation.

1. datapulse_scrape_submit(url="huntington.com", wait=true)
2. datapulse_scrape_submit(url="huntingtonbankonline.com", wait=true)
3. datapulse_scrape_result(url_hash="...", include_dom=true) for each

Compare:
- Page structure and layout patterns
- Copied text (slogans, disclaimers, product descriptions)
- Form fields and input patterns
- CSS class names and styling approaches
- Image URLs and branding assets
- JavaScript patterns and functionality

Rate Limits

  • 100 requests per second per IP (configurable server-side)
  • Burst: 200 requests
  • Sync mode (wait=true) is additionally per-IP rate limited
  • 429 responses include Retry-After: 1 header

Error Handling

Every API failure is returned as an error carrying the API’s HTTP status and its response body verbatim; nothing is reworded by this server.

StatusCause
429Too many requests, wait 1 second
401Invalid API token
404Unknown url_hash/domain_hash or job not yet created

Generated from the live server (DataPulse MCP 1.0.0) on October 1, 2026. Your AI assistant reads this page by calling datapulse_help(topic="scraping").