Web Scraping & Live Lookup Tools
DataPulse provides tools to submit URLs for scraping, perform live DNS/RDAP lookups, and retrieve structured results.
Architecture: Pages are rendered with a headless browser, then an LLM extracts structured entity data. Results are cached for 90 days.
START HERE: Domain Overview
For most domain research, use datapulse_domain_overview — it handles everything in one call.
datapulse_domain_overview(domain="osu.edu")
This single call:
- Runs a live DNS lookup (all record types, DNSSEC-validated)
- Fetches RDAP registration data (registrar, dates, status)
- Scrapes the website with headless browser + LLM entity extraction
- Returns the combined result: scrape + DNS + RDAP
First call for a new domain takes 15-30 seconds (jobs must complete). Subsequent calls return cached results instantly. Use force=true to re-fetch everything.
When to use the other tools instead:
- Bulk operations (100+ domains) where nobody is waiting on the answer: use
datapulse_scrape_bulk+datapulse_live_dns_bulk+datapulse_live_rdap_bulk. Bulk scrapes run at lower priority — see Queue Priority. - DOM analysis: use
datapulse_scrape_result(url_hash=..., include_dom=true)after the overview - Specific URL paths:
datapulse_domain_overviewscrapes the root domain; for/about,/contact, etc. usedatapulse_scrape_submit - DNS-only or RDAP-only: use
datapulse_live_dnsordatapulse_live_rdapif you only need one piece
Queue Priority: Interactive vs Bulk Scrapes
Web scrapes are processed in two classes:
| Submitted via | Class | When it runs |
|---|---|---|
datapulse_scrape_submit, datapulse_domain_overview | Interactive | Processed first — every worker keeps a scrape slot free for it |
datapulse_scrape_bulk | Bulk | Lower priority. Processed only when a worker has a spare slot, behind all interactive work. During large background runs it can take hours |
If someone is waiting on the answer, do not use datapulse_scrape_bulk. Call datapulse_scrape_submit(url=..., wait=true) once per URL, or datapulse_domain_overview(domain=...), instead — both are processed ahead of datapulse_scrape_bulk submissions. Use bulk for background work whose results can arrive later.
This applies to web scrapes only. datapulse_live_dns_bulk and datapulse_live_rdap_bulk share one queue with their single-lookup forms and are not deprioritised.
Bulk work can pause
Bulk scrapes also pause while the extraction model is unavailable. Bulk jobs wait in the queue, or are retried automatically, and resume on their own when it recovers. A bulk batch can therefore stall for as long as an outage lasts, on top of waiting behind interactive work. Interactive submissions are not affected: datapulse_scrape_submit and datapulse_domain_overview keep being processed throughout.
Polling a bulk batch
- A long-pending result is not a failure.
pendingorprocessingmeans the URL is still queued or in progress, however long it has been. Onlystatus: "failed"is a failure. - Poll with backoff. Wait at least 5 seconds, then double the interval each time (5s → 10s → 20s → 40s …), capping at a few minutes. For a large batch, check
status_countsfromdatapulse_scrape_list(batch_id=...), or fetch many hashes in one call withdatapulse_scrape_bulk_results, instead of polling everyurl_hash. - Do not resubmit the batch to unstick it. Resubmitting with
datapulse_scrape_bulkuses the same lower-priority queue, so it will not run the URLs any sooner. - If one URL becomes urgent, resubmit it interactively with
datapulse_scrape_submit(url=..., force=true, wait=true).force=trueis required: without it, a URL still queued from the bulk batch returnsin_flightand stays behind the bulk work.
Sync vs Async Modes
The individual scrape tools have three modes of operation. Understanding these prevents confusion — but if you’re using datapulse_domain_overview, you don’t need to worry about this.
Sync Mode (RECOMMENDED for single URLs)
Use wait=true on datapulse_scrape_submit. The call blocks until the scrape completes (up to 50 seconds) and returns the full result with extraction data in a single call. No polling needed.
datapulse_scrape_submit(url="example.com", wait=true)
→ Returns: complete result with status, extraction data, timing — all in one response
Use sync mode when: scraping 1 URL at a time. This is the simplest path.
Note: The sync response has the same shape as datapulse_scrape_result (including final_url, redirected, tls, quic, server_addr, and link_domains and links unless link_data narrows them). Only cleaned_dom is unavailable here — use datapulse_scrape_result(url_hash=..., include_dom=true) after the sync call completes if you need the DOM.
Async Mode (for bulk or background work)
Without wait=true, datapulse_scrape_submit returns immediately with only a job receipt — NOT the scrape result. You must then poll separately:
Step 1: datapulse_scrape_submit(url="example.com")
→ Returns ONLY: {status: "queued", job_id, domain_hash, url_hash}
→ The page has NOT been scraped yet
Step 2: datapulse_scrape_result(url_hash="...")
→ If status="pending" or "processing" → NOT DONE, poll again in 5-10 seconds
→ If status="completed" → extraction data is ready
→ If status="failed" → check error field
Use async mode when: submitting many URLs via datapulse_scrape_bulk, or when you need to submit and check results later.
Async Status Values (datapulse_scrape_submit only)
In async mode, the response status field indicates what happened:
| Status | Meaning | Action |
|---|---|---|
queued | New job accepted | Poll with datapulse_scrape_result(url_hash=...) |
cached | Previously cached, full result returned inline | No polling needed — result is complete |
in_flight | Already being processed | Poll with datapulse_scrape_result(url_hash=...) |
failed | Scrape failed, returns error and full result inline | No polling needed — error detail is in the response |
How to detect an inline result: When the response contains fields like extraction or completed_at, the result is complete (from cached or failed status). When the response only contains status, job_id, domain_hash, url_hash, it’s an async receipt that requires polling.
Note: This inline result behavior only applies to datapulse_scrape_submit. The datapulse_scrape_bulk tool always returns aggregate counts (accepted/rejected/cached) — it does not return per-URL results inline. To retrieve cached results for bulk-submitted URLs, use datapulse_scrape_bulk_results(url_hashes=[...]) for efficient batch retrieval, or datapulse_scrape_result(url_hash=...) for individual results.
Common Mistakes
| Mistake | Fix |
|---|---|
| Treating the async submit response as the scrape result | Without wait=true, submit only returns a receipt ({status, job_id, domain_hash, url_hash}). The actual content is retrieved via datapulse_scrape_result. Exception: cached and failed statuses return the full result inline. |
| Using async mode for a single URL | Use wait=true — it returns the full result in one call, no polling needed. |
Using datapulse_scrape_bulk for interactive work | Bulk scrapes run at lower priority, can take hours during large background runs, and pause while the extraction model is unavailable. When someone is waiting, call datapulse_scrape_submit(wait=true) per URL, or datapulse_domain_overview. |
| Treating a long-pending bulk result as failed | A bulk-submitted URL can stay pending or processing for hours and still complete. Keep polling with backoff; only status: "failed" is a failure. If one URL is urgent, resubmit it with datapulse_scrape_submit(url=..., force=true, wait=true). See Polling a bulk batch. |
| Not polling after async submit | After async submit, you MUST call datapulse_scrape_result(url_hash=...) to get actual results. Exception: cached status returns results inline. |
| Polling immediately without waiting | Scrapes take 5-30 seconds. Wait at least 5 seconds before the first poll. Use exponential backoff: 5s → 10s → 20s. |
Constructing url_hash yourself | Always use the url_hash from the submit response. URL normalization rules are non-trivial. |
Using force on every request | Only use force when you need a fresh scrape. Without it, cached results return instantly. |
Requesting include_dom by default | DOM is 10KB-1MB+ and consumes context window. Only request it for specific analysis like impersonation detection. |
| Expecting batch_id to include all submitted URLs | Cached URLs (already cached/in-flight) are excluded from the batch. Only newly accepted URLs appear in datapulse_scrape_list(batch_id=...) results. To fetch all results (including cached), pass the url_hashes array from the bulk response to datapulse_scrape_bulk_results(url_hashes=[...]) — it covers all submitted URLs regardless of accepted/cached status. |
| Using an invalid batch_id | Malformed batch_id (not a valid UUID) returns 400 Bad Request. A valid UUID that doesn’t match any batch returns 404 Not Found. Verify your batch_id came from a datapulse_scrape_bulk response. |
URL Normalization
The tool sends the URL as you gave it, with only the scheme and host case-folded
and the host in its A-label form (submitted_url echoes your input when that
changed it); the path, query and fragment go up byte for byte. Everything below
is what the API then does on its side for consistent hashing:
- Scheme is ignored for hashing —
http://andhttps://produce the same hash - URLs without a scheme default to
https://for fetching; if port 443 is unreachable but port 80 is open, falls back tohttp://automatically www.prefix is stripped from hostname- Hostname is lowercased (paths are case-sensitive)
- Default ports are removed (443 for https, 80 for http)
- Query parameters are sorted alphabetically
- Fragment is stripped
- Trailing bare slash is removed
- Unicode/IDN hostnames are converted to punycode
Important: Paths ARE significant. example.com/about and example.com/contact are different scrapes.
These all produce the same url_hash:
https://www.Example.com/http://example.comwww.example.comexample.com
But these produce different url_hash values:
example.comvsexample.com/about(different path)example.com/page?b=2&a=1vsexample.com/page?a=1&b=2(same — query params sorted)
Note: For root URLs (no path), url_hash equals domain_hash — both hash the bare domain string "example.com". This is expected and harmless since they are used in different contexts.
IDN domains are automatically converted per IDNA 2008 (RFC 5890-5893):
münchen.de→xn--mnchen-3ya.de中国.cn→xn--fiqs8s.cn
Emoji domains are not supported. Emoji are DISALLOWED by IDNA 2008 (RFC 5892, Unicode category So). Both Unicode emoji (e.g. 💩.la) and pre-encoded punycode emoji (e.g. xn--ls8h.la) are rejected per SSAC Advisory SAC095. Some ccTLD registries accept emoji registrations, but these cannot be processed consistently by IDNA-conformant applications.
Scrape Tool Reference
datapulse_domain_overview
Recommended starting point for domain research. One call to get everything: web scrape + DNS + RDAP.
Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
domain | string | Yes | Domain name to investigate (e.g. “osu.edu”). Accepts Unicode/IDN. Its spelling is canonicalized and no labels are removed (www. kept); a single label (localhost) is rejected before any request. When canonicalizing changed more than letter case (an IDN folded to punycode, a trailing dot dropped), the response carries submitted_domain with what you sent. See datapulse_help(topic="normalization"). |
force | boolean | No | Re-fetch all data even if cached (default: false) |
include_dom | boolean | No | Include cleaned HTML DOM in scrape results (default: false) |
link_data | string | No | Which link data to return: none (default), domains, links, both |
No link data is returned unless you ask. The page’s links arrive in two shapes that differ in size by two orders of magnitude:
link_data | Returns | Typical size |
|---|---|---|
none (default) | nothing | — |
domains | link_domains | ~500 bytes |
links | the individual links | ~76 KB on a commercial site |
both | both | ~77 KB |
domains answers most link questions. links and both are returned WHOLE,
never truncated — a partial link list is misleading rather than merely
incomplete, because the link that mattered may be the one cut. An unrecognised
value is treated as none and the response says so, so a mistake costs data
rather than silently returning a shape you did not ask for.
Example
{"domain": "osu.edu"}
{"domain": "osu.edu", "link_data": "domains"}
Response
summary fields are conditional. Only resolves, dnssec and
registrar_name are always present. The rest — registrar_iana_id (a number,
the same value as rdap.ianaid), registration_date, expiration_date,
entity_name, entity_type, nature, country — are omitted when the source
had no value at all. An absent field means the source had no answer for it.
dnssec is one of signed, unsigned, broken (the validating resolver
answered SERVFAIL) or unknown. unknown appears together with dns.status: "error": the DNS lookup itself failed or returned nothing usable, so nothing was
observed. resolves is false in that case because no address was seen, not
because the name is known to be absent — check dns.status before treating a
false as a finding.
The four extraction fields are passed through exactly as the extractor wrote
them, including its placeholders: entity_name: "unknown", entity_type: "other" and country: "unknown" are the extractor declining to classify, not
classifications. Read them as “no answer”, never as a finding.
country is the extractor’s own wording for where the site says it operates
("USA", "United Kingdom"), not an ISO 3166 code. RDAP’s country on IP
lookups is a two-letter code; the two are not comparable as codes.
The response is optimized for token efficiency with a top-level summary and compact subsections:
{
"domain": "github.com",
"domain_hash": "a1b2c3d4...",
"summary": {
"resolves": true,
"dnssec": "unsigned",
"registrar_iana_id": 292,
"registrar_name": "MarkMonitor Inc.",
"_note": "registrar_name is \"unknown\" when RDAP failed and \"not assigned (registry without IANA IDs)\" when the registry publishes no IANA registrar IDs (Nominet .uk); registrar_iana_id is then absent",
"registration_date": "2007-10-09T18:20:50Z",
"expiration_date": "2026-10-09T18:20:50Z",
"entity_name": "GitHub, Inc.",
"entity_type": "commercial",
"nature": "Software development platform and code hosting service",
"country": "USA"
},
"scrape": {
"status": "completed",
"url": "https://github.com",
"final_url": "https://github.com/",
"redirected": true,
"tls": {
"protocol": "TLS 1.3",
"issuer": "DigiCert Inc",
"subject": "github.com",
"valid_from": "2025-02-13T00:00:00.000Z",
"valid_to": "2026-03-14T23:59:59.000Z"
},
"server_addr": "140.82.121.3:443",
"quic": {
"supported": true,
"alpn": "h3",
"tls_version": "TLS 1.3",
"server_addr": "140.82.121.3:443",
"handshake_ms": 45
},
"extraction": { "entity_name": "GitHub, Inc.", "entity_type": "commercial", ... },
"render_time_ms": 3200,
"llm_time_ms": 4100,
...
},
"dns": {
"status": "ok",
"dnssec_signed": false,
"records": {
"A": [{"name": "github.com.", "ttl": 60, "data": {"address": "140.82.121.3"}}],
"MX": [{"name": "github.com.", "ttl": 3600, "data": {"preference": 1, "exchange": "aspmx.l.google.com."}}],
"NS": [{"name": "github.com.", "ttl": 900, "data": {"nameserver": "dns1.p08.nsone.net."}}],
"TXT": [...]
},
"query_time_ms": 450
},
"rdap": {
"name": "github.com",
"researchdate": "2026-02-03T21:07:05Z",
"ianaid": 292,
"registrationdate": "2007-10-09T18:20:50Z",
"expirationdate": "2026-10-09T18:20:50Z",
"status": ["client delete prohibited", "client transfer prohibited"],
"dns": ["dns1.p08.nsone.net", "dns2.p08.nsone.net"],
"abuse_email": "abusecomplaints@markmonitor.com"
}
}
Response sections:
summary: Key facts at a glance — includes resolution status, DNSSEC state, authoritative registrar ID/name, registration dates, and extracted entity classification.scrape: Single scrape result object (not array). Same fields asdatapulse_scrape_result. Includesfinal_url(actual URL rendered after redirects/protocol fallback),redirected(boolean),tls(certificate details — absent for HTTP),quic(QUIC/HTTP3 support probe — absent when not probed),server_addr(IP:port), and, whenlink_dataasks for them,link_domainsand/orlinks(see the parameter table above — neither is returned by default). The page content (cleaned_dom) is likewise omitted unlessinclude_dom=true. Always present: a failed scrape carriesstatus: "failed"anderror; if the scrape has not finished yet, a placeholder withstatus: "processing"and anerrornote is returned whilednsandrdapare still populated.dns: Summarized DNS — answer records only, no RRSIG/NSEC3/authority sections. For NXDOMAIN domains, returnsstatus: "nxdomain"with just the SOA. For full raw DNS output, usedatapulse_live_dns.rdap: The RDAP envelope as the API returned it, names and codes case-folded.resultsholds every observation in the API’s order; each has name, ianaid, registrationdate, expirationdate (optional), status, dns, abuse_email, abuse_phone. Null when there is no envelope, withrdap_statusandrdap_errorsaying what the RDAP row reported. Note the expiry field isexpirationdatehere butexpiration_dateinsummaryand indatapulse_dns_rdap_history.
Individual job failures are non-fatal — the response includes whatever succeeded.
datapulse_scrape_submit
Submit a single URL for scraping. This is the tool for interactive use — single submissions are processed ahead of datapulse_scrape_bulk submissions. Use wait=true for single URLs (recommended).
Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
url | string | Yes | URL to scrape (protocol optional) |
wait | boolean | No | Recommended. Sync mode: block until complete, return full result. Default: false |
wait_timeout | integer | No | Max seconds to wait in sync mode (default: 50, max: 50) |
force | boolean | No | Bypass 90-day cache (default: false) |
metadata | object | No | Custom key-value metadata stored with result (returned as metadata in all result responses) |
link_data | string | No | wait=true only. Which link structures the sync result carries: none, domains (link_domains only), links (the links array only) or both (default). Unlike the overview, the sync result returns both by default; pass domains when you do not need every URL, since links can run to hundreds of KB. An unrecognised value is refused. Ignored without wait=true: an async receipt carries no link data, and datapulse_scrape_result always returns both. |
Sync mode example (recommended)
{"url": "example.com", "wait": true}
Response — the completed result:
{
"status": "completed",
"domain": "example.com",
"url_hash": "a379a6f6...",
"extraction": {
"entity_name": "IANA",
"entity_type": "other",
"nature": "Example domain",
"language": "eng",
"location_country": "USA",
...
}
}
If the sync timeout expires before the scrape finishes, the tool returns a timeout error. To recover, resubmit without wait=true — the response will be status: "in_flight" (with url_hash), then poll with datapulse_scrape_result(url_hash=...).
Async mode example
{"url": "example.com"}
Response — just a job receipt (NOT the scrape result):
{
"status": "queued",
"job_id": "550e8400-e29b-41d4-a716-446655440000",
"domain_hash": "a379a6f6eeafb9a55e378c118034e2751e682fab9f2d30ab13d2125586ce1947",
"url_hash": "a379a6f6eeafb9a55e378c118034e2751e682fab9f2d30ab13d2125586ce1947"
}
status:queued(new),cached(cached result returned inline),in_flight(already being processed), orfailed(failed, result inline)job_id: Job identifier (UUID)domain_hash: SHA-256 hash of the normalized domain, computed with the shared dpdomain convention (www. stripped, lowercase, punycode, no trailing dot) — the same valuedatapulse_domain_overviewanddatapulse_live_domain_healthreturnurl_hash: SHA-256 hash of the normalized URL — use withdatapulse_scrape_resultto poll for actual results
Inline results: If status is cached or failed, the full result is returned inline — no polling needed. All statuses include url_hash and domain_hash; receipts (queued, in_flight) carry job_id, while inline results carry the job ID as id. For in_flight, only IDs are returned — poll with datapulse_scrape_result.
datapulse_scrape_bulk
Submit multiple URLs in a single request (up to 1000). Always async — poll results individually or use batch_id to retrieve the whole batch.
Lower priority. Bulk submissions are processed behind single submissions (datapulse_scrape_submit, datapulse_domain_overview) and can take hours during large background runs. They also pause while the extraction model is unavailable and resume automatically when it recovers, so some results can stay pending or processing much longer than usual (see Polling a bulk batch). For interactive or time-sensitive use — someone waiting on the answer — call datapulse_scrape_submit(wait=true) per URL, or datapulse_domain_overview, instead.
Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
urls | array | Yes | Array of URL objects (max 1000) |
Each URL object can have: url (required), force, metadata
Example
{
"urls": [
{"url": "example.com"},
{"url": "another.com", "force": true},
{"url": "third.com", "metadata": {"campaign": "q1"}}
]
}
Response
Returns job receipts with a batch_id for tracking:
{
"accepted": 2,
"rejected": 0,
"cached": 1,
"batch_id": "f47ac10b-58cc-4372-a567-0e02b2c3d479",
"job_ids": ["550e8400-...", "550e8400-..."],
"domain_hashes": ["d1e2f3a4...", "b5c6d7e8..."],
"url_hashes": ["a379a6f6...", "a1b2c3d4...", "c3d4e5f6..."],
"errors": [],
"rejected_urls": [],
"rejected_local": [],
"sent": [
{"index": 0, "url": "https://example.com/"},
{"index": 1, "url": "https://another.com/"},
{"index": 2, "url": "https://third.com/"}
]
}
Every URL is validated before anything is sent, and one bad URL never costs you
the batch. Every index is the position in your urls array:
accepted,cachedandrejectedare the API’s own counts over the URLs that were sent.rejected_urls(index,url,error) lists what the API refused, with its code (for exampleinvalid_domain).rejected_local(index,url,error) lists what was refused before the request — a blank item,ftp:orjavascript:, credentials, a scheme-relative URL — with the reason. These are not counted inrejected, because the API never saw them. To find every failure, read both lists.sentis every URL that was sent, in the form it was sent. Sent is not accepted: a URL the API refuses is insentand inrejected_urls.- Every input is accounted for:
accepted + cached + rejected + len(rejected_local)equals the number of URLs you passed.
Key: url_hashes contains hashes for all submitted URLs (accepted + cached), not just accepted ones. This array is what you pass to datapulse_scrape_bulk_results to fetch results. In the example above, 3 URLs were submitted → 3 hashes returned, even though only 2 were accepted.
After bulk submit:
- Use
datapulse_scrape_bulk_results(url_hashes=[...])to fetch all results (accepted + cached) in a single call — pass theurl_hashesarray from this response directly. - Use
datapulse_scrape_list(batch_id="...")to paginate through newly accepted results with status summary counts. - Use
datapulse_scrape_result(url_hash="...")for individual results or when you needinclude_dom=true.
Important: cached handling. URLs that are already cached or in-flight are counted as cached and are not included in the batch_id tracking. Only newly accepted URLs appear when querying by batch_id. Cached hashes are still included in the url_hashes response array — use datapulse_scrape_bulk_results(url_hashes=[...]) to fetch all results regardless of whether they were accepted or cached.
datapulse_scrape_result
Retrieve a scrape result by url hash. Use this to poll for async results, or to retrieve results with DOM.
Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
url_hash | string | Yes | 64-char SHA-256 hash from submit response |
include_dom | boolean | No | Include cleaned HTML DOM (see warning below) |
Status Values
| Status | Meaning | Action |
|---|---|---|
pending | Job waiting in queue | Poll again in 5-10 seconds. A bulk-submitted URL can stay pending for hours without having failed — back off (see Polling a bulk batch) |
processing | Currently being scraped | Poll again in 5-10 seconds, backing off if it persists |
completed | Successfully scraped and extracted | Extraction data is ready |
cached | Served from storage by a sync datapulse_scrape_submit call (the row itself stays completed) | Result is inline; cached: true says the scrape did not run for this call |
failed | Error occurred | Check error field |
Warning: DOM Size and Context Windows
The include_dom=true option returns the full cleaned HTML DOM. It is now
requested from the API rather than filtered on arrival, so the bytes are never
built unless you ask:
- Simple pages: 10-50KB
- Average pages: 50-200KB
- Complex pages: 200KB-1MB+
This can consume significant context window space. Only request DOM when needed for specific analysis tasks like impersonation detection or content comparison. For most use cases, the extraction data provides the key information without the DOM overhead.
MCP transport overflow: Results exceeding MCP context limits may be stored externally by the MCP client — you may receive a file reference rather than inline content. This is an MCP protocol-level behavior, not a DataPulse API limitation.
Response Fields
Each result includes the following fields:
| Field | Type | Description |
|---|---|---|
id | string | Unique job ID (UUID) |
domain | string | Domain extracted from URL |
domain_hash | string | SHA-256 hash of the normalized domain (dpdomain convention: www. stripped, lowercase, punycode) |
url | string | Normalized URL |
url_hash | string | SHA-256 hash of normalized URL |
status | string | pending, processing, completed, failed; a sync submit reports cached for a stored row |
cached | boolean | true when the row was served from storage rather than scraped for this call. Set by datapulse_scrape_submit (sync) and datapulse_scrape_list; the lifecycle status is unchanged |
final_url | string | Actual URL rendered after redirects and protocol fallback. Compare with url to see what changed. |
redirected | boolean | true when final_url differs from the requested URL |
tls | object | TLS certificate details: protocol (“TLS 1.3”), issuer (CA name), subject (cert CN), valid_from, valid_to (ISO 8601). Absent for HTTP connections. |
quic | object | QUIC/HTTP3 support probe. supported: true with alpn (e.g. “h3”), tls_version, server_addr (IP:port) and handshake_ms; supported: false with error (a timeout means the host does not answer QUIC on UDP 443 — most hosts without HTTP/3 simply drop it) and, on rows scraped since 2026-09-15, reason (timeout, tls_rejected, resolve_failed, version_negotiation, stateless_reset, transport_error, application_error, listen_failed, invalid_args or other); a tls_rejected failure adds tls_alert (e.g. “handshake failure”) and tls_alert_code (e.g. 40), which is what a CDN with HTTP/3 disabled returns. No handshake_ms on a failure. Absent when the probe could not run. |
server_addr | string | IP:port of the server that responded (e.g. “140.82.121.3:443”) |
error | string | Error message (if failed) |
created_at | datetime | When job was first submitted |
completed_at | datetime | When processing finished for this submission |
last_scraped_at | datetime | When the URL was last successfully scraped (may differ from completed_at if result was served from cache — last_scraped_at reflects the original scrape time, completed_at reflects the current job completion) |
render_time_ms | int | Browser rendering time in ms |
llm_time_ms | int | LLM processing time in ms |
total_time_ms | int | Total end-to-end processing time in ms |
batch_id | string | Batch UUID if the URL was submitted via datapulse_scrape_bulk (absent for individual submissions) |
metadata | object | Custom key-value metadata from submission (always present, {} when not set). Shared, not per caller: the corpus is one store for every caller of this deployment, so metadata written by one caller is returned to any caller who retrieves the row, and a cache hit returns the metadata of whoever first scraped the URL. Never put credentials, personal data or anything confidential in it |
extraction | object | Structured extracted data (see below). Omitted for non-HTML content (PDFs): a completed result with no extraction and no llm_time_ms is a PDF or other non-HTML document; pending/processing is “not yet available”. |
crawler_policies | object | Robot/AI crawler policy files discovered for the domain (robots_txt, llms_txt, ai_txt, ai_json — each a string or null). They are collected and returned verbatim, not applied: a Disallow in robots_txt does not stop the scrape. Always present (as {} when no policy files found) in datapulse_scrape_result, datapulse_scrape_bulk_results and datapulse_scrape_list; datapulse_domain_overview omits it when empty. |
link_domains | object | The URLs referenced by the rendered page, split by ownership. Links whose host is an IP literal are in links only: an address is not a registrable domain. The two halves hold different units. internal holds full HOSTNAMES sharing the page’s registrable domain (wholesale.example.com) — you want to know which subdomains a site runs. external holds distinct third-party REGISTRABLE DOMAINS (klaviyo.com), deduplicated to the entity rather than every CDN hostname — the portfolio-discovery signal. Do not deduplicate or count across the two: they are not the same kind of thing. |
links | array | URLs referenced by the rendered page (includes JS-inserted links), up to 1500 records of {url, source, rel?, text?} where source is one of a, link, script, iframe, img, media, form, jsonld, meta. Returned by this tool, by datapulse_scrape_submit with wait=true (by default; link_data can narrow it), and by datapulse_domain_overview when link_data is links or both. List and bulk responses carry link_domains only. |
cleaned_dom | string | Cleaned HTML DOM (only with include_dom=true) |
Link discovery availability: links and link_domains are populated for scrapes performed on or after 2026-07-02. Results scraped earlier lack both fields entirely — absence means “scraped before the feature existed”, not “page has no links” (a genuinely link-free page returns links: []). Re-scrape with force=true to populate them. PDF results have no rendered DOM, so both fields are always absent for PDFs.
Extraction Data
When status is completed, the extraction object contains LLM-extracted data.
Note: For non-HTML content (PDFs, binary files), the extraction field is omitted entirely and so is llm_time_ms (zero values are not sent) — LLM analysis is skipped for these content types. links and link_domains are absent too, since nothing was rendered. PDF results include markdown_content via the full API but this field is not currently surfaced through the MCP tool responses.
| Field | Type | Description |
|---|---|---|
is_error | yes/no | Whether page was an error page |
entity_name | string | Organization/entity name |
entity_type | enum | commercial, non-profit, government, personal, education, other |
nature | string | Page nature classification |
nature_adult | yes/no | Adult content indicator |
nature_domain | string | Domain purpose classification |
location | string | Physical location |
location_country | string | ISO 3166-1 alpha-3 country code |
phone | string | Contact phone number |
email | string | Contact email address |
domain_parked | yes/no | Whether domain is parked |
domain_random | yes/no | Whether domain appears randomly generated |
ecommerce | yes/no | E-commerce site indicator |
language | string | ISO 639-3 language code |
bot_detected | yes/no | Bot/CAPTCHA detection page identified |
Note: Boolean fields return "yes" or "no" strings, not true/false.
datapulse_scrape_bulk_results
Retrieve multiple scrape results by url_hash in a single request (up to 1000). Use this any time you have multiple url_hashes and want to fetch all results efficiently in one call instead of N individual datapulse_scrape_result calls.
Common use cases:
- After
datapulse_scrape_bulk: pass theurl_hashesarray from the bulk response directly to get all results (accepted + cached) in one call - Cross-batch queries: fetch results from multiple batches or submissions at once
- Any scenario where you have a set of known url_hashes
Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
url_hashes | array of strings | Yes | SHA-256 URL hashes to retrieve (max 1000) |
Each hash must be a 64-character hex string. Invalid hashes return 400 Bad Request.
Example
{
"url_hashes": ["a379a6f6...", "a1b2c3d4...", "deadbeef..."]
}
Response
{
"count": 2,
"results": [
{"url_hash": "a379a6f6...", "status": "completed", "extraction": {...}, ...},
{"url_hash": "a1b2c3d4...", "status": "pending", ...}
],
"not_found": ["deadbeef..."]
}
count: Number of results foundresults: Array of scrape results with the same fields asdatapulse_scrape_result(see Response Fields table above), excludingcleaned_dom,semantic_snapshot, thelinksarray, and the per-connection fieldsfinal_url,redirected,tls,quicandserver_addr— the bulk read is served from a lean projection; fetch a single result for those (link_domainsis included). Includeserrorfield on failed results,batch_idon batch-submitted results, andcrawler_policies.not_found: Array of url_hashes that had no matching result in storage
In-flight jobs: Hashes for URLs that are still pending or processing appear in results (with their current status), not in not_found. A hash only appears in not_found if no record exists at all — e.g., a hash that was never submitted or is simply invalid.
datapulse_scrape_list
List scrape results with pagination. Supports filtering by batch_id from bulk submissions. Each result includes an error field (present only when status is "failed") so you can identify failures without individual polling.
The listing is global, not per caller. It covers every row in this deployment’s corpus, including results and metadata submitted by other callers; nothing is scoped to the OAuth identity that made the call.
Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
limit | integer | No | Results per page (default: 100, max: 1000) |
offset | integer | No | Skip N results for pagination |
batch_id | string | No | Filter to a specific bulk submission batch (UUID from datapulse_scrape_bulk response) |
Example — list all results
{
"limit": 50,
"offset": 100
}
Example — list results for a batch
{
"batch_id": "f47ac10b-58cc-4372-a567-0e02b2c3d479",
"limit": 200
}
When batch_id is provided, the response includes additional fields:
{
"count": 200,
"total": 1597,
"batch_id": "f47ac10b-58cc-4372-a567-0e02b2c3d479",
"status_counts": {
"pending": 50,
"processing": 100,
"completed": 1400,
"failed": 47
},
"results": [...]
}
total: Total number of results in the batch (across all pages)status_counts: Breakdown of all results in the batch by status (not just the current page)
Note: The total and status_counts fields are only available when filtering by batch_id. Global list responses (without batch_id) return only count (items on this page) — there is no total field for the global corpus.
DNS/RDAP Live Lookup Tools
Live DNS and RDAP lookup tools return current records as published on the internet.
Always synchronous — no polling required. Unlike datapulse_scrape_submit (which defaults to async), these tools return full results directly in the response. There is no wait parameter and no need to call datapulse_scrape_result afterward.
Bulk variants: For multi-domain submission (up to 1000 per request), use datapulse_live_dns_bulk or datapulse_live_rdap_bulk. These are async — they return job receipts, not results.
datapulse_live_dns
Live DNS lookup — returns current records as published on the internet in a single synchronous call.
Resolver: DNSSEC-validating recursive, no content filtering except that answers in private, link-local or unique-local ranges are dropped (DNS-rebinding protection; see datapulse_help(topic="dns")). The specific resolver is not exposed or configurable.
Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
domain | string | Yes | Domain to look up |
record_type | string | No | Specific record type (e.g., SRV, PTR, TLSA). If omitted, queries 10 default types |
full | boolean | No | Return full raw DNS output (default: false — returns compact summary) |
force | boolean | No | Force a fresh lookup (default: false). DNS lookups are always live — cached only at the resolver layer — so this is rarely needed |
metadata | object | No | Custom key-value metadata |
Example
{"domain": "example.com"}
By default returns a compact summary with status, dnssec_signed, and records grouped by type (A, AAAA, MX, etc.), with DNSSEC metadata stripped. Queries 10 record types (A, AAAA, MX, NS, TXT, SOA, CNAME, CAA, DNSKEY, DS) plus a _dmarc.<domain> TXT lookup (returned in the dmarc field) when record_type is omitted. Use record_type for specific types (SRV, PTR, TLSA, etc.).
Set full=true for the raw DNS output including per-query flags, RRSIG records, TTL, authorities, and additionals.
Caching: DNS lookups are always live — results are cached only at the DNS resolver layer, not the application layer — so force=true is rarely needed (datapulse_live_dns_bulk always reports cached: 0 for the same reason).
datapulse_live_rdap
Live RDAP lookup — returns current registration data as published on the internet in a single synchronous call. Supports domain names and IP addresses (IPv4 and IPv6). The query parameter accepts any of these (e.g., example.com, 8.8.8.8, 2601:249:8080:5296::35).
For full documentation — parameters, the {query, type, datetime, results} envelope and field reference, lookup errors, caching behavior, and how to interpret RDAP output — see datapulse_help(topic="rdap").
Workflow Patterns
Domain research — use domain overview (RECOMMENDED)
datapulse_domain_overview(domain="example.com")
→ Done. Web scrape + DNS + RDAP in one call.
Single URL — use sync mode
datapulse_scrape_submit(url="example.com", wait=true)
→ Done. Full result with extraction in one call.
Single URL with DOM analysis
1. datapulse_scrape_submit(url="example.com", wait=true)
→ Get url_hash from the response
2. datapulse_scrape_result(url_hash="...", include_dom=true)
→ Get full result with cleaned HTML DOM
Bulk scraping — use batch_id to track progress
Background work only: bulk scrapes run at lower priority, can take hours during large background runs, and pause while the extraction model is unavailable (see Queue Priority).
1. datapulse_scrape_bulk(urls=[...up to 1000...])
→ Returns batch_id + url_hashes array
2. Wait at least 5 seconds
3. Check batch progress:
datapulse_scrape_list(batch_id="...")
→ Returns status_counts showing how many are pending/processing/completed/failed
→ Paginate through results with limit/offset
→ If pending/processing remain, repeat with backoff (5s → 10s → 20s …,
capping at a few minutes). A batch that stops moving is paused or
queued behind interactive work, not failed; do not resubmit it
4. For individual results with DOM:
datapulse_scrape_result(url_hash="...", include_dom=true)
5. If one URL becomes urgent while the batch is still pending:
datapulse_scrape_submit(url="...", force=true, wait=true)
→ force=true is required: without it, wait=true waits for the queued bulk job and usually times out
Efficient multi-hash retrieval
Use datapulse_scrape_bulk_results to fetch many results in a single call. Works after any bulk submission — pass the url_hashes array directly:
1. datapulse_scrape_bulk(urls=[...])
→ Returns url_hashes=[...] covering all submitted URLs (accepted + cached)
2. datapulse_scrape_bulk_results(url_hashes=[...up to 1000...])
→ Returns all results + not_found list in one call
This is especially useful when most or all URLs are cached (accepted=0, cached=N), since batch_id only covers accepted URLs. But it works equally well for mixed batches or any set of known hashes.
Force Fresh Scrape
Use force: true to bypass the 90-day cache:
{"url": "example.com", "force": true, "wait": true}
Live DNS + RDAP Lookup
datapulse_live_dns(domain="example.com")
→ Returns comprehensive DNS records
datapulse_live_rdap(query="example.com")
→ Returns RDAP registration data (domains)
datapulse_live_rdap(query="8.8.8.8")
→ Returns RDAP IP network data (IPv4)
datapulse_live_rdap(query="2601:249:8080:5296::35")
→ Returns RDAP IP network data (IPv6)
DOM Comparison for Impersonation Detection
When investigating whether a suspicious domain is impersonating a legitimate brand, scrape BOTH domains, then retrieve with include_dom=true. Compare DOM structures: similar page layouts, copied text, replicated forms, matching CSS/JS patterns, and reused images or branding elements are strong evidence of impersonation.
1. datapulse_scrape_submit(url="huntington.com", wait=true)
2. datapulse_scrape_submit(url="huntingtonbankonline.com", wait=true)
3. datapulse_scrape_result(url_hash="...", include_dom=true) for each
Compare:
- Page structure and layout patterns
- Copied text (slogans, disclaimers, product descriptions)
- Form fields and input patterns
- CSS class names and styling approaches
- Image URLs and branding assets
- JavaScript patterns and functionality
Rate Limits
- 100 requests per second per IP (configurable server-side)
- Burst: 200 requests
- Sync mode (
wait=true) is additionally per-IP rate limited - 429 responses include
Retry-After: 1header
Error Handling
Every API failure is returned as an error carrying the API’s HTTP status and its response body verbatim; nothing is reworded by this server.
| Status | Cause |
|---|---|
| 429 | Too many requests, wait 1 second |
| 401 | Invalid API token |
| 404 | Unknown url_hash/domain_hash or job not yet created |
Generated from the live server (DataPulse MCP 1.0.0) on October 1, 2026. Your AI assistant reads this page by calling datapulse_help(topic="scraping").