datapulse_scrape_submit
datapulse:scrapeOutput: JSONRead-onlySubmit URL for scraping
Description
Submit a URL for scraping and LLM extraction.
INTERACTIVE: Use this tool for interactive work (someone waiting on the answer). Single submissions are processed ahead of datapulse_scrape_bulk submissions, and keep being processed while bulk work is paused.
TIP: For domain research, prefer datapulse_domain_overview(domain=…) which combines scrape + DNS + RDAP in one call. Use this tool for specific URL paths or when you only need the web scrape.
IMPORTANT: This tool has two modes. Choose the right one:
SYNC MODE (recommended for single URLs): Set wait=true. The call blocks until the scrape completes and returns the FULL result with extraction data in a single call. No polling needed.
ASYNC MODE (default — for bulk/background work): Without wait=true, returns ONLY a job receipt {status, job_id, domain_hash, url_hash}. The page has NOT been scraped yet. You MUST then call datapulse_scrape_result(url_hash=…) to poll for the actual result. Wait at least 5 seconds before first poll. Scrapes take 5-30 seconds.
STATUS VALUES (async mode):
- “queued”: New job accepted, poll with datapulse_scrape_result(url_hash=…)
- “cached”: Previously cached result returned inline with full extraction data. No polling needed.
- “in_flight”: Already being processed, returns IDs only. Poll for result.
- “failed”: Scrape failed. Returns the full result inline (status, url_hash, domain_hash, error, timestamps), same shape as datapulse_scrape_result. No polling needed.
CACHE HITS IN SYNC MODE: a URL scraped within the last 90 days is served from storage, with status “cached” and cached=true (a stored failure keeps status “failed” with cached=true). Only a scrape that ran for this call reports “completed”. last_scraped_at is the time of the stored scrape.
What is sent: your URL with the scheme and host case-folded and the host in a-label form; the path, query and fragment go up byte for byte, and submitted_url reports your input when that changed it. The API then normalizes for hashing on its side (scheme ignored, www. stripped, default ports removed, query params sorted, fragment stripped; paths preserved, so example.com/about and example.com/contact are different scrapes). Results are cached 90 days (use force=true to bypass). An invalid URL (unsupported scheme, no host, malformed host) is refused before any request.
For usage guidance, call datapulse_help(topic="scraping")
Parameters
| Parameter | Type | Description |
|---|---|---|
urlrequired | string | The URL to scrape. Protocol is optional (defaults to https://). http:// and https:// produce the same hash (scheme-agnostic). Min length: 1Max length: 2048 |
force | boolean | Bypass 90-day cache and force a fresh scrape. Default: false |
link_data | string | wait=true only. Which link structures the result carries: none, domains (link_domains only), links (the full links array only) or both (default). Pass domains when you do not need every URL — the links array can run to hundreds of KB. nonedomainslinksboth |
metadata | map<string, string> | Custom key-value metadata to store with the result. Stored with the row this submission produces; ignored on a cache hit (the stored row keeps the metadata of the submission that scraped it), and replaced by a force=true re-scrape. The corpus is shared by every caller of this deployment, so metadata is visible to anyone who retrieves the row; never put credentials or personal data here. |
wait | boolean | Synchronous mode: wait for scrape to complete before returning. Eliminates need for polling. Per-IP rate limited. Default: false |
wait_timeout | integer | Max seconds to wait in sync mode. Default 50, max 50. Only used when wait=true. Default: 50Min: 1Max: 50 |
Input schema (JSON)
{
"additionalProperties": false,
"properties": {
"force": {
"default": false,
"description": "Bypass 90-day cache and force a fresh scrape.",
"type": "boolean"
},
"link_data": {
"description": "wait=true only. Which link structures the result carries: none, domains (link_domains only), links (the full links array only) or both (default). Pass domains when you do not need every URL — the links array can run to hundreds of KB.",
"enum": [
"none",
"domains",
"links",
"both"
],
"type": "string"
},
"metadata": {
"additionalProperties": {
"type": "string"
},
"description": "Custom key-value metadata to store with the result. Stored with the row this submission produces; ignored on a cache hit (the stored row keeps the metadata of the submission that scraped it), and replaced by a force=true re-scrape. The corpus is shared by every caller of this deployment, so metadata is visible to anyone who retrieves the row; never put credentials or personal data here.",
"type": "object"
},
"url": {
"description": "The URL to scrape. Protocol is optional (defaults to https://). http:// and https:// produce the same hash (scheme-agnostic).",
"maxLength": 2048,
"minLength": 1,
"type": "string"
},
"wait": {
"default": false,
"description": "Synchronous mode: wait for scrape to complete before returning. Eliminates need for polling. Per-IP rate limited.",
"type": "boolean"
},
"wait_timeout": {
"default": 50,
"description": "Max seconds to wait in sync mode. Default 50, max 50. Only used when wait=true.",
"maximum": 50,
"minimum": 1,
"type": "integer"
}
},
"required": [
"url"
],
"type": "object"
}Generated from the live server (DataPulse MCP 1.0.0) on October 1, 2026.