MCP documentation menu

datapulse_scrape_submit

ScrapingScope: datapulse:scrapeOutput: JSONRead-only

Submit URL for scraping

Description

Submit a URL for scraping and LLM extraction.

INTERACTIVE: Use this tool for interactive work (someone waiting on the answer). Single submissions are processed ahead of datapulse_scrape_bulk submissions, and keep being processed while bulk work is paused.

TIP: For domain research, prefer datapulse_domain_overview(domain=…) which combines scrape + DNS + RDAP in one call. Use this tool for specific URL paths or when you only need the web scrape.

IMPORTANT: This tool has two modes. Choose the right one:

SYNC MODE (recommended for single URLs): Set wait=true. The call blocks until the scrape completes and returns the FULL result with extraction data in a single call. No polling needed.

ASYNC MODE (default — for bulk/background work): Without wait=true, returns ONLY a job receipt {status, job_id, domain_hash, url_hash}. The page has NOT been scraped yet. You MUST then call datapulse_scrape_result(url_hash=…) to poll for the actual result. Wait at least 5 seconds before first poll. Scrapes take 5-30 seconds.

STATUS VALUES (async mode):

  • “queued”: New job accepted, poll with datapulse_scrape_result(url_hash=…)
  • “cached”: Previously cached result returned inline with full extraction data. No polling needed.
  • “in_flight”: Already being processed, returns IDs only. Poll for result.
  • “failed”: Scrape failed. Returns the full result inline (status, url_hash, domain_hash, error, timestamps), same shape as datapulse_scrape_result. No polling needed.

CACHE HITS IN SYNC MODE: a URL scraped within the last 90 days is served from storage, with status “cached” and cached=true (a stored failure keeps status “failed” with cached=true). Only a scrape that ran for this call reports “completed”. last_scraped_at is the time of the stored scrape.

What is sent: your URL with the scheme and host case-folded and the host in a-label form; the path, query and fragment go up byte for byte, and submitted_url reports your input when that changed it. The API then normalizes for hashing on its side (scheme ignored, www. stripped, default ports removed, query params sorted, fragment stripped; paths preserved, so example.com/about and example.com/contact are different scrapes). Results are cached 90 days (use force=true to bypass). An invalid URL (unsupported scheme, no host, malformed host) is refused before any request.

For usage guidance, call datapulse_help(topic="scraping")

Parameters

ParameterTypeDescription
url
required
string
The URL to scrape. Protocol is optional (defaults to https://). http:// and https:// produce the same hash (scheme-agnostic).
Min length: 1Max length: 2048
forceboolean
Bypass 90-day cache and force a fresh scrape.
Default: false
link_datastring
wait=true only. Which link structures the result carries: none, domains (link_domains only), links (the full links array only) or both (default). Pass domains when you do not need every URL — the links array can run to hundreds of KB.
nonedomainslinksboth
metadatamap<string, string>
Custom key-value metadata to store with the result. Stored with the row this submission produces; ignored on a cache hit (the stored row keeps the metadata of the submission that scraped it), and replaced by a force=true re-scrape. The corpus is shared by every caller of this deployment, so metadata is visible to anyone who retrieves the row; never put credentials or personal data here.
waitboolean
Synchronous mode: wait for scrape to complete before returning. Eliminates need for polling. Per-IP rate limited.
Default: false
wait_timeoutinteger
Max seconds to wait in sync mode. Default 50, max 50. Only used when wait=true.
Default: 50Min: 1Max: 50
Input schema (JSON)
{
  "additionalProperties": false,
  "properties": {
    "force": {
      "default": false,
      "description": "Bypass 90-day cache and force a fresh scrape.",
      "type": "boolean"
    },
    "link_data": {
      "description": "wait=true only. Which link structures the result carries: none, domains (link_domains only), links (the full links array only) or both (default). Pass domains when you do not need every URL — the links array can run to hundreds of KB.",
      "enum": [
        "none",
        "domains",
        "links",
        "both"
      ],
      "type": "string"
    },
    "metadata": {
      "additionalProperties": {
        "type": "string"
      },
      "description": "Custom key-value metadata to store with the result. Stored with the row this submission produces; ignored on a cache hit (the stored row keeps the metadata of the submission that scraped it), and replaced by a force=true re-scrape. The corpus is shared by every caller of this deployment, so metadata is visible to anyone who retrieves the row; never put credentials or personal data here.",
      "type": "object"
    },
    "url": {
      "description": "The URL to scrape. Protocol is optional (defaults to https://). http:// and https:// produce the same hash (scheme-agnostic).",
      "maxLength": 2048,
      "minLength": 1,
      "type": "string"
    },
    "wait": {
      "default": false,
      "description": "Synchronous mode: wait for scrape to complete before returning. Eliminates need for polling. Per-IP rate limited.",
      "type": "boolean"
    },
    "wait_timeout": {
      "default": 50,
      "description": "Max seconds to wait in sync mode. Default 50, max 50. Only used when wait=true.",
      "maximum": 50,
      "minimum": 1,
      "type": "integer"
    }
  },
  "required": [
    "url"
  ],
  "type": "object"
}

Generated from the live server (DataPulse MCP 1.0.0) on October 1, 2026.