datapulse_scrape_bulk
datapulse:scrapeOutput: JSONRead-onlyBulk submit up to 1000 URLs (returns batch_id). Lower priority: can take hours or pause, so not for interactive use
Description
Submit multiple URLs for scraping in a single request (up to 1000).
LOWER PRIORITY: Bulk submissions run at lower priority than single submissions. They are processed behind interactive work and can take hours during large background runs. Bulk work also pauses while the extraction model is unavailable, and resumes automatically when it recovers, so some URLs can stay pending or processing much longer than usual. That is not a failure: keep polling with backoff, and do not resubmit the batch, which will not make it run sooner. For interactive or time-sensitive work (someone waiting on the answer), call datapulse_scrape_submit per URL (wait=true) or datapulse_domain_overview instead. To rush one URL already in a bulk batch, resubmit it with datapulse_scrape_submit(url=…, force=true, wait=true); without force=true it waits for the queued job and usually times out.
IMPORTANT: This tool is ALWAYS ASYNC. It returns job receipts only — NOT scrape results. You MUST poll each url_hash individually with datapulse_scrape_result(url_hash=…) to get actual results. Wait at least 5 seconds before first poll. A scrape takes 5-30 seconds once it starts, but may wait in the queue far longer first.
For DNS/RDAP lookups, use datapulse_live_dns_bulk or datapulse_live_rdap_bulk.
Returns:
- accepted/rejected/cached: Submission counts
- batch_id: UUID grouping all accepted URLs in this submission — use with datapulse_scrape_list(batch_id=…) to retrieve all results at once. NOTE: Cached URLs (already cached/in-flight) are excluded from the batch — only newly accepted URLs appear in batch results.
- url_hashes: Array of hashes for all submitted URLs (accepted + cached)
- domain_hashes: Array of domain hashes for tracking
- job_ids: Array of job IDs for tracking
- errors: the API’s own error codes for the URLs it rejected
- rejected_urls: one entry per input the API refused: index (position in your urls array), url as sent, and error code
- rejected_local: one entry per input refused before the request because it is blank or not a URL (unsupported scheme, no host, malformed host): index, url and error (the reason); these are NOT counted in rejected — read both lists to find every failure. One bad item never fails the batch
- sent: index and wire form (scheme and host case-folded, path untouched) of every URL that was submitted
For usage guidance, call datapulse_help(topic="scraping")
Parameters
| Parameter | Type | Description |
|---|---|---|
urlsrequired | object[] | Array of URLs to scrape. Maximum 1000 per request. Min items: 1Max items: 1000 |
urls[].urlrequired | string | The URL to scrape. Min length: 1 |
urls[].force | boolean | Bypass 90-day cache. |
urls[].metadata | map<string, string> | Custom key-value metadata stored with the result. Shared with every caller of this deployment and returned to anyone who retrieves the row; never put credentials or personal data here. |
Input schema (JSON)
{
"additionalProperties": false,
"properties": {
"urls": {
"description": "Array of URLs to scrape. Maximum 1000 per request.",
"items": {
"properties": {
"force": {
"description": "Bypass 90-day cache.",
"type": "boolean"
},
"metadata": {
"additionalProperties": {
"type": "string"
},
"description": "Custom key-value metadata stored with the result. Shared with every caller of this deployment and returned to anyone who retrieves the row; never put credentials or personal data here.",
"type": "object"
},
"url": {
"description": "The URL to scrape.",
"minLength": 1,
"type": "string"
}
},
"required": [
"url"
],
"type": "object"
},
"maxItems": 1000,
"minItems": 1,
"type": "array"
}
},
"required": [
"urls"
],
"type": "object"
}Generated from the live server (DataPulse MCP 1.0.0) on October 1, 2026.