MCP documentation menu

datapulse_scrape_list

ScrapingScope: datapulse:scrapeOutput: JSONRead-only

List results with pagination (supports batch_id filter)

Description

List scrape results with pagination.

Returns a paginated list of scrape results. Without batch_id the listing has NO chronological ordering guarantee: it is a table scan, so page 1 is not “the latest scrapes” and today’s results may sit on any page. With batch_id the listing is that batch’s ordered, paginated result set, which is the only way to get a complete set. The listing is global, not per caller: it includes results and metadata submitted by others. Each item carries cached=true when it was served from storage rather than scraped for this call.

Pagination:

  • limit: Results per page (default 100, max 1000)
  • offset: Skip N results (for paging through large result sets)
  • Response includes: count (results in this page), limit, and offset for building paginators. total (matching results across all pages) is only present on batch listings (batch_id set). When count < limit, you have reached the last page.

Batch filtering:

  • batch_id: Filter to a specific bulk submission batch. Returns status summary counts (pending/processing/completed/failed) and total count. Get the batch_id from datapulse_scrape_bulk response.

Each result includes:

  • id, domain, domain_hash, url, url_hash
  • status: pending/processing/completed/failed
  • error: Error message (only present when status is “failed”)
  • created_at, completed_at, last_scraped_at timestamps
  • render_time_ms, llm_time_ms, total_time_ms performance metrics
  • extraction: Structured data (when completed)
  • link_domains: the internal hostnames and external registrable domains the page references (absent on results scraped before 2026-07-02). The full links array, and the per-connection fields final_url, redirected, tls, quic and server_addr, come only from a single-result read: datapulse_scrape_result, datapulse_scrape_submit with wait=true, or datapulse_domain_overview.

Use datapulse_scrape_result for individual result details including DOM.

For usage guidance, call datapulse_help(topic="scraping")

Parameters

ParameterTypeDescription
batch_idstring
Filter results by batch ID (UUID returned by datapulse_scrape_bulk). Must be a valid UUID. Returns 400 for malformed values, 404 if no results exist for the batch.
Pattern: ^[0-9a-fA-F]{8}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{12}$
limitinteger
Maximum results to return. Default 100, max 1000.
Default: 100Min: 1Max: 1000
offsetinteger
Number of results to skip for pagination. Default 0.
Default: 0Min: 0
Input schema (JSON)
{
  "additionalProperties": false,
  "properties": {
    "batch_id": {
      "description": "Filter results by batch ID (UUID returned by datapulse_scrape_bulk). Must be a valid UUID. Returns 400 for malformed values, 404 if no results exist for the batch.",
      "pattern": "^[0-9a-fA-F]{8}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{12}$",
      "type": "string"
    },
    "limit": {
      "default": 100,
      "description": "Maximum results to return. Default 100, max 1000.",
      "maximum": 1000,
      "minimum": 1,
      "type": "integer"
    },
    "offset": {
      "default": 0,
      "description": "Number of results to skip for pagination. Default 0.",
      "minimum": 0,
      "type": "integer"
    }
  },
  "type": "object"
}

Generated from the live server (DataPulse MCP 1.0.0) on October 1, 2026.