Crawly
Documentation menu
Core API

Batch scrape

Scrape several known URLs in one asynchronous job.

Request

curl -X POST https://app.crawly.cc/api/v1/batch/scrape \
  -H "Authorization: Bearer $CRAWLY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"urls":["https://example.com","https://example.org"],"async":true}'

Progress and complete results

Poll /api/v1/batch/scrape/:id. Progress reports total, queued, active, completed, successful, failed, remaining, and credits used. Follow the authenticated next URL to retrieve complete documents in ordered pages; content is never silently shortened.

DELETE the same batch URL to stop new work while preserving completed pages. Idempotency-Key makes submission retries return the same operation.

Note: One logical batch accepts up to 200 URLs on every plan. Plans control processing throughput and credits, not submission size. Result pages are capped near 8 MiB without splitting a document.

Credits and partial completion

Crawly checks the minimum one-credit-per-URL cost before creating a job. Browser-rendered pages may cost more. If the balance is exhausted while processing, completed results and their charges are preserved; remaining pages are recorded as insufficient_credits failures without starting more network work.

Webhooks

Batch jobs emit signed started, page, completed, failed, and cancelled events. Page events contain status, credits and an authenticated result URL—not scraped content.

  • HTTPS destinations only
  • X-Crawly-Signature uses HMAC-SHA256
  • Set includePageUrl:false to omit target URLs
  • Delivery IDs remain stable across delivery retries