Crawly
Documentation menu
Core API

Scrape

Turn one public URL into clean Markdown, links, and fetch metadata.

Request

Crawly tries a low-cost HTTP fetch first and escalates to managed Playwright rendering when a blocked response or JavaScript shell needs a browser.

curl -X POST https://app.crawly.cc/api/v1/scrape \
  -H "Authorization: Bearer $CRAWLY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url":"https://example.com"}'

Response

The response includes content, links, metadata, the fetch method (http or playwright), and credit information. Cached results may include metadata.cached: true.

Asynchronous scrape

Add async: true to enqueue the request. Crawly returns 202 with a job id; poll the Jobs endpoint until the job succeeds or fails.

GET /api/v1/jobs/:id

Cache and media behavior

cacheMode defaults to default. Use fresh to bypass an existing cached page, or only to return a cached page without contacting the target. A cache-only miss returns 404 cache_miss.

Note: PDF input is rejected explicitly with HTTP 415 unsupported_media_type in this release; Crawly does not silently treat PDF bytes as HTML.