Core API
Scrape
Turn one public URL into clean Markdown, links, and fetch metadata.
Request
Crawly tries a low-cost HTTP fetch first and escalates to managed Playwright rendering when a blocked response or JavaScript shell needs a browser.
curl -X POST https://app.crawly.cc/api/v1/scrape \
-H "Authorization: Bearer $CRAWLY_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url":"https://example.com"}'Response
The response includes content, links, metadata, the fetch method (http or playwright), and credit information. Cached results may include metadata.cached: true.
Asynchronous scrape
Add async: true to enqueue the request. Crawly returns 202 with a job id; poll the Jobs endpoint until the job succeeds or fails.
GET /api/v1/jobs/:idCache and media behavior
cacheMode defaults to default. Use fresh to bypass an existing cached page, or only to return a cached page without contacting the target. A cache-only miss returns 404 cache_miss.
Note: PDF input is rejected explicitly with HTTP 415 unsupported_media_type in this release; Crawly does not silently treat PDF bytes as HTML.
