Scrape SDK

How it works

Routing, retries, timeouts, and capabilities.

View Markdown
  1. Validate the URL (http / https only).
  2. Prefer /llms.txt on a site root or /docs path when the requested options are compatible (skip article URLs). The probe is bounded by the client timeout; HTML scrape if it misses.
  3. Select providers that advertise the required capability (scrape, search, crawl, extract, js, or agent).
  4. Order them by strategy: priority (config order) or cost (cheaper first).
  5. Execute with an AbortSignal timeout. Retryable errors (429, 5xx, timeout) retry on the same provider, then fail over. Auth errors (401/403) do not retry.
  6. Return a unified result with provider, latencyMs, and failedOverFrom when a hop happened — or AllProvidersFailedError / CapabilityError. Goal-based browser work uses the separate agent() operation so paid interactive runs are never mistaken for ordinary page fetches.

Local Cheerio is the only adapter that sanitizes HTML itself. Cloud adapters return the vendor's markdown.

There is no fake Browserbase page content and no crawl that pretends POST /crawl already contains pages.