How it works
Routing, retries, timeouts, and capabilities.
- Validate the URL (
http/httpsonly). - Prefer
/llms.txton a site root or/docspath when the requested options are compatible (skip article URLs). The probe is bounded by the client timeout; HTML scrape if it misses. - Select providers that advertise the required capability (
scrape,search,crawl,extract,js, oragent). - Order them by
strategy:priority(config order) orcost(cheaper first). - Execute with an
AbortSignaltimeout. Retryable errors (429, 5xx, timeout) retry on the same provider, then fail over. Auth errors (401/403) do not retry. - Return a unified result with
provider,latencyMs, andfailedOverFromwhen a hop happened — orAllProvidersFailedError/CapabilityError. Goal-based browser work uses the separateagent()operation so paid interactive runs are never mistaken for ordinary page fetches.
Local Cheerio is the only adapter that sanitizes HTML itself. Cloud adapters return the vendor's markdown.
There is no fake Browserbase page content and no crawl that pretends POST /crawl already contains pages.