Scrape SDK
Guides

Model Context Protocol

Full-page scrape when host WebFetch only summarizes. Also search, map, crawl, extract, and optional goal-based agents.

View Markdown

Cursor, Claude Code, and Codex already have WebSearch — keep using it. Their WebFetch often summarizes the page (Claude Code uses Haiku). This server returns the actual markdown body, plus site map, crawl, and JSON extract.

mcp.json
{
  "mcpServers": {
    "scrape-sdk": {
      "command": "npx",
      "args": ["-y", "scrape-sdk-mcp"],
      "env": {
        "FIRECRAWL_API_KEY": "fc-...",
        "FIRECRAWL_KEYLESS": "1",
        "TINYFISH_API_KEY": "sk-tinyfish-...",
        "TINYFISH_AGENT": "1",
        "TAVILY_API_KEY": "tvly-...",
        "JINA_API_KEY": "jina_..."
      }
    }
  }
}

A copy lives at examples/mcp.json. The skill at skills/scrape-sdk/ ships a sibling mcp.json.

Tools:

  • scrape_url — always. Full page markdown, default 20_000 chars, truncated + charCount + failedOverFrom. Named so it does not collide with host WebFetch. Use this when you need the real page.
  • search_web — when TinyFish, Tavily, Firecrawl, or another search provider is configured. Otherwise use the host WebSearch.
  • map_site / crawl_site / extract_json — when a configured provider supports them. Hosts do not ship these.
  • run_web_agent — only when TINYFISH_AGENT=1; this is a metered interactive operation.

fromEnv() reads FIRECRAWL_API_KEY (or FIRECRAWL_KEY), FIRECRAWL_KEYLESS=1, TINYFISH_API_KEY, TAVILY_API_KEY, JINA_API_KEY, SPIDER_API_KEY, and BROWSERBASE_API_KEY.