# Webfetch > webfetch.io is a web scraping API. Send a URL, get back rendered HTML, plain > text, markdown, or structured JSON extracted with CSS or XPath selectors. Request > parameters, defaults, credit costs and status-code semantics match ScrapingBee, so > an existing integration works after changing the base URL and the key. Operated by Tuxxin LLC (https://tuxxin.com), an IT company in business for nearly 16 years. Webfetch itself is new. Screenshots are captured on our own servers, with sister service webshot.site as the fallback. ## Endpoint GET|POST https://api.webfetch.io/v1/ GET|POST https://webfetch.io/api/v1/ Both hosts serve the same API. Authenticate with `Authorization: Bearer wfk_...`. The `api_key` query parameter also works, for compatibility with ScrapingBee code. curl "https://api.webfetch.io/v1/?url=https%3A%2F%2Fexample.com&render_js=false" \ -H "Authorization: Bearer wfk_your_key" A key is REQUIRED. There is no keyless tier: rate limiting alone does not hold against a large address pool, and an anonymous endpoint leaves nobody to act against when it is abused. A request without a key returns 401 with code `key_required`. A free account starts with 1,000 credits and gets 250 more on the 1st of every month, banking up to 1,000 (unused free credits carry over, but the balance is never topped up past 1,000). It needs only a working email address: https://webfetch.io/signup. Signup sends a six-digit code to the address and that is the whole of it. No card. Upgrading while signed in to the key portal (or paying with the same email address) keeps the same key; the free balance does not carry over into a paid plan, one-time bonus credits do. A canceled paid plan drops back to the free tier at the end of the paid period, starting again at 1,000 credits. If a payment fails, the plan keeps working for 72 hours before it drops back the same way. Free capacity is best effort. When paid traffic is heavy and open capacity falls to the share held back for paying customers, a free request gets HTTP 503 with code `free_tier_busy` and a `Retry-After` header. It is not charged, and paid plans are never turned away to make room for free traffic. ## Credits, not requests A request costs credits, and only a successful one is charged. no JavaScript rendering 1 JavaScript rendering 5 (this is the DEFAULT, as at ScrapingBee) premium proxy 10 static ISP proxies (addresses registered to internet premium proxy + rendering 25 providers, trusted more than datacenter IPs) stealth proxy 75 always rendered, in a browser hardened against common headless-detection checks. Routed through the mobile pool first when it has capacity, then residential or ISP (billed 25, as premium), then a datacenter IP or our own address (base rate); fallbacks return Wf-Proxy-Downgraded: true and Wf-Proxy-Pool names the pool used any failed request 0 per-scan bandwidth ceiling, measured on the wire, not on the returned document: stealth pools (mobile, residential) 5 MB everything else 25 MB reaching it truncates rather than fails: content so far is returned and billed, with Wf-Truncated: true and Wf-Max-Bytes set. `render_js` defaults to **true**, so a call that does not set it costs 5 credits, not 1. A free account's starting 1,000 credits are therefore 200 rendered pages, or 1,000 raw fetches with `render_js=false`. Every response carries `Wf-Cost` with what was actually charged. Read live state at `GET /v1/usage` with a key. Promo codes grant one-time bonus credits at signup. They are spent only after the free balance or monthly allowance runs out, so they do not vanish at a month boundary. The remaining balance appears in `Wf-Bonus-Credits` and in `/v1/usage`. ## MCP server, for AI clients Webfetch speaks the Model Context Protocol, so an MCP-aware client can call it as a tool instead of scraping these pages. POST https://webfetch.io/mcp JSON-RPC 2.0, protocol 2025-06-18 { "mcpServers": { "webfetch": { "url": "https://webfetch.io/mcp", "headers": { "Authorization": "Bearer wfk_your_key" } } } } Two tools: fetch_page(url, render_js?, return_format?) return_format: markdown | text | html extract_data(url, extract_rules, render_js?) selectors to structured JSON Send `Authorization: Bearer `. A key is required, as it is on the HTTP API. Every result reports the credits it cost. The server card is at /.well-known/mcp/server-card.json. `/mcp` is POST only and answers 405 to GET. It is an endpoint, not a page, and is deliberately absent from the sitemap. ## Official SDKs npm install @tuxxin/webfetch Node 18+, no runtime dependencies pip install webfetch-io Python 3.9+, sync and async (imports as webfetch) composer require tuxxin/webfetch PHP 8.1+, cURL directly Source for all three: https://github.com/tuxxin/webfetch-sdk Same shape in all three: a client takes the key, `fetch()` returns the content plus `credits`, `status` and `resolved_url`, and errors raise a three-level hierarchy ending in a rate-limit type that carries `retry_after` (a 429, or a 503 when free capacity is busy). ## The parameters worth knowing url required, URL-encoded render_js default true; false for a plain HTTP fetch at 1 credit extract_rules JSON of selectors to structured JSON, the highest-value feature js_scenario click, fill, scroll and wait before capture return_page_text readable text, page chrome removed return_page_markdown the content as markdown, useful for a model json_response wrap everything in one JSON envelope wait, wait_for, wait_browser settle conditions premium_proxy, stealth_proxy, country_code, own_proxy, session_id block_resources (default true) also strips third-party telemetry: ad exchanges, analytics, tag managers, error reporting, consent banners and social widgets. About 37% fewer bytes with 100% of the text preserved. The main document is never blocked, and apex domains you might scrape are never on the list. mode=auto ScrapingBee's Auto-Mode: tries five configurations cheapest first and stops at the first that works: raw (1), render_js (5), premium_proxy (10), premium_proxy + render_js (25), stealth_proxy (75). Failed ones cost nothing, so you pay once, for the one that succeeded (Wf-Auto-Cost, 0 if none did). Do not send render_js, premium_proxy, stealth_proxy or transparent_status_code with it (400). A 404, 410 or 413 is a real answer and is returned as-is; every other non-2xx status, a timeout, a failed render, or a challenge/CAPTCHA page climbs. max_cost credit ceiling for one mode=auto call (integer, at least 1) block_ads, block_resources, device, cookies, forward_headers timeout, transparent_status_code, tag respect_robots webfetch only; see the acceptable use policy Full reference: https://webfetch.io/documentation ## extract_rules {"title": "h1", "link": "a@href"} {"articles": {"selector": ".card", "type": "list", "output": { "title": "h2", "url": {"selector": "a", "output": "@href"}, "tags": {"selector": ".tag", "type": "list"} }}} `output` takes `text`, `html`, `@attribute`, `table_json`, `table_array`, `text_relevant`, `markdown_relevant`, or a nested object. `selector_type` is `auto`, `css` or `xpath`; under `auto` a leading `/` means XPath. `type` is `item` or `list`. ## Status codes Our own status is reported, not the target's: a target 2xx becomes 200, 404, 410 and 413 pass through, everything else becomes 500. Set `transparent_status_code=true` to get the target's code verbatim. 200, 404, 410 and 413 are billed; 400, 401, 403, 429, 500 and 503 are not. An anti-bot challenge, CAPTCHA or block page is never billed, whatever status it came with (many arrive as 200): it is returned as an error with code `challenge_page` and the vendor in `Wf-Challenge`. Every error body is JSON with a stable `code` field. ## Limits and conduct robots.txt is honored by default. Opting out needs a key with the acceptable use policy accepted and is logged per request. Private, loopback, link-local and reserved addresses are refused, and the check is repeated on every redirect. Concurrent requests are capped per plan, and requests to a single target host are capped both per key and across all callers. Bandwidth is capped on the wire, not on the returned document: 5 MB through the stealth pools, 25 MB everywhere else. - Migrating from ScrapingBee: https://webfetch.io/scrapingbee-alternative - Acceptable use: https://webfetch.io/aup - Report abuse, or ask to have your domain refused for every customer: abuse@tuxxin.com ## Pricing Credit allowances match ScrapingBee's at 38 to 40% below their price. Free $0 1,000 to start, +250 / month 5 concurrent (banks up to 1,000) Starter $11.99 75,000 25 Pro $29.99 250,000 50 Scale $59.99 1,000,000 100 Business $149.99 3,000,000 200 Business+ $359.99 8,000,000 400 Custom from $499 no datacenter credit cap, dedicated hardware, concurrency bounded only by the machine. Stealth is metered bandwidth, so its allowance is sized and priced with the hardware. Quoted per requirement. Every plan is PRICE LOCKED FOR LIFE: the price you sign up at is the price you keep, for as long as the payment method stays active and the subscription does not lapse. https://webfetch.io/pricing ## Not implemented The AI extraction parameters (`ai_query`, `ai_extract_rules`, `ai_selector`), `scraping_config`, and the Google Search endpoint. Sending one is reported back in `Wf-Unknown-Parameters` rather than silently ignored. ## Pages - https://webfetch.io/documentation full API reference - https://webfetch.io/pricing plans and the credit table - https://webfetch.io/key passwordless key portal - https://webfetch.io/aup acceptable use policy - https://webfetch.io/privacy what is stored and for how long - https://webfetch.io/terms terms of service