Every plan is 38 to 40% below ScrapingBee for the same credits, and price locked for life. See the comparison

Documentation

One endpoint. Send a URL and get back rendered HTML, clean text, markdown, a screenshot, or JSON shaped by your own selectors. Headless Chrome and a rotating proxy pool sit behind it. Already integrated against ScrapingBee? Parameter names, defaults and credit costs match theirs, so it's a base URL and a key. See Migrating from ScrapingBee.

Quick start

Every request needs a key. A free account starts with 1,000 credits, gets 250 more every month (banking up to 1,000), and needs only an email address.

curl "https://api.webfetch.io/v1/?url=https%3A%2F%2Fexample.com&render_js=false" \
  -H "Authorization: Bearer wfk_your_key"

Asking for structured data instead of HTML:

curl "https://api.webfetch.io/v1/" \
  -H "Authorization: Bearer wfk_your_key" \
  --data-urlencode "url=https://quotes.toscrape.com/js/" \
  --data-urlencode 'extract_rules={"quotes":{"selector":".quote","type":"list","output":{"text":".text","author":".author"}}}' \
  -G

Both hosts serve the same API:

https://api.webfetch.io/v1/
https://webfetch.io/api/v1/

Authentication

Send your key as a bearer token:

Authorization: Bearer wfk_your_key

The api_key query parameter also works, because ScrapingBee accepts it and migrating code often uses it. It is the weaker option: query strings end up in logs and referrers, so prefer the header.

There is no keyless tier. Rate limiting alone does not hold against a large address pool, and an anonymous endpoint leaves no way to act against whoever abuses it. A request without a key returns 401 with code key_required.

Parameters

Sent as a query string on GET, or as a JSON or form body on POST. Query parameters still apply on a POST.

ParameterTypeDefaultDescription
api_key string required Your key. Prefer the Authorization: Bearer header; this query parameter is accepted for compatibility.
url string required The URL to fetch, URL-encoded.
render_js boolean true Render with headless Chrome. Set false for a plain HTTP fetch at 1 credit instead of 5.
wait integer ms 0 Extra settle time after load. Maximum 35000.
wait_for string "" CSS or XPath selector to wait for before capturing.
wait_browser string domcontentloaded One of domcontentloaded, load, networkidle0, networkidle2.
js_scenario stringified JSON {} Click, fill, scroll and wait before capture. See below.
window_width integer 1920 Viewport width, 200 to 3840.
window_height integer 1080 Viewport height, 200 to 2160.
premium_proxy boolean false Route through static ISP proxies: addresses registered to internet providers rather than hosting companies, which sites trust more than datacenter IPs. 10 credits, or 25 with rendering. With no ISP proxy available the call falls back to the residential pool (same price), then a datacenter IP at the base rate, and says so in Wf-Proxy-Downgraded and Wf-Proxy-Pool. Note for ScrapingBee users: their premium_proxy is residential.
stealth_proxy boolean false For targets that refuse ordinary traffic. Routes through the mobile pool first, when it has a usable proxy, and renders in a browser hardened against common headless-detection checks: a consistent user agent, client hints, platform and GPU, and, when the exit country is known (from country_code or a country-pinned proxy), that country's time zone and language. Implies render_js. Costs 75 credits on a mobile exit and is capped at 5 MB per scrape on the mobile and residential pools. With no mobile proxy available the call falls back to the residential pool, then the ISP pool, billed at the premium rate (25), then a datacenter IP or, as a last resort, our own address, at the base rate; any fallback returns Wf-Proxy-Downgraded: true and Wf-Proxy-Pool names the pool used, so you are never charged for a pool you did not get.
country_code string "" Two-letter ISO 3166-1 code for proxy geotargeting. Use it with premium_proxy or stealth_proxy. A country the pool cannot exit from is refused with 400 country_unavailable rather than served from another country.
own_proxy string "" Use your own proxy: <protocol><user>:<pass>@<host>:<port>. Port defaults to 1080.
session_id integer "" Route requests sharing this id through the same exit IP.
block_resources boolean true Skip images, media, fonts and stylesheets, and strip third-party telemetry: ad exchanges, analytics, tag managers, error reporting, consent banners and social embed widgets. Measured across major news and retail sites this removes about 37% of the bytes while preserving 100% of the text, headings and links, because none of it carries page content. The main document is never blocked whatever its host, and the apex domains customers actually scrape are never on the list. Set false to load the page exactly as a browser would.
block_ads boolean false Strip the same telemetry and ad hosts even when block_resources is off. Redundant at the default, kept because it is a ScrapingBee parameter.
extract_rules stringified JSON "" CSS or XPath selectors to structured JSON. See below.
return_page_source boolean false Return the HTML as delivered, before JavaScript ran.
return_page_text boolean false Return readable text with the page chrome removed.
return_page_markdown boolean false Return the content as markdown.
json_response boolean false Wrap everything in a JSON envelope instead of returning the body raw.
add_html boolean false With json_response=true and a transformation (extract_rules, return_page_text or return_page_markdown), also include the untransformed HTML in body. A request with no transformation already returns it either way.
screenshot boolean false Also capture a screenshot. Returned as a link, at no extra credit cost.
screenshot_full_page boolean false Capture the whole scrollable page.
screenshot_selector string "" Report the bounding box of this selector with the screenshot.
cookies string "" Cookies to send: name=value,domain=example.com;other=value2.
forward_headers boolean false Forward headers you prefix with Wf- (or Spb-).
forward_headers_pure boolean false Forward only your headers, adding none of ours.
device string desktop desktop or mobile.
timeout integer ms 140000 Budget for the fetch, as at ScrapingBee. Whatever it says, a call is answered within 90 seconds in total, mode=auto routes included, because the network edge in front of the API closes longer connections. A request that runs out of time returns a timeout error and is not billed.
mode string "" Set to auto to let webfetch pick the cheapest configuration that works, as ScrapingBee's Auto-Mode does. It tries five configurations in order and stops at the first that succeeds: no rendering (1 credit), render_js (5), premium_proxy without rendering (10), premium_proxy with rendering (25), then stealth_proxy (75). Do not send render_js, premium_proxy, stealth_proxy or transparent_status_code with it: auto mode chooses them, and sending one is a 400. A 404, 410, 413, a robots refusal or a blocked address is a real answer and is never retried. Every other non-2xx target status, a timeout, a failed render, or an anti-bot challenge or CAPTCHA page (whatever status it came with) moves on to the next configuration. Configurations that failed cost nothing, so you are billed once, for the one that worked, reported in Wf-Auto-Cost (0 when none did). The configurations tried come back in Wf-Auto-Attempts.
max_cost integer "" Ceiling in credits on a single mode=auto call: only configurations that cost at most this much are tried, so the charge never exceeds it. max_cost=5 tries a raw fetch and a rendered one but never a proxy; max_cost=25 climbs up to premium_proxy with rendering but never reaches stealth_proxy. A whole number of at least 1. Ignored without mode=auto.
transparent_status_code boolean false Return the target's status code verbatim instead of ours.
tag string "" Your own label, echoed back in Wf-Tag.
respect_robots boolean true Webfetch only. Setting false needs a key with the acceptable use policy accepted, and is logged.

An unrecognized parameter does not fail the request. It comes back listed in Wf-Unknown-Parameters so a typo is visible rather than silent.

Credits

ConfigurationCredits
No JavaScript rendering1
JavaScript rendering5
Premium proxy10
Premium proxy with rendering25
Stealth proxy75
Any failed request0

Identical to ScrapingBee's table. Every response carries Wf-Cost with what was actually charged. If you ask for a proxy pool this deployment cannot provide, the request is served over the best route available, billed for the route it actually got, and marked Wf-Proxy-Downgraded: true.

extract_rules

A JSON object mapping a name to a selector. The response is JSON in that shape instead of HTML.

{"title": "h1", "link": "a@href"}

Expands to the full form, which you can also write out:

{
  "title": {"selector": "h1", "output": "text", "type": "item"},
  "link":  {"selector": "a",  "output": "@href", "type": "item"}
}
FieldValuesDefault
selectorA CSS selector or an XPath expressionrequired
selector_typeauto, css, xpath. Under auto, a leading / means XPath.auto
typeitem for the first match, list for all of themitem
outputtext, html, @attribute, table_json, table_array, text_relevant, markdown_relevant, or a nested object of further rulestext
cleanCollapse runs of whitespace and trimtrue

Nested rules run inside each matched element, and can mix CSS and XPath:

{
  "articles": {
    "selector": ".card",
    "type": "list",
    "output": {
      "title": ".post-title",
      "link":  {"selector": ".post-title a", "output": "@href"},
      "tags":  {"selector": ".tag", "type": "list"}
    }
  }
}

A whole table, keyed by its header row:

{"pricing": {"selector": "#pricing", "output": "table_json"}}

[{"Plan": "Starter", "Price": "$7.99"}, {"Plan": "Pro", "Price": "$19.99"}]

A selector that matches nothing yields null, or an empty array for a list. One bad rule never fails the whole extraction.

js_scenario

Actions performed in the page before it is captured. Requires render_js=true and completes within 40 seconds.

{
  "strict": true,
  "instructions": [
    {"wait_for": ".results"},
    {"click": "li.next > a"},
    {"wait": 1500},
    {"evaluate": "document.querySelectorAll('.item').length"}
  ]
}
InstructionArgument
clickselector
waitmilliseconds
wait_forselector
wait_for_and_clickselector
scroll_x, scroll_ypixels
fill[selector, value]
evaluateJavaScript; results land in evaluate_results
infinite_scroll{max_count, delay, end_click: {selector}}

strict defaults to true and stops on the first failing step. Set it false to continue. With json_response=true the response carries a js_scenario_report with per-step timings and outcomes.

Response

By default the body is exactly what the target returned. Metadata travels in headers, and the target's own headers are mirrored with a Wf- prefix.

HeaderMeaning
Wf-CostCredits charged, 0 for anything not billed
Wf-ChallengeOn a challenge_page error: the anti-bot vendor recognized (for example cloudflare, datadome)
Wf-Initial-Status-CodeThe status the target actually returned
Wf-Resolved-UrlThe URL after redirects
Wf-Proxy-PoolRoute used: direct, datacenter, isp, residential, mobile or own
Wf-Proxy-DowngradedPresent when a requested pool was unavailable; the call was billed for the route it got (see Wf-Cost)
Wf-Screenshot-UrlLink to the capture, when one was requested
Wf-TagYour tag, echoed back
Wf-RateLimit-*Limit (the monthly allowance, or on the free tier the 1,000-credit balance cap), Remaining (credits left, bonus included), Resource, and Reset (Unix time of the next allowance reset or free top-up)

With json_response=true everything arrives in one object:

{
  "headers": { ... },
  "cost": 5,
  "initial-status-code": 200,
  "resolved-url": "https://example.com/",
  "type": "html",
  "body": "<html>...</html>",
  "cookies": [ ... ],
  "evaluate_results": [ ... ],
  "iframes": [ ... ],
  "xhr": [ ... ],
  "metadata": {"json-ld": [ ... ], "microdata": [ ... ]}
}

Status codes

By default we report our own status, not the target's: any target 2xx becomes 200, 404, 410 and 413 pass through, and everything else becomes 500. Redirects are followed first, so the status is the end of the chain. Set transparent_status_code=true to get the target's code verbatim.

An anti-bot challenge, CAPTCHA or block page is never billed, whatever status it arrived with, including the ones served as 200. It comes back as an error with code challenge_page, and mode=auto moves on to the next configuration.

CodeBilledMeaning
200yesSuccess
400noA parameter is missing or invalid
401noInvalid key, or out of credits
403noTarget blocked, redirected somewhere blocked, or refused by robots.txt
404 / 410yesThe target says the URL is gone
413yesThe target refused the request as too large
429noRate limited, or too many concurrent requests
500noWe could not fetch it, the target returned a status we do not pass through, or it answered with an anti-bot challenge, CAPTCHA or block page instead of the content (code challenge_page, vendor in Wf-Challenge)
503noFree tier only: paid traffic is heavy and has priority. Code free_tier_busy; retry after the Retry-After seconds

Every error body is JSON with error, code and docs. The code is stable and worth branching on.

Usage endpoint

curl "https://api.webfetch.io/v1/usage" -H "Authorization: Bearer wfk_your_key"
{
  "plan": "starter",
  "max_api_credit": 75000,
  "used_api_credit": 1284,
  "bonus_credits": 0,
  "max_concurrency": 25,
  "current_concurrency": 0,
  "renewal_subscription_date": "2026-10-01T00:00:00+00:00"
}

On a free account, max_api_credit is the balance plus what was used this month, so max_api_credit - used_api_credit is what is left, as a ScrapingBee client computes it. The response also carries free_credit_balance, free_credit_monthly_topup and free_credit_cap, and renewal_subscription_date is when the next top-up lands.

Official SDKs

MIT licensed, and thin: each wraps the same endpoint documented above. Source is on GitHub.

LanguageInstallRequires
JavaScriptnpm install @tuxxin/webfetchNode 18+, no runtime dependencies
Pythonpip install webfetch-ioPython 3.9+, sync and async, httpx; imports as webfetch
PHPcomposer require tuxxin/webfetchPHP 8.1+, cURL directly
import { WebfetchClient } from '@tuxxin/webfetch';
const client = new WebfetchClient(process.env.WEBFETCH_KEY);

const md = await client.markdown('https://example.com');

const data = await client.extract('https://quotes.toscrape.com/js/', {
  quotes: { selector: '.quote', type: 'list',
            output: { text: '.text', author: '.author' } },
});

Errors raise the same three levels everywhere: WebfetchError, then WebfetchApiError with the status and a stable code, then WebfetchRateLimitError carrying the retry delay (a 429, or a 503 when free capacity is busy). A failed call is never billed, so there is nothing to reconcile afterwards.

MCP server

Webfetch speaks the Model Context Protocol, so an MCP-aware client can call it as a tool rather than scraping this page.

{
  "mcpServers": {
    "webfetch": {
      "url": "https://webfetch.io/mcp",
      "headers": { "Authorization": "Bearer wfk_your_key" }
    }
  }
}
ToolArguments
fetch_pageurl, render_js, return_format (markdown, text or html)
extract_dataurl, extract_rules, render_js

JSON-RPC 2.0 over HTTP POST, protocol version 2025-06-18. A key is required, sent as Authorization: Bearer wfk_... on the HTTP request; calls draw on the key's credits (the free balance, or a paid plan's monthly allowance). Every result reports the credits it cost. The server card is at /.well-known/mcp/server-card.json. The endpoint is POST only and answers 405 to GET.

Migrating from ScrapingBee

  1. Change the base URL from https://app.scrapingbee.com/api/v1/ to https://api.webfetch.io/v1/.
  2. Swap the key.
  3. If you read Spb- response headers, read Wf- instead. Forwarded request headers still accept the Spb- prefix, so those need no change.

Parameter names, defaults, the credit table, the JSON envelope keys and the status-code rules are all the same. Two things are not, and both change what a request actually does rather than what it is called.

The proxy parameters map to different pools. ScrapingBee's premium_proxy is residential; ours routes static ISP proxies, addresses registered to internet providers that sites trust more than datacenter IPs. Their stealth_proxy is a separate pool with anti-bot hardening (they do not publish the proxy type); ours routes mobile exits first, then residential or ISP (billed as premium), in a hardened browser, and says which in Wf-Proxy-Pool and Wf-Proxy-Downgraded. See the comparison.

Not implemented: the AI extraction parameters (ai_query, ai_extract_rules, ai_selector), scraping_config, and the Google Search endpoint. Sending one is reported in Wf-Unknown-Parameters rather than silently ignored.

Limits and fair use

  • Concurrent requests are capped per plan: 5 on free, 25 on Starter, up to 400 on Business+. Custom is bounded only by the hardware it runs on.
  • The free tier is a balance: 1,000 credits at signup, then 250 more on the 1st of each month, never topped up past 1,000. Unused free credits carry over up to that cap; they do not carry over into a paid plan. Paid allowances reset on the 1st.
  • Bonus credits from a promo code are one-time. They are spent only after the free balance or monthly allowance is used up, so they survive a month rollover. The remaining balance is reported in Wf-Bonus-Credits and by /v1/usage.
  • Requests to a single target host are capped per key and across all callers, so nobody can point the service at one victim.
  • robots.txt is honored by default. See the acceptable use policy before opting out.
  • Private, loopback, link-local and reserved addresses are refused, and the check is repeated on every redirect.
  • Each scrape has a bandwidth ceiling, measured on the wire rather than on the returned document, so assets a page loads count toward it: 5 MB through the stealth pools (mobile and residential) and 25 MB everywhere else. A page that reaches the ceiling still returns the content captured up to that point, billed normally, marked Wf-Truncated: true with the ceiling in Wf-Max-Bytes. Ordinary pages are nowhere near it: across a sample of major news, reference and retail sites the median rendered scrape moved under 1 MB.