Documentation
One endpoint. Send a URL and get back rendered HTML, clean text, markdown, a screenshot, or JSON shaped by your own selectors. Headless Chrome and a rotating proxy pool sit behind it. Already integrated against ScrapingBee? Parameter names, defaults and credit costs match theirs, so it's a base URL and a key. See Migrating from ScrapingBee.
Quick start
Every request needs a key. A free account starts with 1,000 credits, gets 250 more every month (banking up to 1,000), and needs only an email address.
curl "https://api.webfetch.io/v1/?url=https%3A%2F%2Fexample.com&render_js=false" \
-H "Authorization: Bearer wfk_your_key"Asking for structured data instead of HTML:
curl "https://api.webfetch.io/v1/" \
-H "Authorization: Bearer wfk_your_key" \
--data-urlencode "url=https://quotes.toscrape.com/js/" \
--data-urlencode 'extract_rules={"quotes":{"selector":".quote","type":"list","output":{"text":".text","author":".author"}}}' \
-GBoth hosts serve the same API:
https://api.webfetch.io/v1/
https://webfetch.io/api/v1/Authentication
Send your key as a bearer token:
Authorization: Bearer wfk_your_key
The api_key query parameter also works, because ScrapingBee accepts
it and migrating code often uses it. It is the weaker option: query strings end up
in logs and referrers, so prefer the header.
There is no keyless tier. Rate limiting alone does not hold against a large
address pool, and an anonymous endpoint leaves no way to act against whoever
abuses it. A request without a key returns 401 with code
key_required.
Parameters
Sent as a query string on GET, or as a JSON or form body on POST. Query parameters still apply on a POST.
| Parameter | Type | Default | Description |
|---|---|---|---|
api_key |
string | required |
Your key. Prefer the Authorization: Bearer header; this query parameter is accepted for compatibility. |
url |
string | required |
The URL to fetch, URL-encoded. |
render_js |
boolean | true |
Render with headless Chrome. Set false for a plain HTTP fetch at 1 credit instead of 5. |
wait |
integer ms | 0 |
Extra settle time after load. Maximum 35000. |
wait_for |
string | "" |
CSS or XPath selector to wait for before capturing. |
wait_browser |
string | domcontentloaded |
One of domcontentloaded, load, networkidle0, networkidle2. |
js_scenario |
stringified JSON | {} |
Click, fill, scroll and wait before capture. See below. |
window_width |
integer | 1920 |
Viewport width, 200 to 3840. |
window_height |
integer | 1080 |
Viewport height, 200 to 2160. |
premium_proxy |
boolean | false |
Route through static ISP proxies: addresses registered to internet providers rather than hosting companies, which sites trust more than datacenter IPs. 10 credits, or 25 with rendering. With no ISP proxy available the call falls back to the residential pool (same price), then a datacenter IP at the base rate, and says so in Wf-Proxy-Downgraded and Wf-Proxy-Pool. Note for ScrapingBee users: their premium_proxy is residential. |
stealth_proxy |
boolean | false |
For targets that refuse ordinary traffic. Routes through the mobile pool first, when it has a usable proxy, and renders in a browser hardened against common headless-detection checks: a consistent user agent, client hints, platform and GPU, and, when the exit country is known (from country_code or a country-pinned proxy), that country's time zone and language. Implies render_js. Costs 75 credits on a mobile exit and is capped at 5 MB per scrape on the mobile and residential pools. With no mobile proxy available the call falls back to the residential pool, then the ISP pool, billed at the premium rate (25), then a datacenter IP or, as a last resort, our own address, at the base rate; any fallback returns Wf-Proxy-Downgraded: true and Wf-Proxy-Pool names the pool used, so you are never charged for a pool you did not get. |
country_code |
string | "" |
Two-letter ISO 3166-1 code for proxy geotargeting. Use it with premium_proxy or stealth_proxy. A country the pool cannot exit from is refused with 400 country_unavailable rather than served from another country. |
own_proxy |
string | "" |
Use your own proxy: <protocol><user>:<pass>@<host>:<port>. Port defaults to 1080. |
session_id |
integer | "" |
Route requests sharing this id through the same exit IP. |
block_resources |
boolean | true |
Skip images, media, fonts and stylesheets, and strip third-party telemetry: ad exchanges, analytics, tag managers, error reporting, consent banners and social embed widgets. Measured across major news and retail sites this removes about 37% of the bytes while preserving 100% of the text, headings and links, because none of it carries page content. The main document is never blocked whatever its host, and the apex domains customers actually scrape are never on the list. Set false to load the page exactly as a browser would. |
block_ads |
boolean | false |
Strip the same telemetry and ad hosts even when block_resources is off. Redundant at the default, kept because it is a ScrapingBee parameter. |
extract_rules |
stringified JSON | "" |
CSS or XPath selectors to structured JSON. See below. |
return_page_source |
boolean | false |
Return the HTML as delivered, before JavaScript ran. |
return_page_text |
boolean | false |
Return readable text with the page chrome removed. |
return_page_markdown |
boolean | false |
Return the content as markdown. |
json_response |
boolean | false |
Wrap everything in a JSON envelope instead of returning the body raw. |
add_html |
boolean | false |
With json_response=true and a transformation (extract_rules, return_page_text or return_page_markdown), also include the untransformed HTML in body. A request with no transformation already returns it either way. |
screenshot |
boolean | false |
Also capture a screenshot. Returned as a link, at no extra credit cost. |
screenshot_full_page |
boolean | false |
Capture the whole scrollable page. |
screenshot_selector |
string | "" |
Report the bounding box of this selector with the screenshot. |
cookies |
string | "" |
Cookies to send: name=value,domain=example.com;other=value2. |
forward_headers |
boolean | false |
Forward headers you prefix with Wf- (or Spb-). |
forward_headers_pure |
boolean | false |
Forward only your headers, adding none of ours. |
device |
string | desktop |
desktop or mobile. |
timeout |
integer ms | 140000 |
Budget for the fetch, as at ScrapingBee. Whatever it says, a call is answered within 90 seconds in total, mode=auto routes included, because the network edge in front of the API closes longer connections. A request that runs out of time returns a timeout error and is not billed. |
mode |
string | "" |
Set to auto to let webfetch pick the cheapest configuration that works, as ScrapingBee's Auto-Mode does. It tries five configurations in order and stops at the first that succeeds: no rendering (1 credit), render_js (5), premium_proxy without rendering (10), premium_proxy with rendering (25), then stealth_proxy (75). Do not send render_js, premium_proxy, stealth_proxy or transparent_status_code with it: auto mode chooses them, and sending one is a 400. A 404, 410, 413, a robots refusal or a blocked address is a real answer and is never retried. Every other non-2xx target status, a timeout, a failed render, or an anti-bot challenge or CAPTCHA page (whatever status it came with) moves on to the next configuration. Configurations that failed cost nothing, so you are billed once, for the one that worked, reported in Wf-Auto-Cost (0 when none did). The configurations tried come back in Wf-Auto-Attempts. |
max_cost |
integer | "" |
Ceiling in credits on a single mode=auto call: only configurations that cost at most this much are tried, so the charge never exceeds it. max_cost=5 tries a raw fetch and a rendered one but never a proxy; max_cost=25 climbs up to premium_proxy with rendering but never reaches stealth_proxy. A whole number of at least 1. Ignored without mode=auto. |
transparent_status_code |
boolean | false |
Return the target's status code verbatim instead of ours. |
tag |
string | "" |
Your own label, echoed back in Wf-Tag. |
respect_robots |
boolean | true |
Webfetch only. Setting false needs a key with the acceptable use policy accepted, and is logged. |
An unrecognized parameter does not fail the request. It comes back listed in
Wf-Unknown-Parameters so a typo is visible rather than silent.
Credits
| Configuration | Credits |
|---|---|
| No JavaScript rendering | 1 |
| JavaScript rendering | 5 |
| Premium proxy | 10 |
| Premium proxy with rendering | 25 |
| Stealth proxy | 75 |
| Any failed request | 0 |
Identical to ScrapingBee's table. Every response carries Wf-Cost with
what was actually charged. If you ask for a proxy pool this deployment cannot
provide, the request is served over the best route available, billed for the route
it actually got, and marked Wf-Proxy-Downgraded: true.
extract_rules
A JSON object mapping a name to a selector. The response is JSON in that shape instead of HTML.
{"title": "h1", "link": "a@href"}Expands to the full form, which you can also write out:
{
"title": {"selector": "h1", "output": "text", "type": "item"},
"link": {"selector": "a", "output": "@href", "type": "item"}
}| Field | Values | Default |
|---|---|---|
selector | A CSS selector or an XPath expression | required |
selector_type | auto, css, xpath. Under auto, a leading / means XPath. | auto |
type | item for the first match, list for all of them | item |
output | text, html, @attribute, table_json, table_array, text_relevant, markdown_relevant, or a nested object of further rules | text |
clean | Collapse runs of whitespace and trim | true |
Nested rules run inside each matched element, and can mix CSS and XPath:
{
"articles": {
"selector": ".card",
"type": "list",
"output": {
"title": ".post-title",
"link": {"selector": ".post-title a", "output": "@href"},
"tags": {"selector": ".tag", "type": "list"}
}
}
}A whole table, keyed by its header row:
{"pricing": {"selector": "#pricing", "output": "table_json"}}
[{"Plan": "Starter", "Price": "$7.99"}, {"Plan": "Pro", "Price": "$19.99"}]A selector that matches nothing yields null, or an empty
array for a list. One bad rule never fails the whole extraction.
js_scenario
Actions performed in the page before it is captured. Requires
render_js=true and completes within 40 seconds.
{
"strict": true,
"instructions": [
{"wait_for": ".results"},
{"click": "li.next > a"},
{"wait": 1500},
{"evaluate": "document.querySelectorAll('.item').length"}
]
}| Instruction | Argument |
|---|---|
click | selector |
wait | milliseconds |
wait_for | selector |
wait_for_and_click | selector |
scroll_x, scroll_y | pixels |
fill | [selector, value] |
evaluate | JavaScript; results land in evaluate_results |
infinite_scroll | {max_count, delay, end_click: {selector}} |
strict defaults to true and stops on the first failing
step. Set it false to continue. With json_response=true the response
carries a js_scenario_report with per-step timings and outcomes.
Response
By default the body is exactly what the target returned. Metadata travels in
headers, and the target's own headers are mirrored with a Wf-
prefix.
| Header | Meaning |
|---|---|
Wf-Cost | Credits charged, 0 for anything not billed |
Wf-Challenge | On a challenge_page error: the anti-bot vendor recognized (for example cloudflare, datadome) |
Wf-Initial-Status-Code | The status the target actually returned |
Wf-Resolved-Url | The URL after redirects |
Wf-Proxy-Pool | Route used: direct, datacenter, isp, residential, mobile or own |
Wf-Proxy-Downgraded | Present when a requested pool was unavailable; the call was billed for the route it got (see Wf-Cost) |
Wf-Screenshot-Url | Link to the capture, when one was requested |
Wf-Tag | Your tag, echoed back |
Wf-RateLimit-* | Limit (the monthly allowance, or on the free tier the
1,000-credit balance cap), Remaining (credits left, bonus included),
Resource, and Reset (Unix time of the next allowance reset or free top-up) |
With json_response=true everything arrives in one object:
{
"headers": { ... },
"cost": 5,
"initial-status-code": 200,
"resolved-url": "https://example.com/",
"type": "html",
"body": "<html>...</html>",
"cookies": [ ... ],
"evaluate_results": [ ... ],
"iframes": [ ... ],
"xhr": [ ... ],
"metadata": {"json-ld": [ ... ], "microdata": [ ... ]}
}Status codes
By default we report our own status, not the target's: any target 2xx becomes
200, 404, 410 and 413 pass through, and everything else becomes 500. Redirects
are followed first, so the status is the end of the chain. Set
transparent_status_code=true to get the target's code verbatim.
An anti-bot challenge, CAPTCHA or block page is never billed, whatever status it
arrived with, including the ones served as 200. It comes back as an error with code
challenge_page, and mode=auto moves on to the next
configuration.
| Code | Billed | Meaning |
|---|---|---|
| 200 | yes | Success |
| 400 | no | A parameter is missing or invalid |
| 401 | no | Invalid key, or out of credits |
| 403 | no | Target blocked, redirected somewhere blocked, or refused by robots.txt |
| 404 / 410 | yes | The target says the URL is gone |
| 413 | yes | The target refused the request as too large |
| 429 | no | Rate limited, or too many concurrent requests |
| 500 | no | We could not fetch it, the target returned a status we do not pass through, or it answered with an anti-bot challenge, CAPTCHA or block page instead of the content (code challenge_page, vendor in Wf-Challenge) |
| 503 | no | Free tier only: paid traffic is heavy and has priority. Code free_tier_busy; retry after the Retry-After seconds |
Every error body is JSON with error, code
and docs. The code is stable and worth branching on.
Usage endpoint
curl "https://api.webfetch.io/v1/usage" -H "Authorization: Bearer wfk_your_key"{
"plan": "starter",
"max_api_credit": 75000,
"used_api_credit": 1284,
"bonus_credits": 0,
"max_concurrency": 25,
"current_concurrency": 0,
"renewal_subscription_date": "2026-10-01T00:00:00+00:00"
}On a free account, max_api_credit is the balance plus what was
used this month, so max_api_credit - used_api_credit is what is left, as a
ScrapingBee client computes it. The response also carries free_credit_balance,
free_credit_monthly_topup and free_credit_cap, and
renewal_subscription_date is when the next top-up lands.
Official SDKs
MIT licensed, and thin: each wraps the same endpoint documented above. Source is on GitHub.
| Language | Install | Requires |
|---|---|---|
| JavaScript | npm install @tuxxin/webfetch | Node 18+, no runtime dependencies |
| Python | pip install webfetch-io | Python 3.9+, sync and async, httpx; imports as webfetch |
| PHP | composer require tuxxin/webfetch | PHP 8.1+, cURL directly |
import { WebfetchClient } from '@tuxxin/webfetch';
const client = new WebfetchClient(process.env.WEBFETCH_KEY);
const md = await client.markdown('https://example.com');
const data = await client.extract('https://quotes.toscrape.com/js/', {
quotes: { selector: '.quote', type: 'list',
output: { text: '.text', author: '.author' } },
});
Errors raise the same three levels everywhere: WebfetchError,
then WebfetchApiError with the status and a stable code, then
WebfetchRateLimitError carrying the retry delay (a 429, or a 503 when
free capacity is busy). A failed call is never billed, so there is nothing to
reconcile afterwards.
MCP server
Webfetch speaks the Model Context Protocol, so an MCP-aware client can call it as a tool rather than scraping this page.
{
"mcpServers": {
"webfetch": {
"url": "https://webfetch.io/mcp",
"headers": { "Authorization": "Bearer wfk_your_key" }
}
}
}| Tool | Arguments |
|---|---|
fetch_page | url, render_js, return_format (markdown, text or html) |
extract_data | url, extract_rules, render_js |
JSON-RPC 2.0 over HTTP POST, protocol version 2025-06-18. A key is
required, sent as Authorization: Bearer wfk_... on the HTTP request; calls
draw on the key's credits (the free balance, or a paid plan's monthly allowance).
Every result reports the credits it cost. The
server card is at /.well-known/mcp/server-card.json. The endpoint
is POST only and answers 405 to GET.
Migrating from ScrapingBee
- Change the base URL from
https://app.scrapingbee.com/api/v1/tohttps://api.webfetch.io/v1/. - Swap the key.
- If you read
Spb-response headers, readWf-instead. Forwarded request headers still accept theSpb-prefix, so those need no change.
Parameter names, defaults, the credit table, the JSON envelope keys and the status-code rules are all the same. Two things are not, and both change what a request actually does rather than what it is called.
The proxy parameters map to different pools.
ScrapingBee's premium_proxy is residential; ours routes static ISP
proxies, addresses registered to internet providers that sites trust more than
datacenter IPs. Their stealth_proxy is a separate pool with anti-bot
hardening (they do not publish the proxy type); ours routes mobile exits first, then
residential or ISP (billed as premium), in a hardened browser, and says which in
Wf-Proxy-Pool and Wf-Proxy-Downgraded.
See the comparison.
Not implemented: the AI extraction parameters (ai_query,
ai_extract_rules, ai_selector),
scraping_config, and the Google Search endpoint. Sending one is
reported in Wf-Unknown-Parameters rather than silently ignored.
Limits and fair use
- Concurrent requests are capped per plan: 5 on free, 25 on Starter, up to 400 on Business+. Custom is bounded only by the hardware it runs on.
- The free tier is a balance: 1,000 credits at signup, then 250 more on the 1st of each month, never topped up past 1,000. Unused free credits carry over up to that cap; they do not carry over into a paid plan. Paid allowances reset on the 1st.
- Bonus credits from a promo code are one-time. They are spent only after the
free balance or monthly allowance is used up, so they survive a month rollover. The
remaining balance is reported in
Wf-Bonus-Creditsand by/v1/usage. - Requests to a single target host are capped per key and across all callers, so nobody can point the service at one victim.
robots.txtis honored by default. See the acceptable use policy before opting out.- Private, loopback, link-local and reserved addresses are refused, and the check is repeated on every redirect.
- Each scrape has a bandwidth ceiling, measured on the wire rather than on
the returned document, so assets a page loads count toward it:
5 MB through the stealth pools (mobile and residential) and
25 MB everywhere else. A page that reaches the ceiling
still returns the content captured up to that point, billed normally,
marked
Wf-Truncated: truewith the ceiling inWf-Max-Bytes. Ordinary pages are nowhere near it: across a sample of major news, reference and retail sites the median rendered scrape moved under 1 MB.