Guides
The CrawlNode API, from first call to production
Every endpoint is a POST with a Token header. After start, every call also carries X-Session-Id. Bodies are JSON objects; send {} when there is nothing to say.
Quickstart
Three steps from nothing to a screenshot on your disk: get a key, make four calls, check what came back.
1. Get a key
- Create an account and confirm your email address.
- On the API keys page, choose how many physical nodes you want and send the request. We confirm the order with you, usually within one business day, and the key appears on that page.
- Reveal the key. It is shown once, with these same calls beside it and the key already filled in.
Then put it in your environment, which is where every sample here reads it from:
export CRAWLNODE_TOKEN='paste-your-key-here'
2. Make the four calls
The base URL is https://api1.crawlnode.com. Start a browser, navigate, take a screenshot, destroy.
# 1. start a browser (warm pool, managed proxy)
SESSION=$(curl -s -X POST https://api1.crawlnode.com/api/start \
-H "Token: $CRAWLNODE_TOKEN" -H "Content-Type: application/json" \
-d '{"proxy":"auto"}' | sed -n 's/.*"session_id":"\([^"]*\)".*/\1/p')
# 2. navigate
curl -s -X POST https://api1.crawlnode.com/api/go \
-H "Token: $CRAWLNODE_TOKEN" -H "X-Session-Id: $SESSION" \
-H "Content-Type: application/json" -d '{"url":"https://example.com"}'
# 3. screenshot of the page (WebP bytes); drop "area" for the whole window
curl -s -X POST https://api1.crawlnode.com/api/screenshot \
-H "Token: $CRAWLNODE_TOKEN" -H "X-Session-Id: $SESSION" \
-H "Content-Type: application/json" -d '{"area":"viewport"}' -o page.webp
# 4. always destroy
curl -s -X POST https://api1.crawlnode.com/api/destroy \
-H "Token: $CRAWLNODE_TOKEN" -H "X-Session-Id: $SESSION" \
-H "Content-Type: application/json" -d '{}'
import json, os, urllib.request
BASE = "https://api1.crawlnode.com"
TOKEN = os.environ["CRAWLNODE_TOKEN"]
def call(path, body, session=None, raw=False):
headers = {"Token": TOKEN, "Content-Type": "application/json"}
if session:
headers["X-Session-Id"] = session
req = urllib.request.Request(BASE + path, json.dumps(body).encode(), headers)
with urllib.request.urlopen(req, timeout=90) as res:
data = res.read()
return data if raw else json.loads(data)
session = call("/api/start", {"proxy": "auto"})["session_id"]
try:
call("/api/go", {"url": "https://example.com"}, session)
with open("page.webp", "wb") as f: # the page alone; {} captures the whole window
f.write(call("/api/screenshot", {"area": "viewport"}, session, raw=True))
finally:
call("/api/destroy", {}, session)
import { writeFile } from "node:fs/promises";
const BASE = "https://api1.crawlnode.com";
const TOKEN = process.env.CRAWLNODE_TOKEN;
async function call(path, body, session) {
const headers = { Token: TOKEN, "Content-Type": "application/json" };
if (session) headers["X-Session-Id"] = session;
const res = await fetch(BASE + path, { method: "POST", headers, body: JSON.stringify(body) });
if (!res.ok) throw new Error(`${path} -> HTTP ${res.status}`);
return res;
}
const { session_id } = await (await call("/api/start", { proxy: "auto" })).json();
try {
await call("/api/go", { url: "https://example.com" }, session_id);
const shot = await call("/api/screenshot", { area: "viewport" }, session_id); // {} = whole window
await writeFile("page.webp", Buffer.from(await shot.arrayBuffer()));
} finally {
await call("/api/destroy", {}, session_id);
}
The session_id comes back both in the body and in the X-Session-Id response header.
3. Check it worked
You now have page.webp in the directory you ran the calls from: the page as the browser painted it. If a call failed instead, its body says why, and Errors and retries says what to do about each status.
To check a key on its own, without starting a browser, ask what it has used. This is the one call that needs no session:
curl -s -X POST https://api1.crawlnode.com/api/usage \
-H "Token: $CRAWLNODE_TOKEN" -H "Content-Type: application/json" -d '{}'
It answers with the limits on the key and what has been spent against them this period. A 401 here means the key itself is wrong, before you have spent any time on a session.
Endpoints
All are POST /api/<name>. The ones marked session need X-Session-Id.
| Endpoint | What it does |
|---|---|
start | Create a session. Options: proxy (auto, a pool type, or your own), proxy_country, proxy_region, node, target, extension for network capture. |
go session | Navigate to a URL. Options for waiting out or solving challenges and for falling back to another exit. |
html session | The rendered DOM after JavaScript, with a quiet window or a selector to wait for. |
screenshot session | WebP by default, or PNG and JPEG. area: viewport captures the page alone; the default is the whole browser window. Re-grabs until the page has painted, up to settle_ms; X-Capture-Area says what you got. |
screenshots session | A low-fps MJPEG stream of the window until you disconnect. |
pdf session | Print the page to PDF. |
wait session | Block until the page settles or a selector appears, instead of polling. |
extract session | Pull named targets out of the page by selector in one call. |
script session | Run a list of steps (go, wait, extract, click, input, and more) as one round trip, buffered or streamed. |
view, click, input, drag session | Inspect the accessibility tree and interact with elements. |
eval, dom, console, state session | Run JavaScript, query the DOM, read console output, read page state. |
network, download session | List captured requests and download any request or response body. Needs extension: true on start. |
fetch, upload session | Fetch a URL with the session's cookies; upload a file into a form. |
solve session | Run the challenge solver ladder against what is on screen. |
proxy, egress session | Switch the upstream proxy without restarting Chrome; read the current exit IP. |
sessions | List your live sessions. |
destroy, clear session | Release the browser; wipe browsing data and keep the session. |
Proxies and targeting
Every session leaves through an exit you choose on start. Managed exits are part of every node and are in the United States today. A proxy of your own is used as given and never counted, and is the way to arrive from anywhere else.
Send on start | What you get |
|---|---|
"proxy": "auto" | An exit drawn from the whole managed pool. The right choice until a site gives you a reason to be specific. |
"proxy": "rotating_residential" | Only that pool type. The others are static_residential, isp and mobile. Use one when a site refuses datacenter traffic. |
"proxy_country": "US" | The country of the exit, as an ISO code, beside auto or a pool type. Every managed exit is in the United States today, so US is the only value that finds one. |
"proxy_region": "..." | A region inside the country, when the vendor exposes one. |
"proxy": "user:pass@host:port" | Your own proxy. host:port and the host:port:user:pass export format work too. SOCKS is not supported and is rejected on start rather than silently dialled. |
A residential exit, for a site that blocks hosting ranges:
SESSION=$(curl -s -X POST https://api1.crawlnode.com/api/start \
-H "Token: $CRAWLNODE_TOKEN" -H "Content-Type: application/json" \
-d '{"proxy":"rotating_residential"}' \
| sed -n 's/.*"session_id":"\([^"]*\)".*/\1/p')
Then ask the session where it is leaving from, before you trust what the site shows you. egress answers with the address the site sees:
curl -s -X POST https://api1.crawlnode.com/api/egress \
-H "Token: $CRAWLNODE_TOKEN" -H "X-Session-Id: $SESSION" \
-H "Content-Type: application/json" -d '{}'
proxy on a live session switches the exit without restarting Chrome, so one browser can look at the same page from two addresses.
Errors and retries
Success is 200. Errors carry a JSON body: {"error": {"code", "message"}} from the dispatcher, or {"detail": "..."} from the browser agent. Branch on the code, not the text.
- 401 unknown session or invalid token. Start a new session.
- 429 rate limit exceeded. Wait for
Retry-Afterseconds and retry. - 503 no free node, or no proxy passed the probe. Also carries
Retry-After; retry after it. - 504 the agent timed out. Retry once; the session is still valid.
- 4xx anything else is yours to fix and will fail identically on a retry.
Good practice
- Always destroy, including on failure. A leaked session holds a Chrome and a node slot until the idle timer reclaims it after 30 minutes.
- Send JSON objects, never arrays. An empty body is
{}. - Use
waitinstead of pollingview. A view walks the whole accessibility tree; wait answers the same question without shipping the page. - Read
X-Screenshot-Settled. It is present with value0only when the page never painted. Absent is the good case. - Prefer
scriptfor multi-step work. One round trip instead of one per step, and a failing step does not fail the run.