Guides

The CrawlNode API, from first call to production

Every endpoint is a POST with a Token header. After start, every call also carries X-Session-Id. Bodies are JSON objects; send {} when there is nothing to say.

Quickstart

Three steps from nothing to a screenshot on your disk: get a key, make four calls, check what came back.

1. Get a key

  1. Create an account and confirm your email address.
  2. On the API keys page, choose how many physical nodes you want and send the request. We confirm the order with you, usually within one business day, and the key appears on that page.
  3. Reveal the key. It is shown once, with these same calls beside it and the key already filled in.

Then put it in your environment, which is where every sample here reads it from:

export CRAWLNODE_TOKEN='paste-your-key-here'

2. Make the four calls

The base URL is https://api1.crawlnode.com. Start a browser, navigate, take a screenshot, destroy.

# 1. start a browser (warm pool, managed proxy)
SESSION=$(curl -s -X POST https://api1.crawlnode.com/api/start \
  -H "Token: $CRAWLNODE_TOKEN" -H "Content-Type: application/json" \
  -d '{"proxy":"auto"}' | sed -n 's/.*"session_id":"\([^"]*\)".*/\1/p')

# 2. navigate
curl -s -X POST https://api1.crawlnode.com/api/go \
  -H "Token: $CRAWLNODE_TOKEN" -H "X-Session-Id: $SESSION" \
  -H "Content-Type: application/json" -d '{"url":"https://example.com"}'

# 3. screenshot of the page (WebP bytes); drop "area" for the whole window
curl -s -X POST https://api1.crawlnode.com/api/screenshot \
  -H "Token: $CRAWLNODE_TOKEN" -H "X-Session-Id: $SESSION" \
  -H "Content-Type: application/json" -d '{"area":"viewport"}' -o page.webp

# 4. always destroy
curl -s -X POST https://api1.crawlnode.com/api/destroy \
  -H "Token: $CRAWLNODE_TOKEN" -H "X-Session-Id: $SESSION" \
  -H "Content-Type: application/json" -d '{}'

The session_id comes back both in the body and in the X-Session-Id response header.

3. Check it worked

You now have page.webp in the directory you ran the calls from: the page as the browser painted it. If a call failed instead, its body says why, and Errors and retries says what to do about each status.

To check a key on its own, without starting a browser, ask what it has used. This is the one call that needs no session:

curl -s -X POST https://api1.crawlnode.com/api/usage \
  -H "Token: $CRAWLNODE_TOKEN" -H "Content-Type: application/json" -d '{}'

It answers with the limits on the key and what has been spent against them this period. A 401 here means the key itself is wrong, before you have spent any time on a session.

Endpoints

All are POST /api/<name>. The ones marked session need X-Session-Id.

EndpointWhat it does
startCreate a session. Options: proxy (auto, a pool type, or your own), proxy_country, proxy_region, node, target, extension for network capture.
go sessionNavigate to a URL. Options for waiting out or solving challenges and for falling back to another exit.
html sessionThe rendered DOM after JavaScript, with a quiet window or a selector to wait for.
screenshot sessionWebP by default, or PNG and JPEG. area: viewport captures the page alone; the default is the whole browser window. Re-grabs until the page has painted, up to settle_ms; X-Capture-Area says what you got.
screenshots sessionA low-fps MJPEG stream of the window until you disconnect.
pdf sessionPrint the page to PDF.
wait sessionBlock until the page settles or a selector appears, instead of polling.
extract sessionPull named targets out of the page by selector in one call.
script sessionRun a list of steps (go, wait, extract, click, input, and more) as one round trip, buffered or streamed.
view, click, input, drag sessionInspect the accessibility tree and interact with elements.
eval, dom, console, state sessionRun JavaScript, query the DOM, read console output, read page state.
network, download sessionList captured requests and download any request or response body. Needs extension: true on start.
fetch, upload sessionFetch a URL with the session's cookies; upload a file into a form.
solve sessionRun the challenge solver ladder against what is on screen.
proxy, egress sessionSwitch the upstream proxy without restarting Chrome; read the current exit IP.
sessionsList your live sessions.
destroy, clear sessionRelease the browser; wipe browsing data and keep the session.

Proxies and targeting

Every session leaves through an exit you choose on start. Managed exits are part of every node and are in the United States today. A proxy of your own is used as given and never counted, and is the way to arrive from anywhere else.

Send on startWhat you get
"proxy": "auto"An exit drawn from the whole managed pool. The right choice until a site gives you a reason to be specific.
"proxy": "rotating_residential"Only that pool type. The others are static_residential, isp and mobile. Use one when a site refuses datacenter traffic.
"proxy_country": "US"The country of the exit, as an ISO code, beside auto or a pool type. Every managed exit is in the United States today, so US is the only value that finds one.
"proxy_region": "..."A region inside the country, when the vendor exposes one.
"proxy": "user:pass@host:port"Your own proxy. host:port and the host:port:user:pass export format work too. SOCKS is not supported and is rejected on start rather than silently dialled.

A residential exit, for a site that blocks hosting ranges:

SESSION=$(curl -s -X POST https://api1.crawlnode.com/api/start \
  -H "Token: $CRAWLNODE_TOKEN" -H "Content-Type: application/json" \
  -d '{"proxy":"rotating_residential"}' \
  | sed -n 's/.*"session_id":"\([^"]*\)".*/\1/p')

Then ask the session where it is leaving from, before you trust what the site shows you. egress answers with the address the site sees:

curl -s -X POST https://api1.crawlnode.com/api/egress \
  -H "Token: $CRAWLNODE_TOKEN" -H "X-Session-Id: $SESSION" \
  -H "Content-Type: application/json" -d '{}'

proxy on a live session switches the exit without restarting Chrome, so one browser can look at the same page from two addresses.

Errors and retries

Success is 200. Errors carry a JSON body: {"error": {"code", "message"}} from the dispatcher, or {"detail": "..."} from the browser agent. Branch on the code, not the text.

  • 401 unknown session or invalid token. Start a new session.
  • 429 rate limit exceeded. Wait for Retry-After seconds and retry.
  • 503 no free node, or no proxy passed the probe. Also carries Retry-After; retry after it.
  • 504 the agent timed out. Retry once; the session is still valid.
  • 4xx anything else is yours to fix and will fail identically on a retry.

Good practice

  • Always destroy, including on failure. A leaked session holds a Chrome and a node slot until the idle timer reclaims it after 30 minutes.
  • Send JSON objects, never arrays. An empty body is {}.
  • Use wait instead of polling view. A view walks the whole accessibility tree; wait answers the same question without shipping the page.
  • Read X-Screenshot-Settled. It is present with value 0 only when the page never painted. Absent is the good case.
  • Prefer script for multi-step work. One round trip instead of one per step, and a failing step does not fail the run.