Fetch any page the way a browser sees it.
One POST renders a URL in headless Chromium and gives you back clean
Markdown, boilerplate-free text, structured
metadata, every link, the rendered HTML,
or a full-page screenshot — with JavaScript executed, trackers blocked,
robots.txt respected and anti-bot blocks retried automatically.
What people use it for
Anything that needs the rendered page rather than the raw HTML the server shipped. These are the patterns that come up most often.
Audit what crawlers actually see
- Rendered vs. raw diff — run the same URL with
renderJstrue and false to find content that only exists after JavaScript. - Client-side links — compare the
linksarray across both modes to spot navigation search engines may never follow. - Canonicals and meta robots — catch tags that JavaScript injects, rewrites or strips.
- JSON-LD validation — read structured data as the browser assembled it, not as the server sent it.
- Soft 404 detection — a 200 status with empty rendered content is the classic silent failure.
- Bulk title and description audits across thousands of URLs.
Feed clean content to models
- RAG ingestion —
formats:["markdown"]turns any page into chunk-ready Markdown. - Context without the noise —
formats:["text"]returns main content only, no nav, no footer, no cookie banner. - Agent tooling — give an assistant or MCP server the ability to read live pages.
- Grounded answers — replace stale training data with the page as it exists right now.
- Richer embeddings — attach title, author, site name and published metadata to every chunk.
- Crawl frontiers — use the returned
linksto drive recursive ingestion.
Debug JavaScript-heavy pages
- Render cost —
timings.gotoMsagainsttotalMsshows where the time goes. - Silent breakage —
debug.consoleErrorsandrequestFailedsurface scripts failing in production. - Third-party impact — block analytics, tag managers or pixels with
blockUrlPatternsand see what changes. - Hydration timing —
waitForSelectorplusextraWaitMspins down how long content really takes. - Lazy-loading checks — confirm deferred content appears for a crawler at all.
Track what competitors ship
- Pricing pages that render entirely client-side.
- Content drift — store the Markdown and diff it on a schedule.
- Positioning signals — watch titles, descriptions and OpenGraph copy change.
- Hiring direction — careers pages are usually JavaScript apps.
- Stock and availability on ecommerce listings.
- Schema adoption — see who is adding structured data and where.
Screenshots on demand
- Full-page captures as PNG or JPEG for visual regression baselines.
- Dark mode — set
colorScheme:"dark"and capture the other half of your design. - Responsive checks — render at any
viewport, withisMobileand touch emulation. - Client reporting — pull fresh screenshots straight into audit decks.
- Link previews and dashboard thumbnails.
- Evidence capture when a page breaks or changes unexpectedly.
Launch without losing anything
- Before and after diffs — Markdown in, Markdown out, compare the two.
- Redirect verification —
finalUrlandstatusconfirm each hop lands where intended. - Metadata carryover — prove titles, canonicals and OpenGraph survived the replatform.
- Structured data survival — the first casualty of most CMS moves.
- Staging smoke tests behind basic auth or a preview header via custom
headers.
Watch content, not just uptime
- Content-aware health checks — a 200 that renders nothing is still an outage.
- Selector alarms — alert when a critical element stops appearing.
- Accidental noindex reaching production.
- Block detection —
debug.blockedandblockReasontell you when a target starts refusing you. - Injection and defacement — catch unexpected scripts or copy changes.
Localisation and structured pulls
- Geo-personalised content — set
geolocation,localeandtimezoneIdinstead of using a VPN. - Redirect behaviour by region and hreflang verification.
- Mobile parity — confirm mobile and desktop serve the same content.
- Field-level scraping —
extract.selectorswithtext,htmlorattr. - Stubborn targets — fingerprint rotation and a sticky per-host proxy pool.
- Glue it in — an HTTP node in n8n, Make or Zapier; a formula in Sheets.
One request
Ask for the formats you want. Anything you do not list is not computed and not returned.
curl -sS -X POST 'https://jscrawler.seoworkflow.online/v1/crawl' \
-H 'Content-Type: application/json' \
-H 'X-Api-Key: YOUR_KEY' \
-d '{
"url": "https://example.com",
"waitUntil": ["domcontentloaded", "load"],
"extraWaitMs": 0,
"maxHtmlBytes": 5,
"screenshot": { "enabled": false },
"formats": ["markdown", "metadata", "links"],
"respectRobots": true,
"maxRetries": 2
}'
Capabilities
Every option is per request, so one deployment serves very different jobs.
Output formats
markdown— main content as clean Markdown (Readability + Turndown).text— the same content as boilerplate-free plain text.metadata— title, description, lang, canonical, favicon, author, site name, OpenGraph, Twitter, JSON-LD and raw meta tags.links— absolute, de-duplicated http(s) links with anchor text and rel.html— the fully rendered DOM.extract.selectors— named CSS selectors returning text, HTML or an attribute.
Rendering control
renderJs— execute JavaScript, or skip it for a fast direct-HTTP fetch.waitUntil— any of domcontentloaded, load, networkidle.waitForSelectorandextraWaitMsfor slow hydration.blockUrlPatterns— drop trackers, ads and pixels by glob pattern.timeoutMs— a single wall-clock budget for the whole request.maxHtmlBytes— hard cap on response size.
Browser identity
- Fingerprint profiles for Windows Chrome, macOS Chrome and Android Chrome.
- Stealth mode hides common headless automation tells by default.
viewport,deviceScaleFactor,isMobile,hasTouch,colorScheme.locale,timezoneId,geolocation.- Custom
userAgentand requestheaders.
Getting through
- Automatic block detection for 403, 429, CAPTCHA and challenge pages.
maxRetriesrotates to a different fingerprint and a different proxy on each attempt.- Proxy pool selected sticky per host, so a target sees one consistent IP.
- Per-host rate spacing, jitter and concurrency caps keep crawling polite.
Safety
respectRobotsrefuses URLs the host disallows, with robots.txt cached per host.- SSRF protection on the initial URL and every redirect hop.
- DNS results validated at connect time, blocking DNS-rebinding.
- API-key auth on every endpoint, plus per-key rate limiting.
- Queue shedding returns 503 rather than growing without bound.
Observability
timings— total and navigation duration per crawl.debug— blocked requests, failed requests, console errors, main document status.debug.attempts,blockedandblockReasonwhen retries run./healthzfor liveness and/readyzfor a real Chromium launch check.
Endpoints
Full reference and schemas live in the docs.
| Endpoint | Purpose |
|---|---|
| POST /v1/crawl | Render a URL and return the formats you asked for. |
| GET /assets/:id | Download a screenshot from a crawl. Expires about 10 minutes after creation. |
| GET /playground | Build and run requests in the browser. |
| GET /docs | Interactive API reference. |
| GET /openapi.json | OpenAPI 3.0 specification. |
| GET /healthz | Liveness check. |
| GET /readyz | Readiness check, verifies Chromium can launch. |
Getting access
This deployment requires an API key. Every endpoint expects an
X-Api-Key header, and there is no self-service sign-up — keys are issued manually.
Want one? Get in touch via mihirnaik.com and say what you are building.
Prefer to run your own? The service is a single container — bring your own
API_KEYS, proxies and concurrency settings.