Subject Matter Expert - Web Scraping & Crawling (Web Data Platform)
Summary
Senior web-scraping SME at pharma analytics consultancy Chryselys, owning a production web-crawling platform (Python, Scrapy/Playwright/httpx) that feeds an AWS Bedrock LLM verification agent. Day to day: hardening legacy crawlers, reverse-engineering payer JSON APIs, handling anti-bot/egress, and ensuring legal/compliance-safe data acquisition.
About Chryselys
Chryselys is a Great Place to Work Certified Pharma Analytics & Business consulting company that delivers data-driven insights leveraging AI-powered, cloud-native platforms to achieve high-impact transformations. We specialize in digital technologies and advanced data science techniques that provide strategic and operational insights.
Role Summary
Own and extend a production-bound POC that combines resilient web-crawling engineering with an AWS Bedrock-based LLM verification agent, in a healthcare-data-compliance-sensitive context.
Responsibilities
- Replace scraped SERPs with a paid search API behind the existing
search_domain()interface. - Move
fetch.pyto async httpx with per-domain concurrency limits and politeness budgets. - Introduce proxy rotation and egress management; retire the single-IP failure mode.
- Make failure loud: typed outcomes, structured logs, metrics, drift and volume alerts.
- Reverse-engineer payer XHR/JSON endpoints to replace Playwright recipes wherever possible.
- Honor robots.txt Disallow and Crawl-delay; build a per-domain ToS and licensing register.
- Containerize and schedule the pipeline; add CI running the offline tests on every change.
- Add HAR replay and golden-file parser tests on real payer HTML, plus daily canary crawls.
Skills
Web & protocol fundamentals — HTTP/1.1, HTTP/2, HTTP/3, TLS, JA3/JA4 fingerprinting, cookies/sessions, OAuth2/JWT/SSO, DOM, encodings [Must]
Legacy stack (real mileage) — urllib, requests, mechanize, cURL/wget, BeautifulSoup, lxml, html5lib, XPath/XSLT, Scrapy (middlewares, pipelines, CrawlSpider), Selenium 3 patterns, Splash/PhantomJS, RSS/Atom/SOAP/XML feeds, ASP.NET __VIEWSTATE postbacks, frame & table-layout scraping, sitemap.xml, WARC/Common Crawl [Must]
Modern stack — Playwright (contexts, tracing, network interception), Puppeteer, Selenium 4/CDP, httpx/aiohttp + asyncio, curl_cffi TLS impersonation, selectolax, scrapy-playwright, Crawlee/Apify, JSON-LD/microdata harvesting [Must]
Reverse engineering — Private/undocumented JSON APIs, XHR & fetch interception, GraphQL, mobile API capture (mitmproxy/Charles), bundled-JS deobfuscation, WASM challenge analysis [Must]
Methodology breadth — API-first vs browser rendering, static vs JS-rendered SPA (__NEXT_DATA__, Nuxt payloads), BFS/DFS frontier management, URL canonicalization & dedup (bloom filters/hashing), full-refresh vs incremental/CDC crawling, distributed queue-based crawling (Redis/Kafka/SQS), authenticated sessions, pagination & infinite scroll, per-domain politeness [Must]
Non-HTML extraction — PDF (pdfplumber, PyMuPDF — in use today), OCR (Tesseract/cloud), Office formats, LLM/vision-assisted extraction and its cost & failure modes [Must]
Anti-bot & reliability — Cloudflare, Akamai, DataDome, PerimeterX, Imperva, Kasada; browser/canvas/WebGL fingerprinting; residential/datacenter/mobile proxy rotation; CAPTCHA landscape; honeypot detection; backoff with jitter, circuit breakers, idempotent retries [Must]
Data engineering — Python expert (async, typing, profiling), SQL, pydantic validation (in use); pandera, entity resolution, Postgres/Mongo/S3/Parquet, Airflow/Prefect/Dagster, Docker/K8s, CI/CD to introduce; Node/TypeScript useful [Must]
Testing & observability — vcrpy/responses, HAR replay, golden-file parser tests, canary crawls, selector-drift alerting, volume-anomaly detection, structured logging & metrics [Must]
Build vs. buy — Evaluating Zyte, Bright Data, Oxylabs, Apify, ScrapingBee, Firecrawl on cost, coverage, and risk [Preferred]
Legal & ethical — robots.txt and rate-limit adherence, ToS/CFAA awareness, GDPR & PII minimization, licensed-API-first sourcing; designs defensible, low-risk acquisition strategies in partnership with Legal [Must]
Experience — Required
- 8+ years in data acquisition; 5+ years owning a scraping platform end to end.
- Proven scale: 1,000+ distinct domains or 10M+ pages/month sustained in production.
- Rescued a brittle legacy scraper, with before/after reliability and cost numbers.
- Mentored engineers; set crawl standards, review practice, and on-call runbooks.
- Degree optional — equivalent practical experience is fully accepted.
Nice-to-Have
- US payer policy, formulary, or prior-authorization document domain knowledge.
- Open-source contributions to Scrapy, Playwright, Crawlee, or a parsing library.
- LLM-assisted extraction at controlled cost — we run AWS Bedrock in
verifier/. - Compliance or legal-review exposure on data acquisition programmes.
Equal Employment Opportunity
Chryselys is proud to be an Equal Employment Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.
Skills
- AI
- Airflow
- Analytics
- API
- AWS
- AWS Bedrock
- CI/CD
- Cloud
- Cloud Native
- Dagster
- Data Engineering
- Data Science
- Docker
- .NET
- Gdpr
- GraphQL
- HTML
- JSON
- JWT
- Kafka
- Kubernetes
- LLM
- MongoDB
- Nuxt
- OAuth
- Observability
- Parquet
- Playwright
- PostgreSQL
- PPC
- Prefect
- Puppeteer
- Python
- Redis
- Reverse Engineering
- Selenium
- SQL
- SQS
- SSO
- TLS
- TypeScript
- WebAssembly
- WebGL
- XML
- Xslt