freehire launches on Product Hunt on 26 August.

Follow →

SENIOR DATA ENGINEER

Summary

Senior Data Engineer builds and maintains a Python/Scrapy pipeline that scrapes legal/regulatory sources, normalizes data, and powers a React frontend for compliance teams.

Svitla Systems Inc. is looking for a Senior Data Engineer for a full-time position (40 hours per week) in Colombia, Mexico.

You'll take full ownership of a Python/Scrapy-based legal/regulatory data ingestion pipeline. This role combines large-scale data engineering, advanced web scraping against hostile sources, complex legal document processing, and cost-efficient cloud infrastructure design. It also includes maintenance of a React frontend application.

Requirements

  • 5+ years of experience in data engineering or backend development, with a focus on production pipelines.
  • Strong experience in Python, including production web scraping (Scrapy or equivalent) and structured-data processing.
  • Practical understanding of defeating or working around anti-bot measures (rate limiting, Cloudflare, paywalled/gated legal databases) within legal and ethical bounds.
  • Understanding of relational and non-relational databases; schema design for hierarchical data.
  • Solid understanding of SQL and data modeling, with the judgment to design schemas that serve both change detection and human-readable display.
  • Experience with cloud infrastructure (AWS, GCP, or Azure), IaC (Terraform/CloudFormation), and containers (Docker).
  • Knowledge of Microsoft Azure, including VMs, Blob storage, App Service plans, and container services, with the ability to provision and tear down resources programmatically.
  • Knowledge of search architectures (full-text, semantic search — Elasticsearch, OpenSearch, or vector DBs).
  • Familiarity with messaging/notification systems (email APIs, Slack/Teams webhooks).
  • Understanding of Infrastructure-as-code for repeatable, automated deployments.
  • Be comfortable owning a system end-to-end with minimal hand-holding, including reading and improving inherited code and documentation.
  • The ability to take ownership of complex legacy systems without extensive documentation.
  • Meticulous attention to detail: this work handles legal text where structural precision is critical.
  • Cost- and efficiency-oriented mindset.
  • Be comfortable working with ambiguity across heterogeneous, changing data sources.

Nice to have

  • Experience with React and Vite for maintaining and extending the display layer.
  • Prior experience in legal tech, regtech, or compliance software.
  • Familiarity with anti-detection scraping techniques (proxy rotation, fingerprinting, simulated human behavior).
  • Familiarity with deterministic hashing and content versioning systems.
  • Experience designing systems for horizontal scale (clustering, multi-VM, multi-threaded, or containerized workloads).
  • Experience AWS alongside Azure (the concepts transfer; multi-cloud).
  • Familiarity with semantic / cross-jurisdictional search techniques.
  • Exposure to LLM or vision-AI integration — e.g., using AI to parse document structure, generate summaries and action statements, or pre-classify regulatory changes as impacting vs. non-impacting for human review.

Responsibilities

  • Take over and fully understand an existing Python/Scrapy ingestion pipeline, its codebase, and its documented database schemas, then make it your own.
  • Build and maintain web scrapers across 50+ jurisdictions, each with its own structure, format, and obstacles — XML feeds, HTML pages, DOCX-only sources (e.g., West Virginia), and content behind Westlaw, LexisNexis, Cloudflare, and similar barriers.
  • Consolidate tooling where possible: prefer one or two robust, broadly capable scraping tools over a sprawl of point solutions, falling back to specialized tooling only for genuinely esoteric sources.
  • Normalize ingested content into a structured, plain-text format for deterministic hashing and change detection, while preserving the original, formatted HTML so documents can be displayed exactly as the issuing agency intended.
  • Faithfully preserve each source's original hierarchy (nested clauses, sub-paragraphs, lettered and numbered subsections, Roman numerals, Unicode markers, appendices, and tables), so stored content remains a true representation of the statute.
  • Maintain and extend a multi-destination architecture, including the nested-hierarchy schema used by the SaaS compliance platform (white-labeled in some deployments) and a flattened schema supporting full-text, semantic, and cross-jurisdictional search.
  • Operate and evolve change detection and notification (email, Slack, Teams), with a path toward more granular, paragraph- and clause-level change tracking and side-by-side visual diffs.
  • Design the infrastructure for cost efficiency: model how many VMs/containers are needed and for how long, automate spin-up and shutdown so nothing runs idle, and right-size compute based on real scraping timings rather than assumptions.
  • Maintain the React and Vite single-page application that displays ingested regulations, including a searchable table-of-contents navigation pattern.

We offer

  • US and EU projects based on advanced technologies.
  • Competitive compensation based on skills and experience.
  • Remote-friendly culture and no micromanagement.
  • Christmas Bonus in the amount of 50% of the monthly payment.
  • Bonuses for article writing, public talks, other activities.
  • Personalized learning program tailored to your interests and skill development.
  • Free tech webinars and meetups organized by Svitla.
  • Fun corporate online/offline celebrations and activities.
  • Awesome team, friendly and supportive community!

See also