freehire launches on Product Hunt on 26 August.

Follow →

Data Engineer

Open 26d

Summary

Builds and operates a cloud-based data platform using AWS, GCP, Kafka, Spark, and Kubernetes to run ETL pipelines, machine learning workflows, and production services.

We are looking for a skilled Data Engineer to join our team and help build, develop, and operate our next-generation data platform. In this role, you will work with modern cloud technologies, data engineering tools, and machine learning pipelines to deliver reliable, scalable, and mission-critical data solutions.

Job Responsibilities:

• Build and run Tradeweb’s data platform using such technologies as public cloud infrastructure (AWS and GCP), Kafka, Spark, databases and containers
• Develop Tradeweb’s data platform based on open source software and Cloud services
• Build and run ETL pipelines to onboard data into the platform, define schema, build DAG processing pipelines and monitor data quality.
• Help develop machine learning development framework and pipelines
• Manage and run mission crucial production services.

Key Details:

  • 🌍 Work Model: 100% Remote

  • 🕒 Working Hours: NYC Time Zone (Minimum 5h overlap required)

  • 💰 Contract & Rate: B2B, 40 - 60 USD



Requirements:

• Strong eye for detail, data precision, and data quality.
• Strong experience maintaining system stability and responsibly managing releases.
• Considerable production operations and support experience.
• Clear and effective communicator who is able to liaise with team members and end-users on requirements and issues.
• Agile, self-starter who is able to responsibly see things through to completion with minimal assistance and oversight.
• Expert level grasp of SQL and databases/persistence technologies such as MySQL, PostgreSQL, SQL Server, Snowflake, Redis, Presto, etc
• Strong grasp of Python and related ecosystems such as conda or pip.
• Experience building ETL and stream processing pipelines using Kafka, Spark, Flink, Airflow/Prefect, etc
• Experience with using AWS/GCP (S3/GCS, EC2/GCE, IAM, etc), Kubernetes and Linux in production.
• Experience with parallel and distributed computing
• Strong proclivity for automation and DevOps practices and tools such as Gitlab, Terraform, Prometheus.
• Experience with managing increasing data volume, velocity and variety.
• Ability to deal with ambiguity in a changing environment.
• At least 5-6 hours overlap starting from 9am US Eastern.

Good to have:

• Familiarity with data science stack: e.g. Jupyter, Pandas, Scikit-learn, Pytorch, MLFlow, Kubeflow etc
• Development skills in Java, Go, or Javascript
• Software builds and packaging on MS Windows
• Experience managing time series data
• Familiarity with working with open source communities
• Financial Services experience

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available