Data Engineer

Summary

Data Engineer responsible for designing, maintaining, and improving data lakes, pipelines, batch and real-time processing systems, and metadata infrastructure using Python, AWS services, and big data technologies.

As a Data Engineer you'd be working with us to design, maintain, and improve various analytical and operational servicesand infrastructure which are critical for many other functions within the organization. These include the data lake,operational databases, data pipelines, large-scale batch and real-time data processing systems, a metadata and lineagerepository, which all work in concert to provide the company with accurate, timely, and actionable metrics and insightsto grow and improve our business using data. You may be collaborating with our data science team to design andimplement processes to structure our data schemas and design data models, working with our product teams to integratenew data sources, or pairing with other data engineers to bring to fruition cutting-edge technologies in the data space.

Our Ideal Candidate

We expect candidates to have in-depth experience in some of the following skills and technologies and be motivated tobuild up experience and fill any gaps in knowledge on the job. More importantly, we seek people who are highly logical,with a balance of respect for best practices and using their own critical thinking, adaptable to new situations, capable ofworking independently to deliver projects end-to-end, communicates well in English, collaborates effectively withteammates and stakeholders, and eager to be on a high-performing team, taking their careers to the next level with us.

Highly relevant: (ideally familiar with at least one of the technologies in most of the below categories)

  • General computing concepts and expertise: Unix environments, networking, distributed and cloud computing
  • Python frameworks and tools: pip, pytest, boto3, pyspark, pylint, pandas, scikit-learn, keras
  • Workflow scheduling and monitoring tools: Apache Airflow, Luigi, AWS Batch
  • Columnar and big data databases: Athena, Redshift, Vertica, Hive/Hadoop
  • Container management and orchestration: Docker, Docker Swarm, ECS, EKS/Kubernetes, Mesos CI / CD tools: CircleCI, Jenkins, TravisCI, Spinnaker, AWS CodePipeline
  • Distributed messaging and event streaming systems: Kafka, Pulsar, RabbitMQ, Google Pub/Sub
  • General AWS or Cloud services: Glue, EMR, EC2, ELB, EFS, S3, Lambda, API Gateway, IAM, Cloudwatch
  • Version control: git commands, branching strategies, collaboration etiquette, documentation best practices Agile/Lean project methodologies and rituals: Scrum, Kanban

See also

Data Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available