Data Engineer
Summary
Data Engineer in Boston building scalable, 24×7 Hadoop/Hive data pipelines that process billions of events per day, supporting real-time bidding, reporting, and data distribution. Core stack includes SQL (MySQL/Postgres/Redshift/Hive), Java, Linux/shell, plus Ruby, Go, R, and JavaScript/D3 for reporting and visualization.
• Experience with Linux or Unix based systems – including Bourne shell, cron and other Unix utilities.
• Strong software development skills, including experience with Java.
• Degree in Computer Science or a related field
• Experience with Hadoop/MapReduce and/or EMR, including experience in developing MapReduce Jobs in Java or developing Hive UDF.
• ETL experience maintaining multiple data systems.
• Experience with Ooozie or other Hadoop workflow solutions and experience developing complex data processing pipelines, including experience developing regressions tests and deployment strategies for such environments.
• Experience with data reporting solutions – either developed in house or with 3rd party solutions.
• Working experience developing and supporting 24×7 production data services and pipelines on Linux systems – including experience being on-call supporting such services. Experience with AWS preferred.
• Processing events in near real-time
• Building for the fragility of cloud and distributed services
• Complex processing of large amounts of data in an efficient manner.
• Reporting, distribution of data, data analysis, data visualization and machine learning algorithms.
• Low latency data stores for use in bidding or algo optimization.