Reporting and Analytics Developer / Data Engineer 1026
Summary
A Reporting and Analytics Developer/Data Engineer designs and deploys data systems, builds ETL pipelines, and collaborates with cross-functional teams to develop scalable data-driven products, working with data warehouses, lakes, and cloud technologies like AWS and GCP.
Key Responsibilities
- Design, develop, and deploy data tables, views, and marts in data warehouses, operational data stores, data lakes, and data virtualization.
- Perform data extraction, cleaning, transformation, and data flow management. Web scraping may also be a part of the work scope in data extraction.
- Design, build, launch, and maintain efficient and reliable large-scale batch and real-time data pipelines with data processing frameworks.
- Integrate and collate data silos in a manner that is both scalable and compliant.
- Collaborate with Project Managers, Data Architects, Business Analysts, Frontend Developers, Designers, and Data Analysts to build scalable, data-driven products.
- Be responsible for developing backend APIs and working on databases to support applications.
- Work in an Agile environment that practices Continuous Integration and Continuous Delivery.
- Work closely with fellow developers through pair programming and code review processes.
Experience and Skills Needed
- Proficient in general data cleaning and transformation (e.g., SQL, pandas, R, etc.) to ensure data accuracy and consistency.
- Proficient in building ETL pipelines (e.g., SQL Server Integration Services (SSIS), AWS Database Migration Service (DMS), Python, AWS Lambda, ECS Container Tasks, EventBridge, AWS Glue, Spring).
- Proficient in database design and various databases (e.g., SQL, PostgreSQL, AWS S3, Athena, MongoDB, PostGIS, MySQL, SQLite, VoltDB, Cassandra, etc.).
- Experience in cloud technologies such as GCP, GCC (i.e., AWS, Azure, Google Cloud).
- Experience and passion for data engineering in a big data environment using cloud platforms such as GCP, GCC (i.e., AWS, Azure, Google Cloud).
- Experience with building production-grade data pipelines and ETL/ELT data integration.
- Knowledge of system design, data structures, and algorithms.
- Familiar with data modelling, data access, and data storage infrastructure such as Data Marts, Data Lakes, Data Virtualization, and Data Warehouses for efficient storage and retrieval.
- Familiar with REST APIs and web requests/protocols in general.
- Familiar with big data frameworks and tools (e.g., Hadoop, Spark, Kafka, RabbitMQ).
- Familiar with W3C Document Object Model and customised web scraping (e.g., BeautifulSoup, CasperJS, PhantomJS, Selenium, Node.js, etc.).
- Familiar with data governance policies, access control, and security best practices.
- Comfortable with at least one scripting language (e.g., SQL, Python).
- Comfortable working in both Windows and Linux development environments.
- Interest in being the bridge between engineering and analytics.
Bonus Experience (Added Advantage)
- Experience building data engineering pipelines that require integration with search indexes.
- Experience with Airflow and RDBMS integration and implementation (e.g., MySQL).
- Experience with either Snowflake, Databricks, or an equivalent provider