Data Engineer
Summary
Build and maintain scalable data pipelines and infrastructure to power analytics and compliance in a cloud-based big data environment.
Job Summary
Design, develop, and maintain scalable data infrastructure and pipelines to support data-driven products. Collaborate cross-functionally to ensure data accuracy, integration, and compliance in a cloud-based big data environment.
Responsibilities
- Design, develop, and deploy data tables, views, and marts across data warehouses, operational data stores, data lakes, and data virtualization platforms to support business needs
- Extract, clean, transform, and manage data flows, including web scraping techniques, to ensure high data quality and availability
- Build, launch, and maintain efficient large-scale batch and real-time data pipelines using data processing frameworks for reliable data delivery
- Integrate and consolidate disparate data silos in scalable and compliant ways to enable unified data access
- Collaborate with project managers, data architects, business analysts, frontend developers, designers, and data analysts to build scalable, data-driven products
- Develop backend APIs and manage databases to support application functionality and performance
- Apply Agile methodologies with continuous integration and delivery to accelerate development cycles
- Engage in pair programming and code reviews to maintain code quality and knowledge sharing
Required competencies and certifications
- Proficient in data cleaning and transformation using tools such as SQL, pandas, or R to ensure data accuracy and consistency
- Skilled in building ETL pipelines with technologies like SQL Server Integration Services (SSIS), AWS Database Migration Services (DMS), Python, AWS Lambda, ECS Container tasks, Event bridge, AWS Glue, or Spring
- Experienced in database design and management across various systems including SQL, PostgreSQL, AWS S3, Athena, MongoDB, Postgres/GIS, MySQL, SQLite, Volt DB, and Cassandra
- Knowledgeable in cloud platforms such as AWS, Azure, or Google Cloud for big data engineering
- Experienced in building production-grade data pipelines and ETL/ELT data integration workflows
- Understanding of system design, data structures, and algorithms to optimize data processing
- Familiar with data modeling, data access methods, and storage infrastructures like Data Mart, Data Lake, Data Virtualization, and Data Warehouse for efficient data retrieval
- Familiar with REST APIs and web protocols to support data integration and application communication
- Familiar with big data frameworks and tools such as Hadoop, Spark, Kafka, and RabbitMQ
- Experienced with web scraping technologies including Beautiful Soup, Casper.JS, Phantom.JS, Selenium, and Node.js for customized data extraction
- Knowledgeable about data governance policies, access control, and security best practices to ensure data compliance
- Comfortable scripting in languages such as SQL and Python for automation and data manipulation
- Proficient in both Windows and Linux development environments for versatile deployment and testing
- Demonstrates interest in bridging engineering and analytics teams to enhance data-driven decision-making