Data Engineer
Posted Updated 1
view
Job Summary
Design, develop, and maintain scalable data infrastructure and pipelines to support data-driven products. Collaborate cross-functionally to ensure data accuracy, integration, and compliance in a cloud-based big data environment.
Responsibilities
Design, develop, and deploy data tables, views, and marts across data warehouses, operational data stores, data lakes, and data virtualization platforms to support business needs
Extract, clean, transform, and manage data flows, including web scraping techniques, to ensure high data quality and availability
Build, launch, and maintain efficient large-scale batch and real-time data pipelines using data processing frameworks for reliable data delivery
Integrate and consolidate disparate data silos in scalable and compliant ways to enable unified data access
Collaborate with project managers, data architects, business analysts, frontend developers, designers, and data analysts to build scalable, data-driven products
Develop backend APIs and manage databases to support application functionality and performance
Apply Agile methodologies with continuous integration and delivery to accelerate development cycles
Engage in pair programming and code reviews to maintain code quality and knowledge sharing
Required competencies and certifications
Proficient in data cleaning and transformation using tools such as SQL, pandas, or R to ensure data accuracy and consistency
Skilled in building ETL pipelines with technologies like SQL Server Integration Services (SSIS), AWS Database Migration Services (DMS), Python, AWS Lambda, ECS Container tasks, Event bridge, AWS Glue, or Spring
Experienced in database design and management across various systems including SQL, PostgreSQL, AWS S3, Athena, MongoDB, Postgres/GIS, MySQL, SQLite, Volt DB, and Cassandra
Knowledgeable in cloud platforms such as AWS, Azure, or Google Cloud for big data engineering
Experienced in building production-grade data pipelines and ETL/ELT data integration workflows
Understanding of system design, data structures, and algorithms to optimize data processing
Familiar with data modeling, data access methods, and storage infrastructures like Data Mart, Data Lake, Data Virtualization, and Data Warehouse for efficient data retrieval
Familiar with REST APIs and web protocols to support data integration and application communication
Familiar with big data frameworks and tools such as Hadoop, Spark, Kafka, and RabbitMQ
Experienced with web scraping technologies including Beautiful Soup, Casper.JS, Phantom.JS, Selenium, and Node.js for customized data extraction
Knowledgeable about data governance policies, access control, and security best practices to ensure data compliance
Comfortable scripting in languages such as SQL and Python for automation and data manipulation
Proficient in both Windows and Linux development environments for versatile deployment and testing
Demonstrates interest in bridging engineering and analytics teams to enhance data-driven decision-making
Skills
- Agile
- Analytics
- API
- Athena
- Automation
- AWS
- Aws Glue
- Azure
- Cassandra
- CI/CD
- Cloud
- Data Engineering
- Data Governance
- Data Lake
- Data Modeling
- Data Pipelines
- Data Quality
- Data Warehousing
- ECS
- ELT
- ETL
- GCP
- GIS
- Hadoop
- Kafka
- Lambda
- Linux
- MongoDB
- MySQL
- Node.js
- pandas
- PostgreSQL
- Python
- RabbitMQ
- REST
- Selenium
- Spark
- Spring
- SQL
- SQL Server
- SQLite
- SSIS
- Virtualization