Data Engineer (Python, SQL & Web Scraping)
Summary
Senior Data Engineer building ETL pipelines and web scraping solutions using Python and SQL to collect, integrate, and optimize data in PostgreSQL.
We are looking for a Senior Data Engineer to join our R&D Center, which is officially approved by the Turkish Ministry of Industry and Technology. The successful candidate will be responsible for collecting data from diverse web sources, integrating it with existing datasets, validating its quality and accuracy, and loading it into our databases.
You will work across web scraping, data integration, ETL pipelines, data quality, and database optimization, with a strong focus on building reliable and scalable data processes.
Responsibilities
- Collect data from various web sources and develop and improve existing scraping processes using tools such as Scrapy, Selenium, Requests, and similar technologies.
- Monitor scraping processes and analyze and resolve issues as they arise.
- Compare and integrate data collected from new sources with existing datasets.
- Perform data matching, duplicate detection, and data quality checks.
- Develop, maintain, and monitor ETL pipelines.
- Improve existing data models and SQL queries.
- Design and implement efficient and scalable data processing solutions, leveraging database capabilities whenever appropriate rather than relying solely on application-level processing.
- Optimize data processing with consideration for data volume, query performance, and scalability.
Technical Requirements
Required
- 5+ years of experience in Data Engineering and/or ETL.
- Strong to advanced SQL skills.
- Experience with PostgreSQL or a similar relational database.
- Hands-on experience with web scraping.
- Experience processing REST APIs, JSON, and HTML data.
- Experience with Git.
- Comfortable working in Linux environments.
Nice to Have
- Scrapy
- Docker
- Elasticsearch / OpenSearch
The Approach We're Looking For
When working with large datasets, we value engineers who go beyond simply processing everything through Python/Pandas.
We are particularly interested in candidates who understand when to leverage SQL and database capabilities to perform operations efficiently at the database layer, resulting in solutions that are performant, scalable, and maintainable.
We are looking for someone who can make thoughtful engineering trade-offs based on data volume, performance requirements, and the capabilities of the underlying database.