freehire launches on Product Hunt on 26 August.

Follow →

Data Engineer

Summary

Design and maintain scalable data pipelines, ETL processes, and storage systems using Python, Spark, and cloud platforms like AWS to ensure data quality and accessibility for business insights.

  • Design, build, and maintain scalable, secure data pipelines and storage systems; ensure data quality through ETL processes and regular checks.

  • Implement policies and practices to control, optimize, and secure data assets, ensuring data integrity and accessibility.

  • Develop and maintain data models, structures, and databases to meet business needs; communicate data architecture effectively.

  • Develop, test, and maintain scripts and programs to automate data processing and pipelines, adhering to industry standards.

  • Create and operationalize data visualization solutions to simplify complex data for stakeholders and decision‑making.

  • Work with cross‑functional teams to gather data requirements, optimize existing processes, and deliver ad hoc reports and insights

JOB SPECIFICATIONS:

  • Education – At least graduate with a Bachelor’s or Master's Degree in IT, Computer Science, Engineering, or any related course.

  • Related Work Experience – at least 5 years of experience in data engineering, data analytics, or related fields.

    • Proficiency in Python and Java, with experience in other programming languages as a plus.

    • Hands‑on experience with data manipulation tools (e.g., pandas, dplyr, or Spark).

    • Proficiency in Extract, Transform, and Load (ETL) processes for efficient data pipeline management.

    • Expertise in both SQL and NoSQL query languages for database interaction and management.

    • Experience with big data storage and processing solutions, such as MongoDB, Spark, Hive, Snowflake, Redshift, or similar technologies.

    • Experience with cloud‑based or server‑based data processing environments (e.g., AWS, Azure, GCP).

    • With experience in APACHE NIFI

  • Knowledge – Knowledgeable in the following:

    • Comprehensive understanding of data manipulation tools such as pandas, dplyr, and Spark.

    • In‑depth knowledge of big data frameworks and tools like Apache Spark and Hadoop.

    • Familiarity with data warehousing services like AWS Redshift, Snowflake, or similar solutions.

    • Proficiency in AWS Cloud Services, particularly AWS Glue and AWS Lake Formation.

    • Familiarity with business intelligence tools such as Tableau, Power BI, QuickSight, or Google Data Studio.

    • Awareness of data visualization libraries and packages like Dash, Plotly, Matplotlib, ggplot, and Folium.

    • Understanding of machine learning libraries and tools (e.g., scikit‑learn, caret, MATLAB) is a plus.

    • Knowledge of data governance, quality control, and security best practices.

  • Skills

    • Problem‑Solving: Ability to address technical challenges and deliver efficient data solutions.

    • Communication: Clear and concise communication skills for collaborating with team members and stakeholders.

    • Collaboration: Ability to work effectively within a team environment to achieve shared goals.

    • Time Management: Capacity to manage tasks and meet deadlines in a structured and timely manner.

    • Attention to Detail: Careful and accurate handling of data to ensure quality and integrity.

    • Client‑Focus: Ability to understand business needs and align data solutions to support decision‑making and strategic objectives.

    • Adaptability: Flexible and open to learning new tools, technologies, and processes in a rapidly changing environment.

See also