freehire launches on Product Hunt on 26 August.

Follow →

Data Engineer – AI, Java, Python, Spark

Summary

Build and optimize scalable data pipelines and cloud-native services using Python/Java, Spark, and Kubernetes to power GenAI applications and LLM workflows.

  • Develop, test and maintain high-quality, production-ready software
  • Design and implement large-scale data pipelines and distributed processing systems
  • Build scalable cloud-native services and platforms
  • Provide technical leadership for cross-team initiatives and complex engineering projects
  • Design and develop reusable libraries, frameworks and platform components
  • Optimize distributed data processing workloads for performance, scalability and reliability
  • Work with Databricks, Apache Spark and Snowflake data platforms
  • Develop and deploy applications using Python and/or Java
  • Build and operate containerized workloads using Kubernetes and cloud-native technologies
  • Design and implement GenAI/LLM-based applications and services
  • Use LangChain and LangGraph for LLM orchestration and agentic workflows
  • Collaborate with data scientists, software engineers, architects and product teams
  • Establish engineering best practices around testing, observability, reliability and deployment

Requirements

  • 5+ years of professional software/data engineering experience
  • Strong hands-on experience with Python and/or Java
  • Strong experience with Apache Spark and distributed data processing
  • Experience with Databricks and/or modern lakehouse platforms
  • Experience with Snowflake or comparable cloud data warehouses
  • Practical experience with Kubernetes and cloud-native technologies
  • Experience designing and maintaining large-scale data pipelines
  • Strong understanding of distributed systems, scalability and production engineering
  • Experience developing ML/AI or GenAI applications
  • Experience with LLM-based applications, RAG, AI agents or LLM orchestration
  • Familiarity with LangChain, LangGraph or similar GenAI frameworks
  • Strong software engineering fundamentals including testing, code quality and system design

Core Competencies

Demonstrates expertise in developing and maintaining high-quality software, with a strong focus on building scalable cloud-native services and optimizing distributed data processing. Proficient in Python, Java, and modern data platforms, with a solid understanding of GenAI applications and engineering best practices.

Highest-signal resume keywords

  • Python Development
  • Java Development
  • Apache Spark
  • Kubernetes
  • GenAI Applications

ATS Optimization Keywords

Hard Skills

  • Software Engineering
  • Data Engineering
  • Distributed Systems
  • Data Pipeline Design
  • Performance Optimization
  • Testing and Code Quality
  • System Design
  • ML/AI Development
  • LLM Orchestration
  • Cloud-Native Technologies

Soft Skills

  • Technical Leadership
  • Collaboration

Industry Keywords

  • Cloud Data Warehouses
  • Lakehouse Platforms
  • Distributed Processing Systems
  • Engineering Best Practices

Tools & Technologies

  • Databricks
  • Snowflake
  • LangChain
  • LangGraph

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available