Data Engineer – AI, Java, Python, Spark
Summary
Build and optimize scalable data pipelines and cloud-native services using Python/Java, Spark, and Kubernetes to power GenAI applications and LLM workflows.
- Develop, test and maintain high-quality, production-ready software
- Design and implement large-scale data pipelines and distributed processing systems
- Build scalable cloud-native services and platforms
- Provide technical leadership for cross-team initiatives and complex engineering projects
- Design and develop reusable libraries, frameworks and platform components
- Optimize distributed data processing workloads for performance, scalability and reliability
- Work with Databricks, Apache Spark and Snowflake data platforms
- Develop and deploy applications using Python and/or Java
- Build and operate containerized workloads using Kubernetes and cloud-native technologies
- Design and implement GenAI/LLM-based applications and services
- Use LangChain and LangGraph for LLM orchestration and agentic workflows
- Collaborate with data scientists, software engineers, architects and product teams
- Establish engineering best practices around testing, observability, reliability and deployment
Requirements
- 5+ years of professional software/data engineering experience
- Strong hands-on experience with Python and/or Java
- Strong experience with Apache Spark and distributed data processing
- Experience with Databricks and/or modern lakehouse platforms
- Experience with Snowflake or comparable cloud data warehouses
- Practical experience with Kubernetes and cloud-native technologies
- Experience designing and maintaining large-scale data pipelines
- Strong understanding of distributed systems, scalability and production engineering
- Experience developing ML/AI or GenAI applications
- Experience with LLM-based applications, RAG, AI agents or LLM orchestration
- Familiarity with LangChain, LangGraph or similar GenAI frameworks
- Strong software engineering fundamentals including testing, code quality and system design
Core Competencies
Demonstrates expertise in developing and maintaining high-quality software, with a strong focus on building scalable cloud-native services and optimizing distributed data processing. Proficient in Python, Java, and modern data platforms, with a solid understanding of GenAI applications and engineering best practices.
Highest-signal resume keywords
- Python Development
- Java Development
- Apache Spark
- Kubernetes
- GenAI Applications
ATS Optimization Keywords
Hard Skills
- Software Engineering
- Data Engineering
- Distributed Systems
- Data Pipeline Design
- Performance Optimization
- Testing and Code Quality
- System Design
- ML/AI Development
- LLM Orchestration
- Cloud-Native Technologies
Soft Skills
- Technical Leadership
- Collaboration
Industry Keywords
- Cloud Data Warehouses
- Lakehouse Platforms
- Distributed Processing Systems
- Engineering Best Practices
Tools & Technologies
- Databricks
- Snowflake
- LangChain
- LangGraph