Software Engineer (AI Training)
Summary
Builds and optimizes AI training pipelines, annotation workflows, and model evaluation tools to improve machine learning performance at scale, collaborating with data scientists and researchers.
- Design, build, and maintain data pipelines that support AI model training and evaluation workflows.
- Develop tooling and automation to streamline data collection, labeling, and preprocessing processes.
- Collaborate with ML researchers and data scientists to understand model requirements and translate them into engineering solutions.
- Implement quality control mechanisms to ensure training data accuracy, consistency, and integrity.
- Build and maintain internal platforms and APIs that support annotation and human feedback workflows (RLHF).
- Monitor and optimize pipeline performance, scalability, and reliability in production environments.
- Contribute to documentation, code reviews, and engineering best practices across the team.
- 3–5 years of software engineering experience with exposure to AI, ML, or data engineering environments.
- Proficiency in Python and at least one additional language (Go, Java, or TypeScript).
- Experience building and maintaining data pipelines using tools such as Apache Airflow, Spark, or equivalent.
- Familiarity with machine learning concepts, model training workflows, and evaluation methodologies.
- Experience working with large datasets, data annotation platforms, or labeling tools.
- Strong understanding of REST APIs, microservices, and cloud infrastructure (AWS, GCP, or Azure).
- Excellent problem-solving skills and ability to work independently in a remote environment.
- Experience with Reinforcement Learning from Human Feedback (RLHF) or similar human-in-the-loop training methodologies.
- Familiarity with LLM training infrastructure and model evaluation frameworks.
- Experience with vector databases, embedding, or retrieval-augmented generation (RAG) pipelines.
- Prior experience at an AI-focused company or research lab.