Machine Learning Engineer - AI
Summary
Fine-tune and optimize Small Language Models (SLMs) for edge, mobile, and local deployment using Hugging Face, LoRA, QLoRA, and PEFT, while building end-to-end MLOps pipelines spanning the full ML lifecycle.
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Machine Learning Engineer - AI based in India.
This is an engineering-focused opportunity for a machine learning professional working at the intersection of small language models, model optimization, and production AI. You will fine-tune and optimize lightweight models designed for efficient inference across edge, mobile, and local environments. The role spans the full machine learning lifecycle, from data ingestion and experimentation through deployment, monitoring, and continuous improvement. You will work with modern frameworks and techniques such as Hugging Face, TRL, LoRA, QLoRA, and PEFT to build efficient AI solutions. A strong focus is placed on performance, with strict attention to latency, model accuracy, and hardware utilization. You will also contribute to scalable MLOps practices that make AI models reliable and production-ready.
Accountabilities:
- Fine-tune and train Small Language Models (SLMs) using Hugging Face, TRL, and parameter-efficient adaptation techniques such as LoRA, QLoRA, and PEFT.
- Experiment with model architectures, training strategies, and datasets to improve model quality and task-specific performance.
- Optimize models for efficient inference using techniques including quantization, pruning, knowledge distillation, and other model compression approaches.
- Prepare and deploy lightweight AI models to edge devices, mobile environments, local servers, and other resource-constrained platforms.
- Design and implement end-to-end MLOps pipelines covering data ingestion, preprocessing, experimentation, model training, validation, packaging, deployment, and monitoring.
- Build reliable and repeatable workflows that support efficient model development and production deployment.
- Monitor deployed models for accuracy, latency, resource consumption, and CPU/GPU utilization.
- Develop and maintain model benchmarking frameworks and custom evaluation suites to measure model quality and performance.
- Analyze production performance and identify opportunities to improve model efficiency, reliability, and scalability.
- Work with cross-functional engineering teams to integrate machine learning models into real-world products and environments.
- Contribute to ML engineering best practices around experimentation, versioning, deployment, observability, and continuous improvement.
- Explore emerging techniques and tooling for efficient AI inference, edge deployment, and production machine learning.
- Hands-on experience developing, training, and fine-tuning Small Language Models or other transformer-based models.
- Strong practical knowledge of Hugging Face and modern model adaptation techniques, including LoRA, QLoRA, and PEFT.
- Experience optimizing machine learning models for efficient inference through quantization, pruning, knowledge distillation, or similar techniques.
- Experience deploying machine learning models to edge devices, mobile platforms, local servers, or other environments with constrained compute and strict latency requirements.
- Strong understanding of end-to-end MLOps practices, from data ingestion and model experimentation through deployment and production monitoring.
- Experience monitoring model accuracy, inference latency, and CPU/GPU or other hardware utilization in production.
- Ability to develop meaningful model evaluation and benchmarking frameworks and use data-driven results to improve model performance.
- Strong software engineering, debugging, analytical, and problem-solving skills.
- Ability to work effectively in a collaborative, fast-moving environment and communicate technical concepts clearly.
- Strong ownership mindset and willingness to continuously learn new machine learning technologies and deployment techniques.
- Experience with ONNX export and cross-platform inference is a plus.
- Experience deploying AI solutions to edge or mobile environments is preferred.
- Familiarity with MLOps tooling for experiment tracking, model registries, and ML-focused CI/CD pipelines is advantageous.
- Opportunity to work on modern AI and machine learning technologies with a strong focus on Small Language Models.
- Hands-on exposure to model fine-tuning, optimization, compression, and efficient inference.
- Opportunity to build production-grade MLOps pipelines spanning the full machine learning lifecycle.
- Experience working with edge, mobile, and local AI deployment environments.
- Exposure to technologies including Hugging Face, TRL, LoRA, QLoRA, PEFT, and modern MLOps tooling.
- Opportunity to contribute to scalable AI solutions designed for real-world production environments.
- Collaborative environment with opportunities to work alongside experienced engineering and technology professionals.
- Strong focus on continuous learning, experimentation, and adoption of emerging AI technologies.
- Inclusive workplace culture that values diverse perspectives, collaboration, and individual contributions.