Senior AI Platform Engineer (MLOps & Data Science Infrastructure
Summary
Builds and maintains end-to-end MLOps infrastructure and AI pipelines — including LLM/Generative AI platforms — so ML models (churn, pricing, anomaly detection, GenAI) run reliably in real-time enterprise workflows. Core stack: Python, SQL, Spark, MLflow/Airflow/Kubeflow, cloud (Azure/AWS/GCP), Docker, Kubernetes.
Model Lifecycle Management: Implement robust MLOps practices, including automated model training, orchestration, deployment, versioning, monitoring, and continuous integration/continuous deployment (CI/CD) for ML workflows.
LLM & Generative AI Platforming: Build infrastructure supporting fine-tuning, evaluation frameworks, prompt engineering pipelines, redteaming, and real-time inference monitoring for Large Language Models (e.g., GPT, LLaMA).
2. Model Governance, Quality & Optimization Governance & Observability: Drive best practices around model governance, explainability, lineage tracking, and performance monitoring to prevent model drift and ensure compliance with enterprise policies and regulations.
Scalable Data Systems: Process and manipulate large-scale multidomain datasets (network, customer, financial, operational) to support real-time inference, anomaly detection, digital twins, and optimization models.
Experimentation Frameworks: Build and optimize continuous experimentation environments, enabling scalable A/B testing, uplift modeling, causal inference testing, and dynamic feature engineering.
3. Engineering Excellence & Collaboration Software Engineering Best Practices: Champion high code quality, modular software design, reproducibility, automated testing, and version control standards across the team.
Technical Mentorship: Provide technical guidance, pair programming, and constructive code reviews for junior team members to raise overall engineering capability.
Cross-Functional Partnering: Translate complex enterprise business challenges (such as pricing optimization, churn prevention, and resource allocation) into scalable AI platform architectures.
Skills & Qualifications Bachelor's or Postgraduate degree in Computer Science, Data Science, Mathematics, Statistics, Software Engineering, or a related quantitative field.
At least 3-5 years of hands-on experience developing, deploying, and maintaining production machine learning models, MLOps platforms, or scalable data science pipelines.
Technical Capabilities ML & Generative AI Frameworks: Deep understanding of machine learning algorithms (supervised learning, time series forecasting, neural networks, clustering) alongside hands-on experience finetuning, serving, and evaluating LLMs.
MLOps & Lifecycle Tools: Proficiency with ML orchestration and tracking tools (e.g., MLflow, Airflow, Kubeflow, Databricks).
Data Processing & Engineering: Expertise in Python, SQL, Spark, and big data technologies (Hadoop/Hive) for efficient data manipulation and feature engineering.
Cloud & DevOps Platforming: Hands-on experience with major cloud platforms (Azure, AWS, or GCP) and software development tools (Git, GitHub, GitLab, Docker, Kubernetes).
Code & Architecture Quality: Solid foundation in modular coding, unit/integration testing, CI/CD practices, and scalable system design.