Point your AI agent at freehire and let it find you a job.

Get the CLI →

Dataeconomy

NewBe an early applicant

MLOps / Serving Engineer

Posted 1 view
Discussion

Summary

Designs and runs production serving infrastructure for fine-tuned LLMs on AWS: deploying models with vLLM, TensorRT-LLM or Triton, building shadow-mode and staged canary rollouts with automated rollback, optimizing inference (KV-cache, INT8), and building monitoring and auto-scaling. Hybrid role in Hyderabad or Pune.

Job Title: MLOps / Serving Engineer
Experience : 5+ years
Location : Hyderabad OR Pune
Notice Period: 0-30 days
Work mode - Hybrid

We are seeking an experienced MLOps / Serving Engineer who can design and operate the production serving infrastructure for fine-tuned LLMs on AWS — optimised inference engines, shadow-mode and staged rollout pipelines, monitoring dashboards, and the path from experimental model to full production traffic.
Key Responsibilities:
  • Deploy fine-tuned LLMs using vLLM, TensorRT-LLM, or Triton with continuous batching on AWS GPU instances
  • Build shadow-mode deployment: run fine-tuned model alongside production, log comparison data without impacting live traffic
  • Execute staged rollout: canary (5%) → gradual ramp (25% → 50% → 100%) with automated rollback on quality degradation
  • Optimize inference for input-heavy workloads (~17K token inputs, ~130 token outputs): prefill throughput, KV-cache, INT8 quantization
  • Build monitoring dashboards: latency, throughput, accuracy metrics, cost per request
  • Design auto-scaling; implement high-availability (2× instances); automated rollback triggers on end-to-end quality metrics

Requirements

  • 5+ years MLOps or ML infrastructure engineering
  • Hands-on with vLLM, TensorRT-LLM, or Triton Inference Server
  • Deep familiarity with g5, p4de, p5 instance families, EC2 auto-scaling
  • Have worked on Deployment patterns like Shadow-mode, canary, A/B traffic routing, automated rollback
  • Experience onto Continuous batching, INT8 quantization, KV-cache management
  • Expertise on Docker, Kubernetes (EKS) for ML workloads
  • Worked on CloudWatch, Prometheus, Grafana



Benefits

  • Comprehensive Medical Coverage:
    Health insurance of INR 5.0 Lakhs for you and your family (up to 6 members), ensuring complete peace of mind.
  • Robust Protection Plans:
    Group Personal Accident Insurance and Group Term Life Insurance to safeguard you and your loved ones.
  • Retirement Benefits:
    PF and Gratuity provided as per standard government regulations.
  • Flexible Work Options:
    Enjoy hybrid work arrangements & flexible working hours
  • Generous Leave Policy:
    21 days of annual leave, in addition to 10 company-declared holidays.
  • Employee Well-being Spaces:
    Access to a dedicated break-out area with round-the-clock refreshments for relaxation and rejuvenation.

Skills

See also

DevOps jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available