AI SRE
Southeast Asian tech company building AI infrastructure and large language model products.
Focused on developing AI solutions tailored to regional markets and use cases.
Project/Goal:
Build and maintain infrastructure powering AI applications and model training, supporting workloads reliably and securely at scale across cloud and on-prem environments.
Key Responsibilities:
Build and maintain infrastructure for model hosting, prompt execution, and training workflows.
Operate and optimize model-serving systems (e.g. vLLM, Triton, OpenAI-style proxies).
Implement secure API gateways; manage token usage, routing, and fallback logic.
Develop and support training pipelines, GPU scheduling, and experiment tracking.
Maintain CI/CD systems, observability tooling, and infrastructure documentation.
Skills Requirements:
6+ years in infrastructure engineering, DevOps, or ML systems.
Strong command of Kubernetes, Terraform, and cloud-native architecture (AWS/Azure/GCP).
Experience with containerization, CI/CD, and API security practices.
Prior exposure to model hosting or ML pipeline orchestration.
Understanding of networking (VPNs, VNets, hybrid connectivity) and cross-platform security best practices.
Experience with on-prem infrastructure (networking, storage hardware).
Good to Have:
GPU resource orchestration or Kubeflow experience.
Familiarity with inference servers (vLLM, Triton, TGI, TorchServe).
Cost telemetry / resource budgeting for model traffic.
Security mindset — IAM, logging, compliance experience.
Familiarity with compliance frameworks (SOC2, GDPR, HIPAA).
Background in database management across platforms.
Work Location:
Mid-to-senior level, 6+ years relevant experience.