Senior Kubernetes Engineer
Summary
Senior Kubernetes engineer designs, deploys, and optimizes GPU-accelerated Kubernetes clusters for AI/ML, HPC, and LLM workloads in hybrid or on-prem environments.
Position
Senior Kubernetes Engineer – NMC² Office, Dallas
Design, implement, and optimise GPU‑accelerated container platforms at scale, enabling high‑performance workloads (AI/ML, HPC, LLM training) across hybrid or on‑prem environments.
Responsibilities
- Architect and operate Kubernetes clusters optimised for GPU workloads (leveraging NVIDIA GPU Operator, Network Operator, DCGM).
- Develop, deploy, and maintain custom Kubernetes operators and controllers to automate infrastructure services.
- Integrate NVIDIA device plugins, Mult‑Instance GPU (MIG), and GPU sharing features into the scheduling layer.
- Optimise GPU utilisation and job placement through scheduler extensions (kube‑scheduler plugins, Slurm, Volcano).
- Collaborate with HPC, ML, and DevOps teams to ensure multi‑tenant, high‑throughput cluster performance.
- Drive observability and telemetry integrations using Prometheus, Grafana, DCGM Exporter, and OpenTelemetry.
- Implement secure multi‑user and multi‑namespace GPU isolation with RBAC and policy enforcement (OPA or Gatekeeper).
- Maintain CI/CD pipelines for Kubernetes infrastructure using GitOps (ArgoCD, FluxCD).
- Contribute to infrastructure‑as‑code using Terraform, Helm, and Kustomize.
- Participate in performance tuning, incident response, and production readiness reviews.
Qualifications
- Extensive experience with Kubernetes in production environments, particularly GPU workloads (GPU Operator, device plugin, NVML, MIG, DCGM).
- Proficiency in Go or Python for operator development and Kubernetes controller logic.
- Deep understanding of Kubernetes internals (CRDs, RBAC, custom controllers, scheduler extensions).
- Experience with GPU‑intensive workloads (LLMs, training pipelines, scientific computing).
- Hands‑on experience with Helm, Kustomize, and GitOps workflows.
- Familiarity with CNI plugins, especially NVIDIA CNI and Multus.
- Experience monitoring GPU metrics and cluster health with Prometheus and DCGM Exporter.
Benefits & Perks
- Company‑Paid Lunch Stipend via GrubHub.
- 100% Employer‑Paid Medical, Dental, and Vision (High Deductible Health Plan) for employees and families.
- 16 weeks Paid Parental Leave.
- Employee Assistance Program; Life, Short‑Term Disability, Long‑Term Disability insurance.
- 401(k) with 100% match up to 6% of contributions.
- Optional Employee‑Paid Medical under PPO plan and other benefits (Health Savings Accounts, Flexible Spending Accounts, Supplemental Life Insurance, etc.).
- 25 days Paid Time Off plus 12 company holidays.
Equal Opportunity Employment
NORTHMARK STRATEGIES LLC IS AN EQUAL EMPLOYMENT OPPORTUNITY EMPLOYER. THE COMPANY DOES NOT DISCRIMINATE ON BASIS OF RACE, COLOR, RELIGION, NATIONAL ORIGIN, GENDER, AGE, SEXUAL ORIENTATION, GENDER IDENTITY, MARITAL STATUS, MENTAL OR PHYSICAL DISABILITY, OR ANY OTHER PROTECTED BASIS.
Required: Must be legally authorized to work in the United States without employer sponsorship.
The Company reserves the right to adjust, add to, or eliminate any aspect of the above description. It is not comprehensive or exhaustive.