Lead Platform & Infrastructure Engineer
Summary
Fully remote (US) lead role heading platform & infrastructure engineering for a partner company's software and AI-driven solutions, spanning cloud, single-tenant, and customer-managed environments. Day to day: Infrastructure-as-Code, Kubernetes/Docker/K3s/Helm, CI/CD automation, security hardening, observability, and incident response.
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Lead Platform & Infrastructure Engineer based in the United States.
This role offers the opportunity to lead the infrastructure foundation behind scalable, secure, and highly available software and AI-driven solutions.
You’ll design and operate modern platforms spanning cloud, single-tenant, and customer-managed environments.
The position combines platform engineering, infrastructure automation, container orchestration, security, and operational excellence.
You’ll establish Infrastructure-as-Code standards and deployment practices that make environments more consistent, reliable, and scalable.
Working closely with engineering, security, data, and AI/ML teams, you’ll help deliver resilient platforms for demanding production workloads.
You’ll also play a key role in observability, incident response, capacity planning, disaster recovery, and production readiness.
This is a high-impact opportunity for an experienced infrastructure leader who enjoys solving complex platform challenges and improving how technology is delivered and operated.
Accountabilities
-
Design, implement, and maintain scalable platform infrastructure supporting cloud, single-tenant, and customer-managed deployment environments.
-
Lead infrastructure architecture and operational practices for secure, highly available production platforms.
-
Build and manage Infrastructure-as-Code solutions, automated provisioning, configuration management, deployment automation, and environment lifecycle processes.
-
Develop and operate containerized platforms using Docker, Kubernetes, K3s, Helm, and related technologies.
-
Establish and maintain CI/CD pipelines, release automation, environment promotion strategies, rollback procedures, and modern DevOps practices.
-
Partner with security teams to implement encryption, network segmentation, secrets management, certificate management, access controls, and infrastructure-hardening standards.
-
Drive platform observability through monitoring, logging, alerting, capacity planning, reliability practices, and operational readiness.
-
Support production deployments, incident response, troubleshooting, and on-call activities while maintaining high standards of service reliability.
-
Contribute to disaster recovery, resilience, and infrastructure continuity strategies.
-
Collaborate with engineering, security, data, and AI/ML teams to ensure infrastructure effectively supports evolving application and platform requirements.
-
Identify opportunities to automate repetitive operational processes and continuously improve platform reliability, scalability, and efficiency.
-
8+ years of experience building, deploying, and operating enterprise infrastructure and platform solutions, including leadership experience in Platform Engineering, Infrastructure Engineering, DevOps, or Site Reliability Engineering.
-
Deep hands-on expertise with Docker, Kubernetes, K3s, Helm, Linux, networking, and production-grade container orchestration.
-
Strong experience with Infrastructure-as-Code and automation technologies, including Terraform, Ansible, Packer, configuration management, automated provisioning, and deployment pipelines.
-
Proven experience designing and operating secure, highly available infrastructure, including secrets management, certificate management, observability, disaster recovery, and operational resilience.
-
Experience with CI/CD platforms and modern DevOps practices, including GitHub Actions, GitLab CI, Gitea Actions, release automation, monitoring, and infrastructure lifecycle management.
-
Strong understanding of infrastructure security, access controls, network architecture, encryption, and platform hardening.
-
Demonstrated ability to troubleshoot complex production environments and respond effectively to incidents and operational challenges.
-
Strong technical leadership and communication skills, with the ability to collaborate across engineering, security, data, AI/ML, and other technical functions.
-
Experience supporting regulated industries, customer-managed or on-premises environments, object storage, GPU infrastructure, or AI/ML platform operations is preferred.
-
Strong analytical, problem-solving, prioritization, and continuous-improvement mindset.
-
Full-time, fully remote position within the United States.
-
Opportunity to lead infrastructure and platform engineering initiatives supporting modern software and AI-driven solutions.
-
High-impact technical leadership role spanning cloud, on-premises, customer-managed, and containerized environments.
-
Exposure to Kubernetes, Infrastructure-as-Code, CI/CD, automation, security, observability, and AI/ML infrastructure.
-
Opportunity to establish engineering standards and influence platform architecture and operational practices.
-
Collaboration with multidisciplinary teams across engineering, security, data, and AI/ML.
-
Opportunity to work on complex enterprise infrastructure challenges with a strong focus on reliability, scalability, and security.
-
Compensation and additional benefits are not specified in the provided job description.
Requirements
Benefits
Skills
As published by lever
Resume/CV, Full name, Email, Phone, Current location, Current company