Manager, Infrastructure
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Manager, Infrastructure based in the United States.
This is a hands-on infrastructure leadership role overseeing a distributed, high-performance bare-metal environment spanning four global data centers.
You will lead an experienced InfraOps and network engineering team while owning operational reliability across the entire infrastructure footprint.
The role combines people leadership with deep technical responsibility for capacity planning, hardware lifecycle, networking, Linux operations, and distributed systems.
You will drive incident response, improve operational maturity, and ensure the platform consistently meets demanding latency and throughput requirements.
You will also manage data center vendors, procurement, budgets, remote-hands operations, and infrastructure planning across physical and hybrid environments.
Working closely with security and engineering leadership, you will help shape infrastructure strategy in a fast-moving, M&A-oriented organization.
This opportunity is ideal for an infrastructure leader who enjoys owning complex systems where performance, reliability, and operational excellence directly impact the business.
Accountabilities:
- Lead and develop a distributed infrastructure operations team across US and APAC time zones, including regional data center owners and network engineering.
- Establish and maintain operational rhythms including DevOps check-ins, alert reviews, on-call coverage, incident reviews, and performance tracking.
- Own P1/P2 incident response from detection through resolution, with accountability for reducing MTTR, improving runbooks, and maintaining high-quality alerting.
- Manage capacity planning and the full hardware lifecycle across four data centers, including Dell and Supermicro procurement, GPU expansion, colocation power and space, and remote-hands logistics.
- Lead annual cloud-versus-colocation evaluations and provide recommendations based on performance, capacity, cost, and operational requirements.
- Oversee the Linux and infrastructure platform, including Ubuntu/systemd fleets, FreeIPA, Ansible/Salt, MAAS provisioning, and Zabbix monitoring.
- Manage high-performance spine-leaf networking, BGP and peering, transit providers, and low-latency network optimization.
- Oversee stateful distributed data platforms such as Aerospike, Kafka, ClickHouse, and Hadoop/HDFS, including capacity management, migrations, evictions, and performance tuning.
- Own colocation, networking, licensing, procurement, and infrastructure vendor relationships and associated budgets.
- Partner with security and compliance teams on infrastructure hardening, access reviews, SOC 2 Type 2 evidence, and related controls.
- Operate effectively across a multi-entity environment and help navigate infrastructure requirements associated with organizational growth and acquisitions.
- Support and develop infrastructure engineers while maintaining a culture of ownership, proactive communication, technical excellence, and continuous improvement.
- 6–8 years of experience in infrastructure or data center operations, including at least 2 years managing engineers or technical teams.
- Strong bare-metal and colocation experience, including capacity planning, hardware procurement, owned-cage operations, remote-hands coordination, and physical data center logistics.
- Solid networking expertise at scale, including spine-leaf architectures, BGP, high-capacity peering, transit providers, and low-latency network optimization.
- Deep Linux operations experience, including systemd, netplan, FreeIPA, Chrony, Ansible and/or Salt, Zabbix, and management of fleets containing hundreds of servers.
- Experience operating large-scale stateful distributed systems such as Aerospike, Cassandra, Scylla, Kafka, or ClickHouse within demanding latency and capacity constraints.
- Demonstrated ownership of P1/P2 incidents, on-call programs, postmortems, operational runbooks, and alert-management practices.
- Experience leading and developing experienced technical teams across multiple geographic regions and time zones.
- Strong understanding of hardware lifecycle management, infrastructure procurement, vendor relationships, and operational budgeting.
- Excellent analytical, troubleshooting, and decision-making abilities, with a practical approach to complex infrastructure problems.
- Strong written and verbal communication skills and the ability to work effectively with engineering, security, leadership, and external vendors.
- Ability to operate independently in a fast-moving environment with changing priorities and organizational ambiguity.
- Experience with AWS, including IAM, Route 53, GuardDuty, and S3, is a plus.
- Kubernetes exposure and experience with hybrid cloud environments are advantageous.
- Familiarity with SOC 2, Okta, Vanta, or security-focused infrastructure practices is preferred.
- Experience in adtech, real-time bidding, or other high-QPS, latency-sensitive environments is a plus.
- Netris, SDN controller, and MAAS provisioning experience is desirable.
- Willingness and ability to travel periodically to domestic and international data center locations and headquarters.
- Must be authorized to work in the United States; the role may involve background screening and does not indicate sponsorship availability.
- Fully remote position for candidates based in the United States.
- Opportunity to own a complex, globally distributed bare-metal infrastructure environment supporting high-volume, latency-sensitive workloads.
- Direct ownership of hardware strategy, GPU expansion, capacity planning, and cloud-versus-colocation decisions.
- Leadership responsibility for an established and experienced global infrastructure team.
- Direct visibility with senior engineering and infrastructure leadership and strong alignment between infrastructure goals and business objectives.
- Exposure to cutting-edge infrastructure challenges across high-performance networking, distributed systems, data centers, and on-premise machine learning infrastructure.
- Periodic travel to data center sites in the United States and internationally, as well as headquarters for operational reviews, team collaboration, and site work.
- Opportunity to contribute to a rapidly scaling organization and play a significant role in shaping infrastructure strategy and operational excellence.