freehire launches on Product Hunt on 26 August.

Follow →

Software Development Snr Manager

Open 46d reposted 2× · 2 open copies

This role reports into senior leadership responsible for global AI/GPU host management strategy, fleet readiness, AIOps adoption, intelligent automation, service reliability, and operational tooling. You will translate that strategy into disciplined execution across day-to-day operations, service queues, triage, incident response, patching, upgrades, runbook maturity, telemetry improvements, and operational readiness for OCI's AI/GPU host fleet.

You will be a hands-on leader during operational escalations, driving crisp communications, rapid diagnosis, and durable corrective actions. You will partner cross-functionally with geographically distributed OCI engineering, platform, networking, data center, compliance, and operations teams to deliver measurable reliability, efficiency, and execution outcomes.

Key Responsibilities

Organizational Leadership & Talent Development

- Lead, grow, and develop Morocco-based teams of operators, developers, and technical leads in a 24x7 operational environment; establish clear ownership boundaries, on-call expectations, and accountability mechanisms.

- Recruit, hire, coach, and retain high-performing talent; set goals, manage performance, develop successors, and raise the technical and operational bar across the team.

- Create a culture of operational excellence, ownership, continuous learning, pragmatic simplification, and disciplined execution while supporting rapid AI/GPU infrastructure growth.

Service Ownership: AI/GPU Host Operations at Cloud Scale

- Own operational outcomes for assigned AI/GPU host management services, including host health, fleet readiness, hardware/software triage, service queues, reliability, and operational readiness.

- Guide teams that diagnose AI compute host issues across hardware, Linux, services, networking, automation, and monitoring layers; ensure issues are resolved quickly and with durable prevention mechanisms.

- Drive operational execution for service patching, upgrades, staged rollouts, change controls, and readiness reviews using metrics-driven planning and governance.

- Partner with global OCI teams to ensure Morocco operations align with broader host management roadmaps, standards, escalation practices, and fleet reliability goals.

Operational Excellence, Metrics, and Governance

- Define and mature operational KPIs and reporting for service health, incident performance, ticket aging, ticket resolution quality, queue backlog, change execution, and operational readiness.

- Lead high-severity incidents and escalations: coordinate rapid triage, communicate status clearly, drive high-quality post-incident reviews, and follow through on corrective actions.

- Improve operating mechanisms for risk management, audit/compliance alignment, change management, handoffs, escalation hygiene, and cross-team execution tracking.

- Use data and trend analysis to identify repeat issues, operational bottlenecks, staffing gaps, process defects, and opportunities for reliability improvement.

Engineering Enablement, AIOps, and Automation

- Drive adoption of AIOps and intelligent automation to reduce manual toil, improve alert quality, accelerate event correlation, standardize remediation, and improve triage accuracy.

- Partner with engineering and platform teams to prioritize operational tooling, telemetry improvements, workflow enablement, runbook automation, and self-healing opportunities.

- Establish measurable automation outcomes such as reduced manual handling, improved MTTR, increased auto-triage coverage, fewer repeat issues, and better operator/engineer effectiveness.

- Manage focused software engineering work aligned to operations outcomes, including scripts, dashboards, workflow tooling, triage aids, monitoring enhancements, and reliability automation.

Cross-Functional & Stakeholder Engagement

- Work closely with senior leaders, peer managers, technical leads, and globally distributed teams to deliver predictable execution across AI/GPU host operations.

- Translate complex technical and operational situations into accurate narratives, decisions, risks, and action plans for senior stakeholders.

- Represent the team in operational reviews, readiness discussions, incident forums, and cross-functional planning sessions with clarity and ownership.

Qualifications / Experience

- BS or MS in Computer Science, Engineering, or a related technical field, or equivalent practical experience.

- 8+ years of experience in software engineering, infrastructure operations, cloud operations, site reliability, production operations, or related technical areas.

- 5+ years of people management and/or technical leadership experience, including experience leading operators, engineers, senior technical contributors, or multiple operational workstreams.

- Experience building and scaling teams, including recruiting, hiring, coaching, performance management, goal setting, and leadership development.

- Strong operational background with incident management, service ownership, queue management, operational readiness, process improvement, and post-incident corrective actions.

- Experience operating or supporting Linux-based infrastructure at scale, including hardware/software troubleshooting and service lifecycle execution.

- Working familiarity with scripting and automation ecosystems such as Python, Bash, or similar tools sufficient to sponsor, review, and guide operational tooling direction.

- Understanding of distributed systems fundamentals and the ability to reason across hardware, operating systems, networking, services, monitoring, automation, and customer impact.

- Familiarity with networking protocols such as TCP/IP and HTTP and with standard cloud infrastructure architectures.

- Strong organizational and planning skills, including prioritization, scheduling, execution tracking, and operational governance.

- Strong written and verbal communication skills, including the ability to communicate technical risks, operational status, tradeoffs, and execution plans to senior stakeholders.

Preferred / Nice to Have

- Experience operating large-scale cloud infrastructure, AI/ML infrastructure, HPC environments, or GPU fleets.

- Experience leading teams in a 24x7 production operations environment with globally distributed stakeholders.

Career Level - M3

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available