Software Development Snr Manager
This role reports into senior leadership responsible for global AI/GPU host management strategy, fleet readiness, AIOps adoption, intelligent automation, service reliability, and operational tooling. You will translate that strategy into disciplined execution across day-to-day operations, service queues, triage, incident response, patching, upgrades, runbook maturity, telemetry improvements, and operational readiness for OCI's AI/GPU host fleet.
You will be a hands-on leader during operational escalations, driving crisp communications, rapid diagnosis, and durable corrective actions. You will partner cross-functionally with geographically distributed OCI engineering, platform, networking, data center, compliance, and operations teams to deliver measurable reliability, efficiency, and execution outcomes.
Key Responsibilities
Organizational Leadership & Talent Development
- Lead, grow, and develop Morocco-based teams of operators, developers, and technical leads in a 24x7 operational environment; establish clear ownership boundaries, on-call expectations, and accountability mechanisms.
- Recruit, hire, coach, and retain high-performing talent; set goals, manage performance, develop successors, and raise the technical and operational bar across the team.
- Create a culture of operational excellence, ownership, continuous learning, pragmatic simplification, and disciplined execution while supporting rapid AI/GPU infrastructure growth.
Service Ownership: AI/GPU Host Operations at Cloud Scale
- Own operational outcomes for assigned AI/GPU host management services, including host health, fleet readiness, hardware/software triage, service queues, reliability, and operational readiness.
- Guide teams that diagnose AI compute host issues across hardware, Linux, services, networking, automation, and monitoring layers; ensure issues are resolved quickly and with durable prevention mechanisms.
- Drive operational execution for service patching, upgrades, staged rollouts, change controls, and readiness reviews using metrics-driven planning and governance.
- Partner with global OCI teams to ensure Morocco operations align with broader host management roadmaps, standards, escalation practices, and fleet reliability goals.
Operational Excellence, Metrics, and Governance
- Define and mature operational KPIs and reporting for service health, incident performance, ticket aging, ticket resolution quality, queue backlog, change execution, and operational readiness.
- Lead high-severity incidents and escalations: coordinate rapid triage, communicate status clearly, drive high-quality post-incident reviews, and follow through on corrective actions.
- Improve operating mechanisms for risk management, audit/compliance alignment, change management, handoffs, escalation hygiene, and cross-team execution tracking.
- Use data and trend analysis to identify repeat issues, operational bottlenecks, staffing gaps, process defects, and opportunities for reliability improvement.
Engineering Enablement, AIOps, and Automation
- Drive adoption of AIOps and intelligent automation to reduce manual toil, improve alert quality, accelerate event correlation, standardize remediation, and improve triage accuracy.
- Partner with engineering and platform teams to prioritize operational tooling, telemetry improvements, workflow enablement, runbook automation, and self-healing opportunities.
- Establish measurable automation outcomes such as reduced manual handling, improved MTTR, increased auto-triage coverage, fewer repeat issues, and better operator/engineer effectiveness.
- Manage focused software engineering work aligned to operations outcomes, including scripts, dashboards, workflow tooling, triage aids, monitoring enhancements, and reliability automation.
Cross-Functional & Stakeholder Engagement
- Work closely with senior leaders, peer managers, technical leads, and globally distributed teams to deliver predictable execution across AI/GPU host operations.
- Translate complex technical and operational situations into accurate narratives, decisions, risks, and action plans for senior stakeholders.
- Represent the team in operational reviews, readiness discussions, incident forums, and cross-functional planning sessions with clarity and ownership.
Qualifications / Experience
- BS or MS in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
- 8+ years of experience in software engineering, infrastructure operations, cloud operations, site reliability, production operations, or related technical areas.
- 5+ years of people management and/or technical leadership experience, including experience leading operators, engineers, senior technical contributors, or multiple operational workstreams.
- Experience building and scaling teams, including recruiting, hiring, coaching, performance management, goal setting, and leadership development.
- Strong operational background with incident management, service ownership, queue management, operational readiness, process improvement, and post-incident corrective actions.
- Experience operating or supporting Linux-based infrastructure at scale, including hardware/software troubleshooting and service lifecycle execution.
- Working familiarity with scripting and automation ecosystems such as Python, Bash, or similar tools sufficient to sponsor, review, and guide operational tooling direction.
- Understanding of distributed systems fundamentals and the ability to reason across hardware, operating systems, networking, services, monitoring, automation, and customer impact.
- Familiarity with networking protocols such as TCP/IP and HTTP and with standard cloud infrastructure architectures.
- Strong organizational and planning skills, including prioritization, scheduling, execution tracking, and operational governance.
- Strong written and verbal communication skills, including the ability to communicate technical risks, operational status, tradeoffs, and execution plans to senior stakeholders.
Preferred / Nice to Have
- Experience operating large-scale cloud infrastructure, AI/ML infrastructure, HPC environments, or GPU fleets.
- Experience leading teams in a 24x7 production operations environment with globally distributed stakeholders.
Career Level - M3