AIOps Support Lead
At BCE Global Tech we are on a mission to modernize global
connectivity, one connection at a time. We aim to build the highway to the
future of communications, media and entertainment, determined to emerge as a
powerhouse within the technology landscape in India team in Bengaluru.
We bring ambitions to life through design thinking that
bridges the gaps between people, devices and beyond, fostering unprecedented
customer satisfaction through technology.
Our core values support a customer-centric approach and the
harnessing of cutting-edge technology to provide business outcomes with
positive societal impact. Guided by innovation and a commitment to progress,
we’re shaping a brighter future for the generations of today and tomorrow.
If you would like to be a part of a team of thought-leaders
pioneering advancements in 5G, MEC, IoT and cloud-native architecture, we’d
love to hear from you.
Requirements
What You'll Do
• Manage the Tier 1 team: directly manage a
team of 14 AIOps Support Engineers performing manual triage of alarms and
alerts across a diverse, 50-to-320-application portfolio including hiring,
coaching, scheduling, and performance management.
• Own triage quality and speed: set and monitor
standards for how quickly and accurately the team detects, classifies, and
routes incidents, and drive continuous improvement in mean-time-to-triage.
• Drive data stewardship: partner with
application teams to standardize alarm and alert data across heterogeneous log
aggregation tools (Dynatrace, New Relic, ManageEngine, Glass box, and others)
into a clean, consistent telemetry backbone built on Open Telemetry.
• Manage the reactive-to-proactive shift: reduce reliance on
reactive, manual triage over time by improving alert quality, correlation, and
early-warning signals laying the groundwork for future automated and Agentic
triage.
• Navigate a diverse, moving application landscape: support applications
spanning different technology stacks and different architecture dispositions
(Invest, Tolerate, Retire, Migrate), reprioritizing team focus as the portfolio
shifts.
• Coordinate onboarding of new apps: run a repeatable
process for bringing new applications into Tier 1 coverage as the program
scales from 50 to 320 applications, including support group and application
owner mapping.
• Manage stakeholders: act as the primary point of contact
for support groups, application owners, and AIOps program leadership on Tier 1
status, incidents, and data-quality issues.
• Manage shift/roster coverage: ensure the team of 14
provides consistent triage coverage across required hours as the application
count grows.
• Report on outcomes: track and report team KPIs triage
time, alert-to-incident accuracy, false-positive rates, coverage growth to
program leadership.
What We're Looking For
• Experience: 7+ years in application/production support
(L1/L1.5/L2) or site reliability, with 2+ years directly managing a technical
support team.
• Hybrid environment expertise: proven experience
supporting applications across both on-premises and cloud environments, with
exposure to modern microservices architectures.
• Observability tooling: hands-on experience with monitoring
and observability platforms such as Dynatrace, New Relic, AWS CloudWatch,
ManageEngine, or Glass box; working knowledge of Open Telemetry and distributed
tracing concepts.
• ITIL discipline: strong grounding in incident, problem, and
change management practices, with ServiceNow or Jira ticket management
experience.
• Technical range: comfortable with Linux and Windows
troubleshooting, basic networking (TCP/IP, DNS, HTTP/HTTPS, SSL, load
balancers), SQL/database query analysis, and API/integration troubleshooting.
• People management: demonstrated ability to hire, coach,
and retain a team of 10+ technical support staff through a period of
significant scale-up (5x application coverage growth).
• Analytical mindset: able to turn noisy, inconsistent alert
data into clear, actionable insight, and to build repeatable frameworks rather
than one-off fixes.
• Comfort with ambiguity: willing to support a
moving target a portfolio spanning
Invest, Tolerate, Retire, and Migrate applications and to adapt priorities as the program
evolves.
• Incident
Management Lifecycle, and working knowledge of Problem, Change Request, and
Service Request concepts (ITIL)
• CMDB
concepts and their use in incident and asset traceability
• Hands-on
experience with log aggregation technologies (Dynatrace, New Relic,
ManageEngine, Glass box, or similar)
• Working
knowledge of JSON and XML, and basic file/task automation
• Understanding
of IT infrastructure and basic networking: VMs, firewalls, load balancers,
containers, OpenShift (OCP), Kubernetes
• Unix
Shell scripting; Windows batch file creation
• Basic
cloud concepts: compute, storage, and security fundamentals
• Security
fundamentals: TLS, SSL, tokens, and secret management
• Familiarity
with API gateways and API testing toolkits (Postman, SOAP UI, or similar)
• Outage
response management and experience leading cross-functional coordination during
major incidents
• Ability
to drive Root Cause Analyses (RCAs) and build reusable knowledge
articles/runbooks
• SLA/SLO
management and reporting, including availability calculation
• Working
knowledge of data concepts: data latency, data fragmentation, data lineage, and
data marts
• Familiarity
with AI concepts such as prompt engineering, knowledge graphs, and
Retrieval-Augmented Generation (RAG)
• Experience
with cloud-native observability on AWS, Azure, or GCP
• Exposure
to Agentic AI or automation-driven triage tooling
Success Looks Like
• A
14-person Tier 1 team that reliably triages alerts across 50+ applications with
clear, standardized data.
• A
measurable, ongoing reduction in reactive manual triage as proactive detection
improves.
• A
clean, well-governed Open Telemetry-based data backbone that the future Agentic
automation layer can build on.
A
repeatable onboarding process ready to scale coverage from 50 to 320 applications.
Benefits
What We Offer
Competitive salaries and comprehensive health benefits
Flexible work hours and remote work options.
Professional development and training opportunities.
A supportive and inclusive work environment
Skills
- Agentic AI
- AI
- API
- Automation
- AWS
- Azure
- Backbone.js
- Bash
- Cloud
- Cloud Native
- CloudWatch
- CMDB
- Data Lineage
- Data Quality
- Design Thinking
- DNS
- Dynatrace
- Firewall
- GCP
- ITIL
- Jira
- JSON
- Kubernetes
- Linux
- Microservices
- Networking
- New Relic
- Observability
- OpenShift
- OpenTelemetry
- Performance Management
- Postman
- Prompt Engineering
- RAG
- ServiceNow
- SQL
- SSL
- TCP/IP
- TLS
- Unix
- XML