Operations Support System Engineer
Your Role
Megaport is looking for an experienced, hands-on engineer in OSS and operations automation. This is a new role within Megaport’s operations department. The purpose of the role is to continue managing our monitoring and customer support systems, implement synergies with existing tools, and develop automation for day-to-day work. Your customers are the front-line support team, CS / TSE / NOC, Network Operations Engineering and Production Development Engineering teams.
To be successful in this role, you will need to be experienced in relevant technologies, have excellent interpersonal skills, be able to project-manage dependencies across technical departments and lead the implementation of operations tools/automation strategy.
We are looking for someone excited to lead the development of this capability. There is an abundance of high-value opportunities in this space for Megaport. Proven experience in designing, implementing and maintaining OSS tools is a must. The ideal candidate would have a genuine interest in analysing measurable operations work activity, exploring opportunities to develop systemic automations and managing the design and implementation of the identified opportunities.
What You’ll Be Doing
Business Analysis, Solution Architecture & Stakeholder Engagement
- Requirements Elicitation & Scoping: Partner with operations managers, business units, and key stakeholders to capture, analyse, and refine operational and network requirements into detailed Business Requirements Documents (BRDs), Epics, User Stories, and Acceptance Criteria (BDD/Gherkin).
- Architecture & Executive Sign-off: Collaborate with internal technology teams to architect scalable end-to-end solutions; present technical proposals and delivery roadmaps to Business Unit leadership for sign-off.
- Technical-Business Bridge: Translate complex network operational data and technical findings into actionable business insights, workflow models, and operational improvements.
End-to-End Delivery & Platform Management
- Lifecycle Ownership: Own solution delivery from concept to production—including solution design, project management, implementation, QA testing, and release management.
- Cloud Infrastructure & System Integration: Design, deploy, and maintain resilient application capabilities across Linux and AWS, leveraging REST APIs, Message Bus architectures, OAuth, and webhooks.
- Process Improvement & QA Strategy: Identify operational bottlenecks, design UAT procedures and test strategies, automate manual tasks, and ensure platform scalability, reliability, and security.
Openware Observability, Monitoring & Network Assurance
- Open-Source Monitoring Management: Deploy, maintain, and optimise openware observability stacks (Prometheus, Grafana, OpenNMS, Nagios, and alerta.io / Alertmanager) to modernise and replace legacy commercial monitoring suites.
- Proactive Assurance & Telemetry: Configure streaming telemetry, SNMP, MIBs, and automated alert thresholds across network elements and cloud applications to ensure rapid failure detection.
- Incident Escalation & RCA: Provide Level 3 troubleshooting and root-cause analysis (RCA) to resolve complex operational, network, and system incidents.
Salesforce Ecosystem & Business Intelligence
- Custom Development & CI/CD: Build and support solutions within Salesforce Service Cloud utilising Apex, LWC, SOQL/SOSL, Agentforce 360, and Data Cloud, backed by GitHub-based CI/CD pipelines.
- Platform Administration: Administer security models (Roles, Profiles, Permission Sets, FLS, MFA/SSO), declarative automations (Salesforce Flow), custom objects, and data governance.
- BI Dashboards & Reporting: Design, build, and maintain executive reports and interactive dashboards in Power BI and Salesforce to drive data-informed decision-making.
Atlassian Ecosystem & Jira Space Administration
- Jira Space & Project Administration: Configure and manage Jira spaces, issue types, custom workflows, field configurations, screen schemes, and permission models across operational teams.
- Workflow Automation: Build and maintain native Jira Automation rules, custom JQL filters, and webhooks to streamline cross-team routing, SLA tracking, and operational handoffs.
Governance, Lifecycle & Operational Support
- Knowledge Base Authoring: Create and maintain technical specifications, architecture diagrams, API docs, and operational playbooks in Confluence.
- Vendor Management & Cost Optimisation: Partner with third-party vendors and partners to manage product lifecycles, integrate external systems, and optimise system costs.
- On-Call Support: Participate in scheduled on-call rotations and maintain flexible working hours to support critical incidents and maintain system uptime.
What We Are Looking For
Professional & Industry Experience with Certifications [Required or Preferred]
- 8–10+ Years of Experience: Proven track record in IT, Telecommunications, or Cloud Systems Engineering within hybrid Business Systems Analysis, OSS/BSS, Salesforce, and Atlassian administration roles.
- Agile & Remote Leadership: Demonstrated ability to work independently within distributed teams, participate actively in Agile frameworks, and manage multi-stakeholder engagements.
- Certifications: Salesforce or Atlassian, Business Analysis & Agile, Cloud & Analytics and Telecom & Network Frameworks
Business Analysis & Business Intelligence
- Requirements & Process Mapping: Expertise modelling "As-Is" and "To-Be" operational workflows, performing gap analysis, and building specs using Lucidchart, Visio, Jira, and Confluence.
- Data Visualisation & Analytics: Hands-on experience constructing dashboards, JQL filters, gadget views, and KPI metrics in Power BI, Salesforce Reports, and Jira.
- Observability Tools: Hands-on experience with open-source monitoring and alerting tools including Prometheus, Grafana, Alerta (alerta.io), Alertmanager, OpenNMS, and Nagios.
- Networking Protocols: Solid understanding of Network Management Systems (NMS), SNMP, MIBs, streaming telemetry, MPLS and AWS networking platforms.
Salesforce Ecosystem Engineering
- Admin & Declarative Tools: Deep knowledge of Salesforce Service Cloud, Flow Builder, Data Loader, AppExchange, Omni-Channel, and security controls.
- Programmatic Development: Practical experience with Apex, Triggers, Lightning Web Components (LWC), SOQL, SOSL, Data Cloud, and Agentforce 360.
- DevOps & Release Management: Hands-on release management using DevOps Centre, Change Sets, and GitHub-based CI/CD pipelines.
Atlassian Ecosystem Administration
- Jira Admin & Automation: Advanced skills in native Jira Space administration, custom workflow/scheme design, Jira Automation rules, webhooks, and advanced JQL.
Software Engineering, DevOps & Data
- Programming & Scripting: Strong coding skills in Python and Shell scripting; proficiency in Groovy, JavaScript, Node.js, Vue, and working knowledge of Perl.
- Databases: Hands-on experience querying and managing PostgreSQL, MongoDB, and MySQL.
- Cloud & IaC: Functional knowledge of AWS cloud services, Linux sysadmin, and Infrastructure as Code (IaC) tools.
Working Conditions, Location and Hours
- Full-time office-based role, at our Gurugram office
- The working day is 8 hours, and the working week is 40 hours.
- 24/7 on-call availability for emergencies
- You will get 2 days weekly off (on-call availability for emergencies)
- 90-day notice period for resignation after the probation period
- Working exclusively for Extreme Infocom Pvt. Ltd. and not for any other companies
What We Offer
Subject to the internal policies of Megaport, which may be updated from time to time:
- 18 days per year's privilege/earned leave
- Family health insurance according to company policy
- A motivated team combining industry experts and emerging talent.
- Recognition programs – including Legend and Kudos Awards.
- Health & wellness programs and mental well-being support.