Application /Production Support
Production Support & Incident Management
· Act as the primary Ll /L2 support contact for digital platforms, e-commerce systems, websites, and customer-facing services.
· Monitor incident queues, service requests, alerts, and support tickets, ensuring adherence to SLAs and operational procedures.
· Lead incident triage, troubleshooting, escalation, and resolution activities.
· Perform impact assessment and coordinate with relevant stakeholders during service disruptions.
· Support incident management activities and facilitate communication during critical outages.
· Conduct post-incident reviews and root cause analysis (RCA) to prevent recurrence.
· Develop and maintain operational runbooks, support procedures, and knowledge base articles.
System Monitoring & Reliability
· Monitor application, infrastructure, and business service health using observability and monitoring tools.
· Analyse system performance, availability, error trends, and capacity utilization.
· Configure and tune alerts to reduce noise and improve operational visibility.
· Collaborate with engineering teams to improve system reliability and operational resilience.
· Support Site Reliability Engineering (SRE)practices, including reliability metrics, incident reduction, and service availability improvements.
Cloud & Infrastructure Support
· Provide operational support for cloud-hosted applications and infrastructure, primarily on AWS.
· Perform first-level troubleshooting on:
o Compute services (EC2, ECS, Lambda)
o Networking
o Load Balancers
o CDN services
o Storage services
· Investigate infrastructure-related issues affecting application performance or availability.
· Support deployment verification and post-release monitoring activities.
Application & Integration Support
· Troubleshoot application issues across web, mobile, APIs, and middleware platforms.
· Analyse application logs, monitoring data, and system traces to identify root causes.
· Support integrations with external systems, partners, payment gateways, and other enterprise platforms.
· Work closely with L3 to reproduce issues and validate fixes.
· Support release and deployment activities, including late-night and weekend implementations when required.
Continuous Improvement
· Identify recurring incidents and operational inefficiencies.
· Drive automation opportunities to reduce manual effort and repetitive support activities.
· Recommend improvements to monitoring, alerting, deployment processes, and support workflows.
· Contribute to operational excellence initiatives and service reliability improvements.
Stakeholder &Vendor Management
· Collaborate with internal teams, external vendors, and partners across different geographies and time zones.
· Communicate effectively with technical and non-technical stakeholders.
· Provide timely updates during incidents and service disruptions.
· Participate in operational reviews, governance meetings, and service improvement discussions.
Required Skills & Experience Technical Skills
· 5+ years of experience in Application Support, Production Support, Technical Operations, TechOps, or related roles.
· Strong knowledge of AWS cloud services and operational support.
· Good understanding of cloud-native application architecture.
· Experience supporting:
o Digital platforms
o E-commerce systems
o Customer-facing web applications
o Mobile applications
· Knowledge of CDN technologies and content delivery architecture.
· Experience with monitoring and observability platforms such as:
o Datadog
o New Relic
o Dynatrace
o AppDynamics
o Grafana
o CloudWatch
· Familiarity with APM (Application Performance Monitoring) concepts.
· Understanding of SRE principles and operational best practices.
· Strong knowledge of:
o APIs
o Microservices
o Web services
o System integrations
o Authentication and authorization flows
· Experience using ticketing and ITSM platforms(Jira Service Management, ServiceNow, etc.).
· Understanding of Incident, Problem, and Change Management processes.
· Familiarity with log analysis tools such as Elasticsearch, Kibana, Splunk, or Cloud WatchLogs.
Soft Skills
· Strong troubleshooting and analytical thinking abilities.
· Naturally curious and investigative, with adesire to understand the full context behind incidents and operational events.
· Excellent problem-solving and root cause analysis skills.
· Strong ownership mindset and accountability.
· Ability to work independently in a fast-paced operational environment.
· Good communication and stakeholder management skills.
· Ability to remain calm and methodical during high-severity incidents.
· Strong documentation and knowledge-sharing practices.
· Continuous improvement mindset with a focus on automation and operational efficiency.
Good to Have Skills
· Knowledge of WeChat Mini Program ecosystem and integrations.
· Experience supporting SAP Commerce, Adobe Experience Manager (AEM), Magento, Shopify, or similar e-commerce platforms.
· Basic scripting skills (Python, Shell, Bash, PowerShell).
· Experience with API Gateway and event-driven architectures.
· AWS Certifications (Cloud Practitioner, Solutions Architect Associate, SysOps Administrator).
· ITIL Foundation certification.
· Experience supporting payment gateways and digital commerce ecosystems.
Working Conditions
· Participate in a 24x7 support and on-call rotation model.
· Support late-night, weekend, and public holiday deployments where required.
· Work closely with internal teams, vendors, and stakeholders across multiple time zones.
· Respond to critical production incidents outside business hours when necessary.