AI DevOps Engineer
Summary
Provides L2/L3 production support for enterprise Java/J2EE applications on Linux/Unix, Oracle Database, and WebLogic, including 24x7 on-call, incident troubleshooting, RCA, and CI/CD-backed deployments. The core focus is building Agentic AI automation—self-healing systems and automated runbooks—integrated with monitoring, alerting, and ITSM platforms.
Job Summary
We are looking for an experienced AI DevOps Engineer to provide L2/L3 application support and AI-driven automation across enterprise applications running on Linux/Unix, Java/J2EE, Oracle Database, and WebLogic Server.
The ideal candidate will have proven hands-on experience in Agentic AI and intelligent automation, with the ability to design and implement AI-driven solutions that improve operational efficiency, automate repetitive support activities, enable proactive incident management, and reduce manual intervention.
The role requires a strong combination of Application Support, DevOps, enterprise technology, and Agentic AI automation skills.
Key Responsibilities Application Support & Operations
- Provide L2/L3 production support for critical enterprise applications.
- Participate in 24x7 production support as required.
- Monitor application health, availability, and performance.
- Troubleshoot and resolve complex production incidents.
- Ensure high availability, operational stability, and adherence to defined SLAs.
- Perform detailed Root Cause Analysis (RCA) for recurring and critical incidents.
- Implement permanent fixes and remediation to prevent recurrence.
- Support application deployments, environment management, and production releases.
- Collaborate with Development, Infrastructure, Database, and other technology teams to resolve technical issues.
- Proactively identify potential incidents and performance issues.
DevOps Responsibilities
- Support application deployment and release activities across environments.
- Integrate application support processes with CI/CD pipelines.
- Assist with environment configuration and management.
- Automate repetitive operational and deployment activities.
- Improve application reliability and operational efficiency through automation.
- Support monitoring, alerting, and incident-management processes.
Agentic AI & Intelligent Automation
- Design, develop, and implement Agentic AI solutions for IT Operations.
- Build AI-powered automation to reduce manual operational effort.
- Develop self-healing systems capable of detecting issues and initiating automated remediation.
- Create intelligent and automated runbooks for common operational incidents.
- Integrate AI agents with:
- Monitoring tools
- Alerting platforms
- IT Service Management (ITSM) tools
- Application support processes
- Use AI agents to support proactive incident detection, diagnosis, and resolution.
- Identify opportunities where Agentic AI can improve operational efficiency.
- Deliver measurable improvements in areas such as:
- Incident reduction
- Mean Time to Resolution (MTTR)
- Manual effort reduction
- Operational overhead
- Application availability
Required Technical Skills Application Support
Hands-on experience supporting enterprise applications running on:
- Linux / Unix
- Java / J2EE
- Oracle Database
- WebLogic Server
DevOps
- Experience with application deployment and environment management.
- Understanding of CI/CD integration and release processes.
- Experience with production monitoring and incident management.
- Strong troubleshooting and problem-solving skills.
- Experience working with development, infrastructure, and database teams.
Agentic AI / AI Automation
- Proven hands-on experience with Agentic AI.
- Experience designing and implementing AI agents for IT Operations.
- Experience building AI-driven automation and intelligent workflows.
- Knowledge of AI-powered incident management and remediation.
- Experience developing automated runbooks or self-healing solutions.
- Ability to integrate AI agents with monitoring, alerting, and ITSM platforms.
Key Deliverables
- Stable and reliable production application support.
- Timely resolution of L2/L3 incidents.
- Effective RCA and permanent remediation.
- Improved application availability and SLA compliance.
- Automated operational processes and runbooks.
- AI-driven self-healing and proactive incident management.
- Measurable reduction in manual support effort and operational overhead.
- Continuous improvement of application support and DevOps processes.