Lead Infrastructure Engineer - Infrastructure (HPE NonStop/Tandem) with AI/Automation
Summary
Lead a team maintaining HPE NonStop payment infrastructure, integrating AI/automation for monitoring, troubleshooting, and secure key management to ensure high availability and compliance.
Assume a vital position as a key member of a high-performing team that delivers infrastructure and performance excellence. Your role will be instrumental in shaping the future at one of the world's largest and most influential companies.
Job responsibilities
- Configure, maintain, and troubleshoot HPE NonStop hardware and architecture.
- Manage and configure Enterprise Secure Key Managers to ensure robust security for sensitive data.
- Set up and maintain Etinet servers, ensuring optimal performance and integration with HPE NonStop systems.
- Understand and support HPE NonStop subsystems (Mediacom, TMF, KMSF), and their relationship to hardware for effective troubleshooting and planning of upgrades or installations.
- Utilize and manage SCF, ZZSTO, ZZZCIP, and ZZKRN utilities for system configuration, monitoring, and maintenance.
- Uses enterprise-authorized AI capabilities within the work environment to accelerate infrastructure analysis and design documentation, validating outputs and handling operational data according to sensitivity and security requirements.
- Applies reuse-first, AI-assisted practices within delivery and automation routines to identify recurring issues and validate remediation options, ensuring changes are traceable/auditable and aligned to resiliency and security expectations.
- Leverage AI-assisted operations (AIOps) techniques to improve incident triage, reduce MTTR, and proactively detect infrastructure risks (e.g., anomaly detection on system/EMS logs, event correlation, early-warning indicators).
- Build and curate high-quality operational knowledge (KB articles, runbooks, known-error records) that can be used by AI assistants to provide accurate, auditable troubleshooting guidance.
- Partner with SRE/Observability and Cyber teams to evaluate, implement, and govern AI-enabled monitoring and alerting, ensuring model outputs are explainable, traceable, and compliant with security controls.
- Document configurations, procedures, and troubleshooting steps for knowledge sharing and compliance, collaborate with cross-functional teams to plan and execute system upgrades and installations.
Required qualifications, capabilities, and skills
- Formal training or certification on software engineering concepts and 5+ years applied experience
- In-depth knowledge of HPE NonStop hardware and architecture, including system configuration and maintenance.
- Experience with Enterprise Secure Key Managers: ability to configure and manage secure key solutions in an enterprise environment.
- Proficiency in configuring Etinet servers and integrating them with HPE NonStop systems.
- Strong understanding of HPE NonStop subsystems (Mediacom, TMF, KMSF) and their interaction with hardware, especially for troubleshooting and upgrade/install planning.
- Hands-on experience with SCF, ZZSTO, ZZZCIP, and ZZKRN for system configuration and management.
- Demonstrated experience using enterprise-authorized AI capabilities within the work environment to support infrastructure engineering workflows with strong validation habits and awareness of data sensitivity.
- Ability to review and validate AI-assisted recommendations before implementation, escalating when uncertain and ensuring outcomes align to resiliency, security, and auditability expectations.
- Experience applying AIOps / ML-driven monitoring concepts (anomaly detection, alert correlation, noise reduction, trend forecasting, predictive capacity/health signals) in a production infrastructure environment.
- Ability to work with telemetry pipelines (logs/metrics/traces), define data quality expectations, and operationalize signals for automation and reliability outcomes.
- Practical experience using AI assistants for troubleshooting and documentation, with an emphasis on validation, secure handling of sensitive data, and producing audit-ready outputs.
- Experience implementing or operating AIOps platforms and integrating them with incident/ticketing workflows.
- Exposure to LLM governance concepts in enterprise settings (data classification, access controls, audit logging, model risk considerations).
- Experience building operational analytics (e.g., Python/SQL) to mine EMS/application logs for recurring patterns, failure modes, and leading indicators.
- Familiarity with reliability practices (SLOs/SLIs, error budgets, blameless postmortems) and using AI to improve these processes.