Enterprise Tools & Observability Engineer (L2)

Open 23d

Function: Infrastructure Engineering — Tools & Observability
Experience Level: 2–5 Years

Role Overview

We are seeking a skilled Enterprise Tools & Observability Engineer (L2) to manage and optimize enterprise monitoring, security, and IT service management platforms. This role focuses on hands-on administration, configuration, and integration of observability and tooling platforms to ensure high availability, performance, and security of IT services.

The engineer will also contribute to automation, CI/CD pipelines, and continuous improvement of monitoring and alerting systems to enhance operational efficiency.

Key Responsibilities

Monitoring & Observability Tools

  • Administer and manage Splunk for log ingestion, index management, SPL queries, alerting, and dashboards.
  • Configure and maintain SolarWinds DPM for infrastructure monitoring, node onboarding, thresholds, and performance reporting.
  • Manage Cisco ThousandEyes for synthetic monitoring, network visibility, and performance insights.
  • Implement and maintain log aggregation pipelines across on-premise and cloud environments.

Security & Compliance Tools

  • Administer Imperva DAM for database activity monitoring, policy configuration, and compliance reporting.
  • Manage Trend Micro Deep Security including policy configuration, agent deployment, and security monitoring.
  • Oversee BigFix patch management including baseline setup, patch deployment, and compliance tracking.

ITSM & Platform Integration

  • Support and administer ITSM platforms (BMC Helix or equivalent) including incident, change, problem, and CMDB modules.
  • Integrate monitoring and security tools with ITSM systems for automated ticketing and event correlation.

Automation & DevOps

  • Develop and maintain automation using Ansible for configuration management and operational tasks.
  • Support Infrastructure as Code (Terraform) for provisioning and managing tool infrastructure.
  • Maintain and optimize CI/CD pipelines using GitHub Actions and Jenkins.

Alerting & Performance Optimization

  • Configure alert thresholds and perform alert tuning to reduce noise and improve signal quality.
  • Analyze system metrics and logs to proactively identify issues and optimize performance.
  • Maintain operational dashboards and generate SLA and performance reports.

Operations & Support

  • Perform system health checks, troubleshooting, and L2 support for enterprise tools.
  • Participate in incident resolution and root cause analysis (RCA).
  • Collaborate with cross-functional teams to improve observability and platform reliability.

Required Skills & Competencies

  • 2–5 years of experience in monitoring, observability, or enterprise tools administration
  • Hands-on experience with Splunk (SPL, dashboards, alerting)
  • Experience with SolarWinds, ThousandEyes, or similar monitoring tools
  • Familiarity with security tools (Imperva DAM, Trend Micro, BigFix)
  • Knowledge of ITSM platforms (BMC Helix or similar)
  • Experience with Ansible and Terraform (basic to intermediate)
  • Exposure to CI/CD tools (GitHub Actions, Jenkins)
  • Understanding of log management, alerting, and observability concepts
  • Strong troubleshooting and analytical skills

Preferred Skills

  • Experience with multi-tool integration and event correlation
  • Exposure to cloud environments (AWS, Azure)
  • Basic scripting (Python, Shell)
  • Understanding of ITIL processes