Data Engineer
Summary
Performs 24/7 first-line monitoring and incident response for Teradata and Cloudera environments, using SQL and Linux to validate data and troubleshoot batch workflows.
Summary
We are looking for an Data Engineer to join our 24x7 operations team. The role will be responsible for first-line monitoring, incident identification, ticket management, daily system health checks, and shift handovers. The ideal candidate should be comfortable working with monitoring tools, basic SQL, Linux environments, enterprise schedulers, and operational runbooks.
Responsibilities
- Perform 24x7 first-line monitoring and response based on defined runbooks and operational procedures.
- Monitor Teradata Viewpoint, Cloudera Manager, job monitoring tools, and batch processes for failures, alerts, and performance issues.
- Identify, log, categorize, prioritize, and elevate incidents according to defined SLA and operational procedures.
- Perform daily system health checks and ensure issues are appropriately documented and followed up.
- Monitor scheduled jobs and batch workflows and respond to failures according to established runbooks.
- Perform basic troubleshooting by reviewing application/system logs and identifying potential issues.
- Execute basic SQL queries for operational checks and data validation.
- Support basic operational activities across Teradata and Cloudera environments.
- Maintain accurate shift handover notes and communicate outstanding incidents, risks, and operational activities to the incoming team.
- Follow ITIL-based incident management and SLA requirements.
- Escalate complex or recurring issues to L2/L3 support teams as per the defined escalation matrix.
- Maintain operational documentation and ensure adherence to standard operating procedures.
Requirements
- 1–3 years of experience in IT Operations, Production Support, Application Support, NOC, or a similar environment.
- Basic knowledge of SQL and ability to perform simple queries.
- Monitoring-level familiarity with Teradata and Cloudera environments.
- Exposure to at least one enterprise job scheduler such as Control‑M, Autosys, Tidal, or similar.
- Basic knowledge of Linux/Unix commands and log‑file analysis.
- Understanding of ITIL concepts, incident management, and SLA‑based support.
- Good troubleshooting, analytical, and problem‑solving skills.
- Strong communication and documentation skills.
- Ability to follow runbooks, standard operating procedures, and escalation processes.
- Willingness and availability to work in a 24x7 rotational shift environment.