L2 Engineer
L2
Engineer
Location: Bangalore
Experience
Required: 3 – 5 Years
Employment
Type: Full-Time | Shift-Based (24x7 Rotational)
About Zybisys
Zybisys is a technology company
that helps banks, financial institutions, and FinTech businesses build and run
secure, reliable, and high-performance technology platforms. We work closely
with some of India's leading stock brokers to manage their cloud infrastructure,
cybersecurity, platform operations, and observability. With deep expertise in
the Capital Markets domain, we focus on simplifying complex technology,
improving operational resilience, and helping our customers innovate with
confidence.
Job Description
The L2 Engineer is the
second-tier technical support layer within Zybisys's 24x7 managed cloud
operations. This role owns the escalation queue from L1, performs deep-dive
troubleshooting and root cause analysis, implements approved infrastructure
changes, and manages the day-to-day configuration and administration of cloud
and hybrid on-premises environments.
This is a hands-on technical
role requiring strong Azure administration skills, solid networking and
security platform knowledge, and the ability to independently resolve complex
infrastructure incidents. The L2 Engineer is also a key contributor to operational
improvement — configuring monitoring rules, writing runbooks, and driving
automation of repeatable tasks.
Key Responsibilities
Incident Management &
Escalation Handling
· Own all L1-escalated incidents —
perform structured deep-dive troubleshooting, establish scope of impact, and
drive resolution within SLA thresholds.
· Lead root cause analysis (RCA)
for P1 and P2 incidents; produce clear, structured RCA reports with timeline,
contributing factors, and corrective actions within agreed timelines.
· Coordinate war-room responses
during major incidents — liaising across L1, specialist, DevOps, and vendor
support teams to restore service as quickly as possible.
· Maintain accurate and timely
ticket documentation throughout the incident lifecycle; ensure all escalations
have clear handover notes and status updates.
· Identify recurring incident
patterns and proactively recommend permanent fixes, monitoring improvements, or
runbook additions to reduce repeat escalations.
Infrastructure Administration
& Change Management
· Administer Azure resources
across multi-region environments — VM provisioning and decommissioning, disk
and snapshot management, VNet configuration, NSG rule management, and resource
tagging.
· Manage Windows Server
(2019/2022) and Linux (RHEL/Ubuntu) at an advanced level — OS hardening,
service configuration, performance tuning, log analysis, and patch management.
· Implement and validate approved
changes through the change management process; participate in change advisory
board (CAB) reviews and post-implementation testing.
· Manage hybrid connectivity —
ExpressRoute, Site-to-Site VPN, and Azure Arc-connected on-premises resources;
troubleshoot connectivity degradation and failover scenarios.
· Administer identity and
directory services: Azure AD, ADFS, hybrid identity, Conditional Access
policies, and group policy management for enterprise user and device
environments.
· Oversee storage operations:
NetApp ONTAP volume management, snapshot scheduling, SnapMirror replication
health, and Azure-native storage account administration.
Security Platform Operations
· Configure and manage Palo Alto
NGFW firewall rules, security policies, NAT policies, and URL filtering
profiles per approved change requests and client infosec policy directives.
· Administer Zscaler ZIA and ZPA —
manage forwarding rules, user policies, SSL inspection profiles, and connector
health for secure internet and private application access.
· Manage Trend Micro EDR policy
groups, agent deployment, exclusion lists, and alert response actions;
coordinate with the infosec team on policy tuning requirements.
· Operate CyberArk PAM — manage
safes, account onboarding, access provisioning, and session recording in line
with privileged access policies.
· Implement and document all
security configuration changes with full change record traceability; ensure no
security platform change is made without an approved change record.
Monitoring, Observability &
Automation
· Configure and maintain
Prometheus alerting rules, scrape targets, and recording rules; build and
manage Grafana dashboards for infrastructure, application, and network
observability.
· Develop and tune Azure Monitor
alert rules and Log Analytics (KQL) queries for hybrid environments; reduce
false positives and improve signal-to-noise ratio across all alert streams.
· Configure sFlow and NetFlow
collectors and visualizations for network traffic baseline analysis, anomaly
detection, and capacity reporting.
· Evaluate APM tool data
(Dynatrace, AppDynamics, Azure Application Insights, or equivalent) and work
with application teams to resolve performance degradation issues.
· Write and maintain operational
runbooks, SOPs, and post-incident playbooks; contribute automation scripts
(PowerShell, Python, Bash) to eliminate repetitive manual tasks.
Vendor & OEM Coordination
· Raise and manage support cases
with Microsoft, Palo Alto, Zscaler, NetApp, Trend Micro, CyberArk, and other
OEM vendors; drive cases to resolution with accurate reproduction data and
clear escalation paths.
· Track open vendor cases,
maintain status communication with internal stakeholders, and coordinate
vendor-recommended fixes through the change management process.
· Participate in vendor
maintenance windows, firmware upgrades, and product lifecycle activities for
all managed platforms.
Documentation & Knowledge
Contribution
· Produce high-quality RCA
reports, change records, and technical documentation that can be consumed by L1
teams and operations leads.
· Maintain and continuously
improve the operations knowledge base — document new solutions, update runbooks
after every major incident, and identify knowledge gaps.
· Mentor and upskill L1 engineers
through guided troubleshooting sessions, knowledge transfer, and structured
feedback on escalated tickets.
Required Skills & Experience
Cloud & Hybrid
Infrastructure
· 3 - 5 years of overall
IT/infrastructure experience; minimum 3 years in cloud infrastructure or
managed cloud operations roles.
· Strong hands-on Azure
administration: VMs, VNets, NSGs, Route Tables, Application Gateways, Load
Balancers, Azure Storage, and Azure AD across multi-region environments.
· Solid hybrid infrastructure
experience — integrating on-premises datacenter environments with Azure via
ExpressRoute, Site-to-Site VPN, and Azure Arc.
· Advanced Windows Server and
Linux (RHEL/Ubuntu) administration — performance tuning, log analysis, service
troubleshooting, and OS-level security hardening.
· NetApp ONTAP administration:
volume and aggregate management, snapshot and replication operations,
performance monitoring, and storage capacity planning.
· Enterprise identity management:
Azure AD, ADFS, hybrid identity, Conditional Access, Privileged Identity
Management, and Group Policy administration.
Networking
· Strong networking fundamentals:
BGP, OSPF, VLANs, VXLAN, subnetting, DNS, DHCP, NTP, IPSec VPN, and SD-WAN
concepts.
· Practical experience
troubleshooting multi-layer network issues — spanning Azure VNets, on-premises
routing, and WAN connectivity paths.
· Familiarity with hub-and-spoke
and Virtual WAN topologies; experience managing network peering, UDRs, and
traffic flow analysis.
Security Platform Operations
· Palo Alto NGFW — firewall policy
management, NAT configuration, security profile tuning, and Panorama-managed
rule deployments.
· Zscaler ZIA/ZPA — policy
administration, forwarding rules, user-based access configuration, and health
monitoring.
· Trend Micro Apex One / XDR —
policy management, agent health, exclusion management, and alert triage.
· CyberArk PAM — safe and account
administration, connector management, and access session governance.
· Ability to implement all
security configurations per client infosec directives with full change
documentation and audit traceability.
Monitoring, Observability &
Automation
· Prometheus and Grafana —
alerting rule configuration, dashboard creation, label management, and data
source integration.
· Azure Monitor and Log Analytics
— KQL query development, diagnostic settings, alert rule management, and
workbook creation.
· sFlow / NetFlow monitoring —
collector configuration, traffic baseline analysis, and anomaly identification.
· APM platform familiarity
(Dynatrace, AppDynamics, New Relic, or Azure Application Insights) —
interpreting traces, service maps, and performance baselines.
· Scripting for operational
automation: PowerShell, Python, or Bash — capable of writing scripts to
automate routine tasks and reduce MTTR.
ITSM & Service Delivery
· Strong ITSM process ownership:
incident, problem, change, and service request management using platforms such
as ServiceNow, Zoho Desk, or Freshservice.
· Experience producing structured
RCA reports and presenting findings to operations leads and client
stakeholders.
· Comfortable with 24x7
shift-based operations, on-call duties, and major incident coordination.
Certification (Preferred)
Domain | Certification |
Microsoft Azure | AZ-104 (Administrator) | AZ-305 (Solution Architect — advantageous) |
Networking | CompTIA Network+ | Cisco CCNA | Palo Alto PCNSA |
Security Platforms | Zscaler ZCCA | CyberArk Defender (advantageous) |
Monitoring | Grafana Certified
Associate | Azure Monitor Associate |
Service Management | ITIL v4 Foundation |