QA Automation Engineer - Network and Distributed Systems
Summary
QA Automation Engineer focused on network and distributed-systems testing, building chaos and failure-injection automation frameworks for Kubernetes, cloud networking, and service-to-service communication using Python and Go.
Role Overview
We are looking for a QA Automation Engineer with strong networking and distributed-systems expertise who can operate at the intersection of quality, reliability, and system-level testing.
This goes far beyond a traditional QA role. In this role, you’ll build automation that tests the limits of complex systems under real-world stress, network chaos, and massive scale. You will work closely with infrastructure, Kubernetes, and distributed networks to break things in staging so they never break in production.
What You’ll Work On
Kubernetes-based distributed systems and multi-cluster environments
Network-heavy distributed systems and service-to-service communication
L2/L3 networking and cloud networking scenarios
Load balancers, VPCs, routing, DNS and connectivity
Network failures, latency, packet loss and service degradation
Observability and alerting pipelines
Reliability validation and failure-injection testing
API, integration, system and end-to-end automation
Infrastructure and deployment automation
CI/CD pipelines and reliability gates
Key Responsibilities
Design and build scalable automation frameworks for network, API, integration and system-level validation
Build automated tests for service-to-service communication and distributed workflows
Validate system behavior under network failures, latency, packet loss, connectivity issues and partial failures
Develop failure-injection and reliability tests for distributed systems
Validate Kubernetes networking and multi-cluster behavior
Integrate automated validation into CI/CD pipelines
Analyze failures across logs, metrics, traces and network signals
Partner with SRE, Platform and Backend teams to make systems more testable and observable
Perform deep-dive root cause analysis on system and network failures
Build reusable testing utilities and infrastructure
Help establish reliability and quality standards across the platform
Technical Expectations
Deep Expertise In
Networking fundamentals and distributed systems
TCP/IP, DNS, HTTP/HTTPS and service communication
Kubernetes networking
Network failure modes and troubleshooting
Distributed-system failure scenarios
Observability and production debugging
Strong Hands-on Experience With
Kubernetes / multi-cluster environments
AWS, Azure or GCP networking
VPCs, load balancers, routing and IAM
Network troubleshooting and debugging
API and integration testing
Chaos / failure-injection testing
CI/CD automation
Programming
Strong coding skills in Python and Go (mandatory)
Experience building automation frameworks and system-level tooling
Proficiency in Shell scripting and infrastructure automation
What Makes This Role Different
You won't just test whether an API returns the expected response. You'll ask:
"What happens when the network behaves badly?"
Can the system detect it?
Can it recover?
Does it fail safely?
Are the right signals generated?
Can we automatically validate that behavior?
That's the kind of QA engineering we're looking for.
About Ciroos
We are an early-stage AI startup focused on Site Reliability Engineering (SRE). Rather than being another observability platform, its goal is to act as an AI SRE teammate that works alongside SRE, DevOps, Platform Engineering, Cloud Operations, and IT Operations teams to investigate incidents, determine root causes, and automate remediation.
Our team includes experienced entrepreneurs and engineers who have built multiple billion-dollar products from scratch. As a well-funded US-based company backed by top-tier VCs, we have offices in the US, India, and Europe. Join us in our fast-paced environment where you’ll have a front-row seat to shape the future of AI-driven Observability solutions.
As published by ashby
Name, Email, Resume, Location
- Mobile/Phone
- LinkedIn Profile Link
- What excites you most about the opportunity to join an early-stage startup like Ciroos? written answer
- What value do you hope to gain by working at an early-stage startup like Ciroos? written answer
- What value would you bring to Ciroos? written answer
- Total Work Experience
- Current / Last Drawn CTC
- Availability to Join choose one