Site Reliability Engineer
Summary
Maintain and automate cloud infrastructure for an AI-driven healthcare platform, ensuring reliability and scalability of Kubernetes clusters and CI/CD pipelines.
Site Reliability Engineer
Location: Bengaluru, India
Department: Site Reliability Engineering
Experience: 2-4
Site Reliability Engineer (SRE)
About nference
Our People
The Opportunity
What You'll Do
- Monitor the health, availability, and performance of cloud infrastructure, Kubernetes clusters, CI/CD systems, and platform services.
- Assist in production incident response by collecting logs, metrics, and diagnostic information to support rapid troubleshooting and recovery.
- Configure, maintain, and improve monitoring dashboards, alerting systems, and observability platforms.
- Develop automation scripts and operational tooling to eliminate manual processes and improve infrastructure efficiency.
- Contribute to Infrastructure-as-Code (IaC) modules for provisioning and managing cloud resources.
- Support the reliability, maintenance, and continuous improvement of CI/CD pipelines and deployment workflows.
- Participate in production operations while learning and applying Site Reliability Engineering principles, including SLIs, SLOs, error budgets, and incident management.
- Collaborate with software engineers to improve system reliability, scalability, and deployment processes.
- Troubleshoot infrastructure, networking, and platform-related issues across cloud environments.
- Maintain operational documentation, runbooks, and knowledge repositories to improve incident response and operational consistency.
- Continuously identify opportunities to improve platform reliability, automation, and operational excellence.
What We're Looking For
Required Qualifications
- Bachelor's or Master's degree in Computer Science, Information Technology, Software Engineering, or a related technical discipline.
- 1–3 years of professional experience in Site Reliability Engineering, DevOps, Cloud Infrastructure Engineering, or a related systems engineering role.
- Strong understanding of Linux systems administration, command-line tools, process management, system diagnostics, and troubleshooting.
- Proficiency in at least one scripting or programming language such as Python, Bash, or Shell for automation and tooling.
- Hands-on experience with at least one major cloud platform (AWS, GCP, or Azure) and familiarity with core infrastructure services including compute, networking, and storage.
- Working knowledge of containerization and orchestration technologies such as Docker and Kubernetes.
- Familiarity with Infrastructure-as-Code (IaC) concepts and tools such as Terraform or similar automation frameworks.
- Experience with monitoring, logging, and observability platforms such as Prometheus, Grafana, ELK Stack, Datadog, or similar tools.
- Good understanding of networking fundamentals including DNS, TCP/IP, load balancing, and distributed systems concepts.
- Familiarity with CI/CD pipelines and modern software delivery practices.
- Experience using version control systems such as Git.
- Strong analytical, troubleshooting, and problem-solving skills.
- Excellent verbal and written communication skills.
Preferred Qualifications
- Exposure to incident management and production support in cloud-native environments.
- Familiarity with reliability engineering concepts including SLIs, SLOs, and error budgets.
- Experience with configuration management or automation tools such as Ansible.
- Knowledge of messaging technologies such as Kafka or RabbitMQ.
- Exposure to distributed systems and microservices architectures.
- Experience supporting highly available production environments.
- Familiarity with security best practices for cloud infrastructure.
- Contributions to open-source projects are a plus.
Why Join nference?
Benefits & Perks
- Industry Prestige: Build your career at the "Google of Biomedicine" (as recognized by The Washington Post), working alongside exceptional software engineers, physicians, scientists, and researchers.
- Cutting-Edge Innovation: Solve complex healthcare challenges using advanced AI, machine learning, and large-scale clinical, molecular, and imaging datasets.
- Meaningful Impact: Contribute to technologies that accelerate drug discovery and biomedical research, with opportunities to be recognized as a contributing author on high-impact scientific publications where applicable.
- Growth & Flexibility: Thrive in a collaborative, innovation-driven culture with continuous learning opportunities and a hybrid work model for eligible employees after successful completion of the three-month probation period.
- Wellness & Perks: Enjoy reimbursements for gym memberships, technology gadgets, high-speed internet, professional development, comprehensive health insurance, and complimentary breakfast, lunch, and snacks at our Bangalore office.
