Site Reliability Engineer (SRE)
Summary
Site Reliability Engineer at iCareManager responsible for ensuring reliability and availability of production systems in a healthcare environment, focusing on automation, observability, and incident management.
Role Summary
The Site Reliability Engineer (SRE) at iCareManager (iCM) is responsible for ensuring the reliability, scalability, performance, and availability of production systems. The role blends software engineering and operations, with a strong focus on automation, observability, incident management, and proactive reliability engineering.
SREs enable engineering teams to deliver features rapidly without compromising system stability, especially in regulated healthcare environments.
Key Responsibilities
1. Reliability Engineering
Define, implement, and enforce Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets
Design and review resilience patterns (redundancy, failover, graceful degradation)
Perform capacity planning, load modeling, and scalability analysis
Conduct chaos testing and failure injection to identify system weaknesses
Reduce Mean Time to Recovery (MTTR) through architectural improvements and tooling
2. Observability & Monitoring
Instrument systems with metrics, logs, and distributed traces
Build and maintain dashboards that reflect system health and performance
Design alerting strategies that are actionable and minimize alert fatigue
Identify leading indicators of failure before customer impact
3. Incident Management & Postmortems
Participate in and lead production incident response
Coordinate with engineering and infrastructure teams during incidents
Lead blameless postmortems and document root cause analysis
Track and remediate reliability debt and systemic risks
Requirements
Required Qualifications
Bachelor’s degree in Computer Science, Engineering, or equivalent experience
3+ years experience in SRE, DevOps, Platform Engineering, or similar roles
Strong programming experience in one or more languages (e.g., Dot net and or node JS
Hands-on experience with Linux-based systems
Experience with cloud platforms (Azure preferred; AWS/GCP acceptable)
Solid understanding of networking, distributed systems, and system design
Experience with monitoring and observability tools (e.g., Prometheus, Grafana, ELK, datadog)
Preferred Qualifications
Experience in healthcare or regulated environments
Familiarity with containerization and orchestration (Docker, Kubernetes)
Experience with CI/CD pipelines and infrastructure as code
Understanding of security best practices in production systems
Experience supporting SOC2-compliant environments
Key Competencies
Strong problem-solving and analytical skills
Calm and effective during high-pressure incidents
Excellent documentation and communication skills
Ownership mindset and bias toward automation
Collaborative and proactive approach
Benefits
- A dynamic and collaborative work environment.
- Opportunities for professional growth and skill development.
- Competitive salary and benefits package.
- The chance to play a key role in revolutionizing the healthcare technology industry.