SRE
Key Responsibilities
Reliability & Operations
- Ensure high availability, scalability, and performance of production systems.
- Define and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs).
- Proactively identify and resolve system bottlenecks and performance issues.
- Perform capacity planning and infrastructure optimization.
Monitoring & Incident Management
- Implement and manage monitoring, logging, and alerting solutions.
- Lead incident response, root cause analysis (RCA), and post-incident reviews.
- Develop automated remediation and self-healing mechanisms.
- Manage on-call support rotations and production support activities.
Automation & Infrastructure
- Automate operational tasks using scripting and Infrastructure as Code (IaC).
- Design and implement CI/CD pipelines to enhance deployment efficiency.
- Standardize infrastructure provisioning and configuration management.
- Drive infrastructure modernization initiatives.
Cloud & Platform Engineering
- Manage cloud infrastructure across AWS, Azure, or GCP environments.
- Optimize cloud resource utilization, security, and cost management.
- Implement containerization and orchestration solutions using Docker and Kubernetes.
- Support hybrid and multi-cloud deployments.
Security & Compliance
- Ensure platform compliance with organizational security standards.
- Implement security best practices, vulnerability remediation, and access controls.
- Participate in disaster recovery planning and business continuity initiatives.
Required Skills
Technical Skills
- Strong experience with Linux/Unix administration.
- Proficiency in one or more programming/scripting languages:
- Python
- Shell Scripting
- Go
- Java
- Experience with cloud platforms:
- AWS
- Microsoft Azure
- Google Cloud Platform (GCP)
- Hands-on experience with:
- Kubernetes
- Docker
- Terraform
- Ansible
- Experience with CI/CD tools:
- Jenkins
- GitHub Actions
- GitLab CI/CD
- Azure DevOps
Monitoring & Observability
- Prometheus
- Grafana
- ELK Stack (Elasticsearch, Logstash, Kibana)
- Splunk
- Datadog
- New Relic
Database Knowledge
- SQL Server
- PostgreSQL
- MySQL
- MongoDB
- Redis
Qualifications
- Bachelor's degree in Computer Science, Information Technology, or related field.
- 5+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering.
- Experience supporting large-scale enterprise applications.
- Understanding of networking concepts, DNS, load balancing, and security principles.
Preferred Qualifications
- AWS Certified Solutions Architect / DevOps Engineer.
- Azure Administrator or Azure DevOps Engineer Certification.
- Google Professional Cloud DevOps Engineer Certification.
- Kubernetes certifications (CKA/CKAD).
- Experience in enterprise retail, eCommerce, or digital transformation projects.
Soft Skills
- Strong troubleshooting and analytical skills.
- Excellent communication and stakeholder management abilities.
- Ability to work in a fast-paced production environment.
- Strong collaboration and cross-functional teamwork skills.
- Continuous learning and improvement mindset.
Experience
5-10+ Years
Location
Bangalore / Hyderabad / Chennai / Pune (Hybrid/Remote)
Employment Type
Full-Time
Provide your feedback on BizChat
Key Responsibilities
Reliability & Operations
- Ensure high availability, scalability, and performance of production systems.
- Define and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs).
- Proactively identify and resolve system bottlenecks and performance issues.
- Perform capacity planning and infrastructure optimization.
Monitoring & Incident Management
- Implement and manage monitoring, logging, and alerting solutions.
- Lead incident response, root cause analysis (RCA), and post-incident reviews.
- Develop automated remediation and self-healing mechanisms.
- Manage on-call support rotations and production support activities.
Automation & Infrastructure
- Automate operational tasks using scripting and Infrastructure as Code (IaC).
- Design and implement CI/CD pipelines to enhance deployment efficiency.
- Standardize infrastructure provisioning and configuration management.
- Drive infrastructure modernization initiatives.
Cloud & Platform Engineering
- Manage cloud infrastructure across AWS, Azure, or GCP environments.
- Optimize cloud resource utilization, security, and cost management.
- Implement containerization and orchestration solutions using Docker and Kubernetes.
- Support hybrid and multi-cloud deployments.
Security & Compliance
- Ensure platform compliance with organizational security standards.
- Implement security best practices, vulnerability remediation, and access controls.
- Participate in disaster recovery planning and business continuity initiatives.
Required Skills
Technical Skills
- Strong experience with Linux/Unix administration.
- Proficiency in one or more programming/scripting languages:
- Python
- Shell Scripting
- Go
- Java
- Experience with cloud platforms:
- AWS
- Microsoft Azure
- Google Cloud Platform (GCP)
- Hands-on experience with:
- Kubernetes
- Docker
- Terraform
- Ansible
- Experience with CI/CD tools:
- Jenkins
- GitHub Actions
- GitLab CI/CD
- Azure DevOps
Monitoring & Observability
- Prometheus
- Grafana
- ELK Stack (Elasticsearch, Logstash, Kibana)
- Splunk
- Datadog
- New Relic
Database Knowledge
- SQL Server
- PostgreSQL
- MySQL
- MongoDB
- Redis
Qualifications
- Bachelor's degree in Computer Science, Information Technology, or related field.
- 5+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering.
- Experience supporting large-scale enterprise applications.
- Understanding of networking concepts, DNS, load balancing, and security principles.
Preferred Qualifications
- AWS Certified Solutions Architect / DevOps Engineer.
- Azure Administrator or Azure DevOps Engineer Certification.
- Google Professional Cloud DevOps Engineer Certification.
- Kubernetes certifications (CKA/CKAD).
- Experience in enterprise retail, eCommerce, or digital transformation projects.
Soft Skills
- Strong troubleshooting and analytical skills.
- Excellent communication and stakeholder management abilities.
- Ability to work in a fast-paced production environment.
- Strong collaboration and cross-functional teamwork skills.
- Continuous learning and improvement mindset.
Experience
5-10+ Years
Location
Bangalore / Hyderabad / Chennai / Pune (Hybrid/Remote)
Employment Type
Full-Time
Provide your feedback on BizChat
Key Responsibilities
Reliability & Operations
- Ensure high availability, scalability, and performance of production systems.
- Define and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs).
- Proactively identify and resolve system bottlenecks and performance issues.
- Perform capacity planning and infrastructure optimization.
Monitoring & Incident Management
- Implement and manage monitoring, logging, and alerting solutions.
- Lead incident response, root cause analysis (RCA), and post-incident reviews.
- Develop automated remediation and self-healing mechanisms.
- Manage on-call support rotations and production support activities.
Automation & Infrastructure
- Automate operational tasks using scripting and Infrastructure as Code (IaC).
- Design and implement CI/CD pipelines to enhance deployment efficiency.
- Standardize infrastructure provisioning and configuration management.
- Drive infrastructure modernization initiatives.
Cloud & Platform Engineering
- Manage cloud infrastructure across AWS, Azure, or GCP environments.
- Optimize cloud resource utilization, security, and cost management.
- Implement containerization and orchestration solutions using Docker and Kubernetes.
- Support hybrid and multi-cloud deployments.
Security & Compliance
- Ensure platform compliance with organizational security standards.
- Implement security best practices, vulnerability remediation, and access controls.
- Participate in disaster recovery planning and business continuity initiatives.
Required Skills
Technical Skills
- Strong experience with Linux/Unix administration.
- Proficiency in one or more programming/scripting languages:
- Python
- Shell Scripting
- Go
- Java
- Experience with cloud platforms:
- AWS
- Microsoft Azure
- Google Cloud Platform (GCP)
- Hands-on experience with:
- Kubernetes
- Docker
- Terraform
- Ansible
- Experience with CI/CD tools:
- Jenkins
- GitHub Actions
- GitLab CI/CD
- Azure DevOps
Monitoring & Observability
- Prometheus
- Grafana
- ELK Stack (Elasticsearch, Logstash, Kibana)
- Splunk
- Datadog
- New Relic
Database Knowledge
- SQL Server
- PostgreSQL
- MySQL
- MongoDB
- Redis
Qualifications
- Bachelor's degree in Computer Science, Information Technology, or related field.
- 5+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering.
- Experience supporting large-scale enterprise applications.
- Understanding of networking concepts, DNS, load balancing, and security principles.
Preferred Qualifications
- AWS Certified Solutions Architect / DevOps Engineer.
- Azure Administrator or Azure DevOps Engineer Certification.
- Google Professional Cloud DevOps Engineer Certification.
- Kubernetes certifications (CKA/CKAD).
- Experience in enterprise retail, eCommerce, or digital transformation projects.
Soft Skills
- Strong troubleshooting and analytical skills.
- Excellent communication and stakeholder management abilities.
- Ability to work in a fast-paced production environment.
- Strong collaboration and cross-functional teamwork skills.
- Continuous learning and improvement mindset.
Experience
5-10+ Years
Location
Bangalore / Hyderabad / Chennai / Pune (Hybrid/Remote)
Employment Type
Full-Time
Provide your feedback on BizChat
Skills
- Ansible
- Automation
- AWS
- Azure
- Azure DevOps
- Bash
- CI/CD
- Cloud
- Containerization
- Datadog
- DevOps
- DNS
- Docker
- E-commerce
- Elasticsearch
- ELK
- GCP
- GitHub
- GitHub Actions
- GitLab
- Grafana
- Infrastructure as Code
- Java
- Jenkins
- Kibana
- Kubernetes
- Linux
- Logstash
- MongoDB
- MySQL
- Networking
- New Relic
- Observability
- PostgreSQL
- Prometheus
- Python
- Redis
- Splunk
- SQL
- SQL Server
- Stakeholder Management
- Terraform
- Unix