Software Engineer- AI-Driven SRE & Cloud SRE
Summary
Junior SRE role at CLOUDSUFI (Noida) embedded in a client engagement supporting a regulated consumer-lending platform on AWS. Day to day: AI-assisted incident response, Datadog monitoring as code via Terraform, cloud reliability and cost engineering, plus Python/Bash automation under senior guidance.
About Us
Responsibilities - AI-Driven Initiatives
1. AI-Augmented Incident Response & Root Cause Analysis
- Support incident response as a shadow or secondary responder, using AI-driven detection, correlation, and root-cause analysis tooling under the guidance of senior engineers to speed up triage.
- Help assemble incident timelines from metrics, logs, traces, and deploy history, learning to separate the alert that fired from the change that actually caused it.
- Help operate and monitor the existing AI SRE sub-agents (covering areas such as incident summarization, monitor-gap detection, and usage attribution), flagging anomalies or failures for senior review.
- Assist in documenting AI-assisted incident findings and postmortems, including tracking corrective and preventive actions through to closure.
2. AI-Driven Observability & Alert Quality
- Build and maintain monitors, dashboards, and SLO definitions in Datadog as code via Terraform, rather than clicking through the UI.
- Assist in reducing alert noise using AI/ML-assisted detection (anomaly, outlier, and forecast monitors, composite conditions, and dynamic thresholds) so that pages stay actionable and rare.
- Run recurring monitor-hygiene and coverage-gap reviews - no-data monitors, missing no-data notification, orphaned monitors on decommissioned services, and monitors with no owning team tag or valid notification target.
- Support SLI/SLO and error-budget reviews for critical customer journeys, escalating services trending toward budget exhaustion to senior team members.
3. AI-Enabled Cloud Reliability, Capacity & Cost Engineering
- Support day-to-day reliability of AWS workloads - ECS/Fargate, EKS, Lambda, RDS/Aurora, ALB, SQS/SNS and Step Functions - under senior guidance.
- Use predictive analytics and intelligent monitoring dashboards to spot capacity, saturation, and cost anomalies across host, container, log, APM, and serverless usage, escalating notable trends with evidence.
- Assist in observability and cloud cost governance, attributing spend and telemetry volume to owning teams and services and helping identify low-value, high-cost signals.
4. AI Governance, Risk & Reliability Reviews (Learning Track)
- Participate in architecture and design reviews, learning how reliability risk and AI-agent risk (e.g., prompt injection, authorization gaps, secrets exposure in agent and tool configuration, blast radius of autonomous actions) is assessed and documented.
- Learn how production-readiness, change management, and audit expectations apply in a regulated environment (PCI-DSS, SOC 2, SOX, GLBA), and why observability evidence matters to auditors.
- Assist in tracking action items from the AI-risk and reliability remediation roadmap and help prepare status updates for stakeholders.
5. AI-Enabled DevOps, CI/CD & Automation
- Support the implementation of reliability and safety controls within CI/CD pipelines, learning automation-first and AI-augmented approaches and shift-left practices from senior engineers.
- Help integrate AI-assisted checks (e.g., change-risk summaries, infrastructure drift detection, dependency and configuration checks) into the delivery lifecycle.
- Write and maintain automation in Python and Bash against platform APIs (Datadog, AWS, GitHub, PagerDuty, Jira) to replace recurring manual work, and contribute Terraform modules and pull requests under review.
- Assist with runbook creation and the progressive automation of runbook steps toward self-healing.
6. AI-Enabled Program & Workflow Support
- Support multiple concurrent reliability initiatives using AI-enabled workflow and tracking tools shared across SRE and Security teams.
- Help promote a reliability-first culture by learning and sharing data-driven, AI-backed operational practices with the wider team.
SRE & Cloud Capabilities
- Interest in observability and monitoring fundamentals - metrics, logs, traces, and APM - and a willingness to learn predictive analytics and anomaly detection.
- Exposure to at least one observability platform (Datadog preferred; Grafana/Prometheus, New Relic, CloudWatch, or ELK/OpenSearch also relevant).
- Foundational understanding of SLI, SLO, and error-budget concepts, and of why alert fatigue is a reliability problem rather than an annoyance.
- Hands-on exposure to AWS (or an equivalent hyperscaler) across compute, networking, storage, and managed database services, with interest in automated reliability tooling.
- Foundational knowledge of container and orchestration concepts (Docker, ECS or Kubernetes) and of serverless execution models.
- Beginner-to-intermediate familiarity with Infrastructure as Code (Terraform preferred; Ansible or CloudFormation acceptable) and an eagerness to build deeper IaC expertise.
- Basic Linux troubleshooting and networking fundamentals, including DNS, TLS, load balancing, timeouts and retries.
- Some hands-on exposure to automation or scripting (Python, Bash, or similar), including consuming REST APIs and parsing JSON.
- Comfort with Git and pull-request-based workflows, and exposure to CI/CD tooling (GitHub Actions, Jenkins, GitLab CI, ArgoCD or similar).
- Familiarity with incident management and on-call concepts, including severity models, escalation policies, and paging tools such as PagerDuty or Opsgenie.
- Basic exposure to cloud posture and continuous monitoring concepts, and to change and release management discipline.
- Practical curiosity about LLM-based assistants and agents applied to operations, and about how to verify whether their output is actually correct.
About You
- 1–3 years of experience in SRE, DevOps, cloud infrastructure, platform, or production-support engineering, with an eagerness to build deeper expertise in infrastructure as code.
- Some exposure to cloud-native logging and monitoring tools, and interest in AI-driven alerting and noise reduction.
- You debug from evidence: when something breaks, your instinct is to look at the data before offering a theory, and you are comfortable saying “I don’t know yet, here is what I am checking.”
- You would rather automate a task the second time you do it than the tenth.
- Coursework or hands-on exposure to monitoring and observability platforms.
- Exposure to CI/CD pipelines and interest in integrating reliability controls and shift-left practices.
- Basic understanding of capacity, performance, and reliability cost concepts, including risk-based prioritization of remediation work.
- Awareness of common compliance frameworks (PCI-DSS, SOC 2, SOX, HIPAA/GLBA) and comfort working with production-change discipline.
- Willingness to learn incident management processes, including AI-assisted triaging, and the judgment to escalate early rather than sit quietly on an uncertain production signal.
- Strong analytical and problem-solving skills, with curiosity to assess simple architectures for failure modes under senior guidance.
- Clear written communication - you can explain an incident, a metric, or a trade-off to someone who was not in the room.
- Interest in learning how reliability standards, operational policies, and governance frameworks are developed and maintained, including for AI-agent usage.
Core Competencies
- Site Reliability Engineering fundamentals (SLI/SLO, error budgets, toil reduction)
- Observability and Monitoring (metrics, logs, distributed tracing, APM, RUM/Synthetics)
- AIOps basics - anomaly detection, event correlation, predictive analytics, alert noise reduction
- Incident Response and Postmortem/RCA practice
- Cloud Infrastructure (AWS core services, multi-account and IAM basics)
- Containers and Orchestration (Docker, ECS, Kubernetes)
- Infrastructure as Code and Configuration Management (Terraform, Ansible)
- CI/CD, Release Engineering and Progressive Delivery
- Automation and Scripting (Python, Bash, platform APIs)
- Capacity, Performance and Reliability Cost Engineering (FinOps fundamentals)
- Resilience Patterns (retries, timeouts, circuit breakers, graceful degradation)
- Chaos/Failure-Injection and Disaster Recovery concepts
- Runbook Automation, Automated Remediation and Self-Healing (exposure)
- DevSecOps and secure-by-default operations fundamentals
- AI Agent Risk & Governance basics (authorization models, prompt-injection risk, secrets exposure in AI tool configs, human-in-the-loop controls)
Preferred Certifications
- AWS Certified Cloud Practitioner / Associate-level AWS certification (preferred, not required)
- HashiCorp Certified: Terraform Associate (preferred, not required)
- Datadog Fundamentals or an equivalent observability-platform certification (preferred, not required)
- Certified Kubernetes Administrator (CKA) or KCNA (preferred, not required)
- Exposure to AI/ML applications in operations and reliability is a plus (preferred, not required)
Skills
- Agentic AI
- AI
- Analytics
- Anomaly Detection
- Ansible
- API
- Argo CD
- Aurora
- Automation
- AWS
- Bash
- CI/CD
- Cloud
- Cloud Native
- CloudFormation
- CloudWatch
- Data Engineering
- Datadog
- DevOps
- DevSecOps
- DNS
- Docker
- ECS
- EKS
- ELK
- Event Driven Architecture
- FinOps
- Git
- GitHub
- GitHub Actions
- GitLab
- Grafana
- Hipaa
- IAM
- Infrastructure as Code
- Jenkins
- Jira
- JSON
- Kubernetes
- Lambda
- Linux
- LLM
- Machine Learning
- Microservices
- Networking
- New Relic
- NLP
- Observability
- OpenSearch
- PagerDuty
- Pci Dss
- Predictive Analytics
- Prometheus
- Python
- RDS
- REST
- Serverless
- SNS
- SOC 2
- SQS
- Terraform
- TLS