freehire launches on Product Hunt on 26 August.

Follow →

SRE Specialist II

Summary

Lead a Site Reliability Engineering team to define reliability strategy, observability, and automation for a critical financial-data product using Kubernetes, AWS, and observability tools.

Company Description

Experian is a global data and technology company that powers opportunities for people and businesses around the world. We operate in diverse markets, such as financial services, healthcare, automotive, agribusiness, insurance, among others. Experian invests in people and new advanced technologies to unlock the power of data. We have an incredible team of 25,200 employees in 32 countries.

Our uniqueness is valuing yours. Experian's people-centric, inclusive, and purpose-driven culture is recognized by numerous awards — including World’s Best Workplaces™ 2025 (Fortune's Top 25 global) and Great Place To Work™ in 26 countries, among others. Check out Experian Life on social media or explore our careers site to understand why. Experian is also proud to be an equal opportunity employer and affirmative action employer.

Job Description

Job Description

We are looking for a Principal Site Reliability Engineer to technically lead the Reliability Engineering discipline for a critical company product.

This professional will be responsible for defining the SRE technical strategy, evolving engineering standards, reliability, observability, and automation, in addition to acting as a technical reference for the team and a manager for the people who make up the team.

We expect a professional with strong leadership capacity, systemic vision, excellent decision-making, and the ability to influence different technology and business areas, ensuring high availability, resilience, and operational efficiency of services.

Daily Responsibilities

  • Lead the SRE team technically and managerially, promoting technical and career development for engineers;

  • Define and evolve the strategy for reliability, observability, and operational excellence of products;

  • Act as the primary technical reference in critical incidents, coordinating responses and root cause analyses;

  • Work together with Architecture, Development, Security, and Product teams on platform evolution;

  • Conduct operational rituals, incident reviews, capacity planning, and risk management;

  • Promote a culture of automation and continuous improvement;

  • Ensure the adoption of SRE, DevOps, and Platform Engineering best practices;

  • Support the prioritization of technical debt and reliability initiatives with business areas.

Key Deliverables

  • Evolution of product availability, reliability, and performance indicators;

  • Definition and monitoring of SLIs, SLOs, and Error Budgets;

  • Evolution of the observability platform (Logs, Metrics, Traces, and APM);

  • Automation of operational processes, reducing manual activities;

  • Reduction of MTTR and increase in incident prevention capacity;

  • Structuring of blameless Post Mortems;

  • Implementation of Chaos Engineering, Disaster Recovery, and Capacity Planning practices;

  • Definition of technical standards for infrastructure, monitoring, and operation;

  • Technical development of the team through mentoring, feedback, dojos, and technical communities.

What we are looking for in you

Leadership

  • Proven experience leading technical SRE, DevOps, or Platform teams;

  • Experience as a people manager;

  • Ability to influence different areas without direct authority;

  • Excellent executive and technical communication;

  • Experience in managing critical incidents and crises;

  • Strong analytical capacity and data-driven decision-making.

Qualifications

Qualifications

Technical Knowledge

  • Kubernetes (EKS/OpenShift)

  • Docker

  • AWS

  • Terraform

  • GitHub Actions, Jenkins, or other CI/CD tools

  • Observability (Dynatrace, Datadog, Grafana, Prometheus, ELK, OpenTelemetry)

  • Linux

  • Networks, DNS, HTTP, TLS, and load balancers

  • Automation using Python, Go, or Shell Script

  • Infrastructure as Code (IaC)

  • Performance Engineering

  • Resilience and High Availability

  • Capacity Management

  • Disaster Recovery

  • Chaos Engineering

  • Vulnerability Management

  • DevSecOps practices

It will be a plus

  • AWS Professional or Specialty certifications;

  • Kubernetes certification (CKA/CKS);

  • Experience with Service Mesh (Istio);

  • Experience with Backstage or Platform Engineering;

  • Knowledge of FinOps;

  • Experience in regulated environments and mission-critical products;

  • Experience with event-driven architecture (Kafka, RabbitMQ).

Additional Information

Serasa Experian is much more than you imagine. With the purpose of creating a better future by expanding opportunities for people and businesses, in Brazil we are more than 4,000 people working in diverse teams and specialties. Here, each knowledge and diversity complements each other and you can work on what you love most; we are committed to building an inclusive culture and an environment in which people can balance their careers with their personal commitments and interests, valuing well-being.

We are very dedicated to being one of the best and most innovative companies to work for in the country, enabling incredible experiences and careers for our people. Our strong people-first approach is recognized externally through various market certifications: we were awarded by Great Place To Work™ in 24 countries and by the international Top Employers certification, in addition to being recognized as one of the best companies for young professionals and having a 4.6 rating on Glassdoor. Each recognition indicates that we are on the right path, providing an increasingly better work environment for our talent.

Experian Careers - Creating a better tomorrow together

Find out what its like to work for Experian by clicking here

See also