Intermediate SRE / DevOps Engineer
We're looking for an infrastructure-minded engineer who is passionate about technology, automation, and reliability. If you are ready to take on exciting challenges, build robust cloud-native systems, and grow within a supportive environment oriented towards continuous improvement, we would love to hear from you.
Company Overview
Kriterion is a Pretoria-based SaaS company at the cutting edge of predictive maintenance and intelligent condition monitoring. As a rapidly growing scale-up, Kriterion has proven its business model and is now focused on expanding our market presence and refining our operations to meet increasing demand.
Our clients are those who take care of the heavy machinery and infrastructure that underpin our modern society. We at Kriterion enable our clients' maintenance teams to perform at their very best by providing actionable AI-driven decisions. By applying AI insight to sensor-rich assets we ensure that they remain healthy and perform optimally. We help reduce emissions by making these assets perform efficiently through lean maintenance while mitigating downtime and extending their useful lives.
We achieve this by incorporating deep learning with engineering insights in a modern AI-centric framework. Our core product is Cerberus; a cloud-native predictive maintenance platform. Our client base is in the telecommunications and mining space, with 30 000+ assets in our portfolio.
Job Description
As an Intermediate SRE / DevOps Engineer at Kriterion, you'll join a vibrant, collaborative team that thrives on innovation. Because we are a smaller scale-up, roles naturally overlap. You will bridge the gap between core infrastructure and software engineering. Your primary focus will be designing, scaling, and maintaining the highly available cloud environments that power the Cerberus platform, while collaborating closely with developers and data scientists to streamline ML and application deployments.
Your day-to-day job will focus on:
- Infrastructure & Reliability: Design, build, and monitor highly available, scalable infrastructure on Google Cloud Platform and Kubernetes to ensure Cerberus performs flawlessly for our clients.
- CI/CD & Automation: Build, maintain, and optimize robust deployment pipelines (GitLab CI, ArgoCD) to enable seamless, automated delivery of software and machine learning models.
- Developer Experience (DevEx): Act as the bridge between development and operations. Write clean, tested code to build internal tooling that empowers the wider engineering team to deploy safely and autonomously.
- Observability & Incident Response: Implement comprehensive monitoring, logging, and alerting systems to proactively identify and resolve bottlenecks or failures.
- Continuous Learning: Stretch your boundaries in a supportive environment where you will gain a broad skill set across cutting-edge cloud and MLOps technologies.
Requirements
- 2+ years of experience in DevOps, Site Reliability Engineering, or Software Engineering with a strong infrastructure focus.
- Hands-on experience with containerization and orchestration (Kubernetes, Docker).
- Experience building and maintaining CI/CD pipelines and utilizing Infrastructure as Code (IaC) principles.
- A strong focus on system architecture, cloud-native design, and platform security.
- Ability to write clean, maintainable automation code (Python, Bash, or similar).
- Proficient verbal and written communication.
- Relevant bachelor’s degree (Comp. Sci., Engineering, etc.).
Technology Stack
Because you will be working across our tech ecosystem, experience in the following (or related) technologies is advantageous. We emphasize our infrastructure stack for this role, but you will interact with the entire system:
- Infrastructure & Orchestration: Google Cloud Platform (GCP), Kubernetes, Linux
- CI/CD & GitOps: GitLab CI, ArgoCD, Crossplane
- Backend & Scripting: Python, FastAPI, FireStore
- AI/Data Ops (MLOps): Apache Beam, Tensorflow (TFX), Kubeflow, DBT, Prefect, BigQuery
- Frontend (Good to know): React.js, TypeScript
Benefits
- Flexible working hours
- Career growth opportunities in a rapidly expanding company
- Chance to contribute to significant company milestones and achievements
- Exposure to multiple aspects of a scaling businessHighly competitive compensation