Technology Infrastructure Services Engineer
Summary: The Technology Infrastructure Services Engineer is crucial for designing and managing enterprise-scale Ceph storage platforms essential for business and research workloads in hybrid and cloud-native settings. The role emphasizes building, supporting, and automating Ceph infrastructure to ensure reliability, performance, and reduced operational overhead.
Main Responsibilities:
Design, implement, and improve enterprise-scale Ceph storage platforms.
Build, support, and operationalize Ceph infrastructure focusing on high availability and scalability.
Develop automation, infrastructure-as-code, and monitoring capabilities.
Work with complementary storage technologies such as Weka and Qumulo.
Maintain robust DevOps and SRE practices to enhance platform reliability.
Enjoy operating large-scale distributed systems to improve automation.
Key Requirements:
Minimum 2 years of hands-on experience with Ceph storage clusters.
Proven capability in operationalizing Ceph including health monitoring and troubleshooting.
Experience in automating operations using Ansible, Terraform, Python, or similar frameworks.
Strong understanding of distributed storage concepts including CRUSH maps and CephFS.
Ability to integrate Ceph with Kubernetes or similar platforms.
Nice to Have:
Experience with other HPC storage platforms.
Knowledge of hybrid cloud infrastructure, especially AWS.
Experience in working with AI/ML workloads and Kubernetes environments.