Technology Infrastructure Services Engineer

Summary: The Technology Infrastructure Services Engineer is crucial for designing and managing enterprise-scale Ceph storage platforms essential for business and research workloads in hybrid and cloud-native settings. The role emphasizes building, supporting, and automating Ceph infrastructure to ensure reliability, performance, and reduced operational overhead.

Main Responsibilities:

  • Design, implement, and improve enterprise-scale Ceph storage platforms.

  • Build, support, and operationalize Ceph infrastructure focusing on high availability and scalability.

  • Develop automation, infrastructure-as-code, and monitoring capabilities.

  • Work with complementary storage technologies such as Weka and Qumulo.

  • Maintain robust DevOps and SRE practices to enhance platform reliability.

  • Enjoy operating large-scale distributed systems to improve automation.

Key Requirements:

  • Minimum 2 years of hands-on experience with Ceph storage clusters.

  • Proven capability in operationalizing Ceph including health monitoring and troubleshooting.

  • Experience in automating operations using Ansible, Terraform, Python, or similar frameworks.

  • Strong understanding of distributed storage concepts including CRUSH maps and CephFS.

  • Ability to integrate Ceph with Kubernetes or similar platforms.

Nice to Have:

  • Experience with other HPC storage platforms.

  • Knowledge of hybrid cloud infrastructure, especially AWS.

  • Experience in working with AI/ML workloads and Kubernetes environments.