Ai platform engineer
Summary
Builds and optimizes the AI platform, integrating NVIDIA GPUs for high-performance computing and deploying scalable AI infrastructure using Python, cloud tech, and AI/ML frameworks.
The AI Platform Engineer at Cassava AI Factory develops the AI platform, enabling users to build and deploy models efficiently. This role integrates the platform with NVIDIA GPUs for high-performance computing, ensuring scalability and user satisfaction. Responsible for the development and optimisation of AI infrastructure utilising NVIDIA products, i.e H200 GPUS.
Design and implement the infrastructure for building, deploying, and managing AI applications. The focus will be on creating scalable, reliable, and secure systems that support AI workflows. This role will involve building data pipelines, integrating machine learning models, and ensuring the efficiency of AI operations.
Working with our AI Solutions Architects who design the solutions, the AI Platform Engineers need to have a firm understanding of all the technologies being deployed and the ability to implement the proposed solutions. The Platform Engineer will be involved throughout the client engagement process, from initial scoping and surveys, through project implementation and delivery, and if required for ongoing support and maintenance of the customer environment.
Role Requirements Bachelor’s degree in Computer Science, Engineering, or a related field is required; Masters degree advantageous. Relevant certifications in AI infrastructure management or similar fields. Certifications in cloud computing or software engineering. Familiarity with GPU computing i.e. NVIDIA Blackwell-architecture GPUs or similar high-performance hardware and optimisation techniques for AI workloads, is essential for platform development. Familiarity with NVIDIA AI Enterprise software and experience building AI platforms, developer tools, or similar products highly advantageous. Contribute to open-source projects related to AI, cloud computing, or platform development. At least 2-3 years’ experience in platform engineering or software development. Strong programming skills in languages like Python, Java, or Go, with experience building scalable applications. Experience with cloud technologies, microservices architecture, and API development is a critical requirement. Knowledge of AI/ML workflows, tools, and frameworks, such as Tensor Flow or Py Torch.