DevOps Manager
DevOps Engineer
Company: Invigilo AI
Location: Singapore / Hybrid or Remote
Employment Type: Full-time
About Invigilo AI
Invigilo AI builds AI-powered safety and operations technology for high-risk industrial environments including construction, manufacturing, logistics, oil and gas, mining and maritime.
Our platform turns existing CCTV and camera infrastructure into an intelligent safety layer, using computer vision and AI to identify risks such as unsafe behaviour, PPE violations, equipment interactions, restricted-area entry, falls, fire and smoke, and other operational events in real time.
Our systems operate across a mix of on-premise edge servers, cloud infrastructure and hybrid deployments, often in challenging industrial environments where reliability matters.
We are looking for a DevOps Engineer to help us build, deploy and operate the infrastructure that powers these systems globally.
What You'll Do
You will work closely with our AI, backend and deployment teams to ensure Invigilo's platform can be reliably deployed and operated across customer environments.
Your responsibilities will include:
Build and maintain CI/CD pipelines for our backend, AI and edge applications.
Containerize and deploy applications using Docker and related orchestration technologies.
Manage cloud infrastructure, primarily across platforms such as Azure and AWS.
Support deployment of Invigilo's AI stack on customer on-premise GPU servers.
Automate provisioning, configuration and deployment of new servers and environments.
Manage Linux-based infrastructure and troubleshoot production systems.
Build monitoring, logging and alerting systems for application, server, GPU and infrastructure health.
Monitor uptime and performance across distributed customer deployments.
Manage networking requirements including VPNs, firewalls, ports, RTSP streams and secure connectivity between customer sites and Invigilo systems.
Support GPU infrastructure used for computer vision and AI inference workloads.
Improve deployment reliability, rollback processes and disaster recovery procedures.
Maintain infrastructure-as-code and configuration management practices.
Work with engineering teams to improve application observability and production readiness.
Implement security best practices around access control, secrets, encryption, patching and infrastructure hardening.
Help troubleshoot deployments together with field engineers and customers when required.
Continuously improve the way we deploy and manage hundreds of cameras, AI services and edge devices across multiple sites.
What We're Looking For
You should have strong practical experience running production systems rather than only theoretical DevOps knowledge.
Ideally, you have experience with several of the following:
Linux administration
Docker and containerized applications
CI/CD using GitHub Actions, GitLab CI, Jenkins or similar
Azure, AWS or another major cloud platform
Terraform, Ansible or similar infrastructure automation tools
Kubernetes or container orchestration
Nginx and reverse proxies
Networking, VPNs, DNS, firewalls and TCP/IP
Monitoring tools such as Prometheus, Grafana, ELK, Loki or similar
PostgreSQL, Redis or other production databases
Bash and/or Python scripting
Git and modern software development workflows
Experience working with GPU servers, NVIDIA drivers, CUDA, computer vision systems, video streaming or RTSP infrastructure would be particularly valuable.
Bonus Points
Experience in any of the following would be a strong advantage:
NVIDIA GPUs, CUDA and GPU monitoring
Edge AI or computer vision deployments
CCTV, IP cameras, NVRs and RTSP streaming
NVIDIA DeepStream, GStreamer or FFmpeg
Large-scale distributed edge deployments
Azure IoT or edge computing platforms
High-availability infrastructure
Cybersecurity and infrastructure hardening
ISO 27001 environments
Deploying systems into enterprise or industrial customer networks
What Success Looks Like
Within the role, you will help us reach a point where:
A new customer deployment can be provisioned rapidly and consistently.
Software updates can be safely deployed across distributed sites with minimal manual intervention.
Infrastructure failures are detected before customers report them.
GPU, camera and application health can be monitored centrally.
Engineering teams can confidently release new versions without worrying about deployment complexity.
Our infrastructure scales from tens of deployments to hundreds and eventually thousands of sites.
Who You'll Work With
You will work directly with our:
AI and Computer Vision Engineers
Backend Engineers
Product Team
Project and Deployment Teams
Customer Success Team
Because our technology operates in real industrial environments, you will also occasionally work through unusual infrastructure challenges that you would rarely encounter in a conventional SaaS company, from isolated customer networks and unreliable connectivity to on-premise GPU servers processing dozens of live video streams.
Why Invigilo
At Invigilo, DevOps is not simply about keeping a web application online.
The infrastructure you build supports AI systems monitoring real construction sites, factories, logistics facilities and other high-risk workplaces. Reliability directly affects whether our platform can identify safety risks when they matter.
You'll have significant ownership over how Invigilo's infrastructure evolves as we scale internationally and expand our AI platform.