DevOps & Infrastructure Engineer (On-Premises)
Summary
Owns and automates on-premises infrastructure, including GPU servers for AI/ML, while building secure CI/CD pipelines and observability stacks for local production systems.
Fortek (Private) Limited is expanding itson-premises ICT infrastructure and data center footprint, including GPU-enabledcompute for AI/ML workloads. The DevOps & Infrastructure Engineer will bethe operational backbone of this environment — owning the full lifecycle ofcompany-owned physical and virtual servers, from provisioning and hardeningthrough to automation, monitoring, and disaster recovery. This is a hands-on,individual-contributor role for someone who is equally comfortable in a serverroom, a Kubernetes cluster, and a runbook.
The incumbent will work closely with the Head ofEngineering/CTO and cross-functional engineering teams to build repeatable,secure, and well-documented infrastructure practices, while directly supportinglocal production systems, CI/CD pipelines, and GPU-accelerated deployments thatunderpin Fortek's service delivery to clients across Pakistan's ICT and datacenter sector.
Position Structure
Department
IT
Line Manager
Stream
DevOps & Infra
Job Requirements
Skills & Tools
Educational Qualification:
- Bachelor'sdegree in Software Engineering, Computer Science, Information Technology, or arelated field.
Certifications (Preferred)
Other Skills & Tools
- Linuxadministration; bare-metal and virtualized infrastructure (VMware, Proxmox,Hyper-V); Docker/Kubernetes; Ansible; Git; Nginx/HAProxy.
- Grafana,Prometheus, Loki, Tempo/OpenTelemetry; networking fundamentals (VLANs, DNS,DHCP, VPNs, firewalls).
- NVIDIA GPUservers, drivers, CUDA and container toolkit; storage, backups, access control,patching and disaster recovery.
Job Experience Required
- 2+ yearsof hands-on experience in DevOps, systems administration, infrastructureengineering, platform engineering, or a related role.
- Experiencesupporting local production servers, building CI/CD pipelines, automatinginfrastructure and troubleshooting live systems.
- Experiencedeploying applications or AI/ML workloads on GPU-enabled servers, including GPUcontainers and monitoring, is a strong advantage.
At Fortek, we believe in empowering our people. We offer:
Competitive salary based on experience and qualifications
Annual performance-based increments & bonuses
Medical facility
Paid annual, casual & sick leaves
Gratuity and EOBI benefits
Collaborative, learning-driven work environment, and much more
Let’s build the future together!
Duties & Responsibilities
- Design,install, configure and maintain secure production infrastructure oncompany-owned physical and virtual servers.
- AdministerLinux systems and virtualization clusters (VMware, Proxmox or Hyper-V); managecompute, memory, storage, OS patching and capacity.
- Maintainworking knowledge of Windows Server for mixed environments.
2. DevOps Automation & CI/CD
- Build andoperate self-hosted CI/CD runners, source-control integrations, artifactrepositories and container registries.
- Automateprovisioning, configuration and repeatable operations using Ansible, scripting(Bash/Python) and infrastructure-as-code practices.
- Deploycontainerized applications and AI/ML workloads on GPU-enabled local serversusing Docker/Kubernetes.
- MaintainNVIDIA GPU drivers, CUDA and container runtimes; monitor GPU utilization,memory, temperature, health and workload performance.
4. Network & Security Administration
- Manageinternal networking: VLANs, routing, DNS, DHCP, VPNs, reverse proxies, loadbalancers (Nginx/HAProxy) and firewalls.
- Applyleast-privilege access controls, secrets management, hardening, vulnerabilityremediation and audit logging.
- MaintainTLS certificate lifecycle across internal and client-facing services.
5. Monitoring, Backup & Disaster Recovery
- Implementand maintain observability using Grafana, Prometheus, Loki andTempo/OpenTelemetry, with centralized logging.
- Ownbackups, restore testing, replication, high availability, incident response,root-cause analysis and disaster recovery.
6. Documentation & Stakeholder Coordination
- Maintaininventory, architecture diagrams, SOPs and runbooks.
- Coordinateinfrastructure changes with developers and provide on-call support as required.
- Log allincidents, service requests, risks and escalations in Odoo per Fortek's agreedtracking protocol.
Reporting Responsibilities (Daily, Weekly & Monthly)
Daily Reporting
- Monitor server, GPU, storage, network, application, database, backup, deployment and pipeline health; respond to alerts and incidents.
- Resolve incidents and service requests; log actions, risks and escalations in Odoo Helpdesk/Task modules.
Weekly Reporting
- Submit deployment, incident, reliability and capacity summaries to the Line Manager and engineering stakeholders via Odoo.
- Review infrastructure changes, patching, backup/restore status, vulnerabilities, access and capacity items.
Monthly Reporting
- Submit KPI report in Odoo covering uptime, incidents, MTTR, deployment performance, resource utilization, backup success and security posture.
- Review CPU/GPU/server/storage capacity, lifecycle needs, DR readiness, documentation and the infrastructure roadmap.