freehire launches on Product Hunt on 26 August.

Follow →

Linux Systems Engineer

Summary

Maintains and secures Linux servers, Kubernetes clusters, and scheduler-backed compute for a US hedge fund, automating ops and troubleshooting incidents.

Project description

We are looking for a Linux Systems Engineer for Large US hedge fund to keep critical Linux-based research and platform infrastructure reliable, secure, and efficient. This role supports production systems across Linux servers, Kubernetes platforms, scheduler-backed compute environments, and user access services. The engineer will troubleshoot incidents, automate operational work, improve observability, and partner with cross-functional teams to deliver stable infrastructure changes.

Responsibilities

  • Operate and support Linux servers and shared infrastructure used by research and platform teams.
  • Troubleshoot production issues involving system performance, availability, access, configuration, and networking.
  • Support Kubernetes and other container-based platforms, including node health, service behavior, and rollout activities.
  • Support scheduler-backed compute environments such as Slurm, including node readiness, maintenance, and incident recovery.
  • Improve user-facing Linux access services such as SSH, shared shell environments, and session-based platforms.
  • Manage OS lifecycle work: provisioning, patching, kernel and package updates, and hardening.
  • Build scripts and automation in Python, Bash, or similar tools to reduce manual work and improve reliability.
  • Use configuration management and version-controlled workflows to implement infrastructure changes safely.
  • Enhance monitoring, alerting, documentation, and operational processes.
  • Participate in incident response and occasional on-call support.

SKILLS

Must have

  • 3+ years of experience in Linux systems engineering, SRE, DevOps, or infrastructure support.
  • Strong Linux administration skills, including systemd, package management, permissions, filesystems, log analysis, and performance troubleshooting.
  • Good understanding of networking fundamentals such as DNS, NTP/PTP, routing, and general host connectivity.
  • Experience with automation, scripting, and operational tooling.
  • Familiarity with Kubernetes, virtualization, or clustered platforms.
  • Experience with configuration management or infrastructure-as-code tools such as Ansible, Salt, or Terraform.
  • Ability to troubleshoot production issues methodically and communicate clearly during incidents.
  • Experience with Git-based workflows and maintainable documentation.
  • Hands-on, practical problem solver with a strong ownership mindset.
  • Comfortable working close to production and balancing support with continuous improvement.
  • Collaborative communicator who works well across compute, storage, networking, and application teams.

Nice to have

• Experience with Slurm, HPC-style environments, GPU infrastructure, or researcher-facing Linux platforms. • Working familiarity with shared storage clients such as NFS, autofs, or GPFS / IBM Storage Scale from a host and application perspective. • Experience with observability tools such as Prometheus, Grafana, or equivalent platforms. • Exposure to identity and access services such as LDAP, Kerberos, SSSD, or PAM. • Exposure to on-premises datacenter operations, hardware lifecycle support, or vendor escalations. • Interest in using AI/ML techniques for infrastructure optimization, anomaly detection, or predictive operations.

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available