freehire launches on Product Hunt on 26 August.

Follow →

Staff Software Engineer, Infrastructure

Open 36d

You'll play a key role in shaping engineering practices, ensuring the reliability and performance of the platform across a multi-cloud environment, and empowering the engineering team to deliver innovative features with speed and efficiency. You will drive improvements in automation, observability, and overall platform stability.

Responsibilities

  • Design, build, and operate scalable, resilient infrastructure across cloud providers including Azure, AWS, GCP, and IBM Cloud.
  • Lead infrastructure architecture and reliability improvements for critical custody services.
  • Implement and improve monitoring, alerting, logging, and observability across distributed systems.
  • Own and evolve blockchain node infrastructure, including high availability, failover, and provider management.
  • Build automation for infrastructure provisioning, deployments, testing, failover, incident response, and operational maintenance.
  • Drive deployment operations and release reliability for platform and product services.
  • Proactively identify performance bottlenecks, reliability risks, and operational gaps before they impact customers.
  • Participate in on-call rotations, support production incidents, and help improve incident response and post-incident learning.
  • Contribute to internal platform tools, services, and developer workflows.
  • Create clear documentation, runbooks, and operational procedures for critical systems.
  • Mentor engineers and provide technical guidance on infrastructure, reliability, and platform engineering decisions.

Requirements

  • 10+ years of experience in software engineering, platform engineering, infrastructure engineering, or systems operations for highly available production systems.
  • Experience managing blockchain nodes, including high availability, failover, and operational resilience.
  • Experience operating infrastructure in high-traffic, customer-critical, or security-sensitive environments.
  • Strong production experience with PostgreSQL.
  • Proven experience provisioning and managing infrastructure with Terraform or similar infrastructure-as-code tools.
  • Deep experience with containerized infrastructure and Kubernetes in highly available environments.
  • Proficiency with programming and scripting languages such as .NET, Go, Bash, or TypeScript.
  • Experience with messaging and queueing technologies such as RabbitMQ or AMQP.
  • Experience with observability tools such as OpenTelemetry, Grafana, Loki, Prometheus, or similar platforms.
  • Familiarity with GitOps practices using platforms like Argo CD or Flux.
  • Experience working with confidential computing solutions like AWS Nitro Enclaves, IBM Hyper Protect Virtual Servers, or GCP Confidential Computing.
  • Excellent ability in solving problems and the ability to diagnose complex distributed system issues.
  • Strong written and verbal communication skills, with a track record of working effectively across teams.
  • A focus on clear procedures, strong documentation habits, and a dedication to operational excellence.

Benefits

  • Professional development budget
  • Flexible in-office collaboration days
  • Bi-weekly all-company meeting with Leadership Team
  • Team offsites, team bonding activities, happy hours
  • Competitive bonuses and equity
  • Health, retirement, family forming, and family support benefits
  • Employee giving match
  • Mobile phone stipend
  • R&R days
  • Wellness reimbursement and weekly onsite & virtual programming
  • Generous vacation policy
  • Parental leave policies and family planning benefits
  • Catered lunches, fully-stocked kitchens with premium snacks/beverages

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available