Point your AI agent at freehire and let it find you a job.

Get the CLI →

Synopsys

NewBe an early applicant

Site Reliability, Staff - 18460

Posted Updated
Discussion

We Are

Synopsys is the leader in engineering solutions from silicon to systems, enabling customers to rapidly innovate AI-powered products. We deliver industry-leading silicon design, IP, simulation and analysis solutions, and design services. We partner closely with our customers across a wide range of industries to maximize their R&D capability and productivity, powering innovation today that ignites the ingenuity of tomorrow.

You Are

You have spent years keeping large-scale compute environments running, not just available but actually performant, under the kind of load that makes most infrastructure buckle. You understand that when an engineer's simulation job sits in queue for six hours instead of six minutes, that is not a scheduler problem, it is a planning problem, a tuning problem, or a resource allocation problem, and you are the person who figures out which one and fixes it before it becomes a pattern.

You think in systems, not tickets. When LSF or Slurm throws an error, you do not just restart the service, you trace it back to the workload spike, the misconfigured policy, or the storage bottleneck that triggered it. You have built enough automation to know that the script is the easy part, the hard part is designing it so it does not create three new problems when the environment shifts next quarter.

You are comfortable in a room with R&D engineers who need 10,000 cores by Tuesday and cloud architects who want to move half the farm to Azure, and you can translate between those conversations without losing the thread of what actually has to work. At Synopsys, you will work on HPC infrastructure that powers the chip design tools the world depends on, and the decisions you make will directly affect how fast our engineers can build the next generation of silicon.

What You'll Be Doing

  • Design, build, and optimize large-scale HPC compute farm platforms that support thousands of engineering workloads across global sites
  • Administer and tune IBM Spectrum LSF, Slurm, or equivalent workload schedulers to maximize resource utilization and minimize job queue times
  • Lead capacity planning and workload optimization efforts, translating engineering demand into infrastructure requirements and deployment timelines
  • Drive automation using Python and Shell scripting to eliminate manual toil, improve reliability, and accelerate incident response
  • Lead complex troubleshooting and root cause analysis for platform issues, working across storage, networking, LDAP, NFS, and scheduler layers
  • Collaborate with R&D, Cloud, Infrastructure, and Security teams on strategic initiatives including cloud-integrated HPC and AI/ML workload enablement
  • Mentor junior engineers and provide technical leadership across global teams, setting standards for operational excellence and engineering rigor
  • Participate in 24x5 support operations and lead major infrastructure projects from design through deployment

The Impact You Will Have

  • Enable faster product development cycles by ensuring high availability and performance of the compute infrastructure that powers Synopsys EDA tools
  • Maximize infrastructure utilization and license efficiency, directly reducing costs and improving engineering productivity across the company
  • Reduce incident response time and platform downtime through automation, monitoring, and proactive capacity management
  • Accelerate cloud adoption and modernization efforts, helping Synopsys scale compute resources dynamically to meet global engineering demand
  • Improve workload throughput and job completion times, giving engineers more iterations per day and faster feedback loops
  • Build operational resilience into the platform, ensuring that infrastructure scales reliably as the business grows
  • Mentor and elevate the technical capabilities of the global infrastructure engineering team, raising the bar for how we operate and support critical systems

What You'll Need

  • 8+ years of Linux/UNIX systems administration experience with deep expertise in performance tuning, troubleshooting, and large-scale operations
  • 5+ years of hands-on HPC or compute farm administration, including workload scheduling, resource management, and capacity planning
  • Strong expertise in IBM Spectrum LSF, Slurm, or equivalent schedulers, including policy configuration, job prioritization, and performance optimization
  • Advanced knowledge of LDAP, NFS, DNS, enterprise storage systems, and networking in the context of distributed compute environments
  • Proven experience with Python and Shell scripting for automation, monitoring, and infrastructure orchestration
  • Solid understanding of monitoring and observability tools such as Grafana, Prometheus, Elastic, or Splunk for proactive incident detection and analysis
  • Experience with EDA environments is a strong plus, as is familiarity with cloud-integrated HPC platforms on Azure or AWS, Kubernetes, Docker, Ansible, or Terraform

Who You Are

  • You take ownership of problems end to end, from the first alert to the post-incident review, and you do not consider something fixed until you understand why it broke
  • You are self-driven and proactive, the kind of engineer who sees a pattern in the logs and writes the script to prevent it before anyone files a ticket
  • You can explain a complex infrastructure tradeoff to a VP in two sentences without losing the technical nuance, and you can turn that same conversation into a design doc for your team
  • You are collaborative and team-oriented, comfortable working across time zones and disciplines to align on solutions that work for everyone
  • You have strong analytical and problem-solving instincts, you do not guess, you instrument, measure, and validate before you change production
  • You are focused on service excellence and customer impact, you measure success by whether engineering teams can do their jobs faster and more reliably because of the infrastructure you built

The Team You'll Be Part Of

The Compute Farm Infrastructure Engineering team operates and enhances Synopsys' global HPC and compute farm platforms. You will work with a distributed team responsible for providing reliable, scalable engineering compute services that support mission-critical EDA workloads. The team drives automation, modernization, and cloud adoption while maintaining high availability and operational excellence across all global sites.

Rewards and Benefits

We offer a comprehensive range of health, wellness, and financial benefits to cater to your needs. Our total rewards include both monetary and non-monetary offerings. Your recruiter will provide more details about the salary range and benefits during the hiring process.

Skills

See also

SRE jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available