IT Site Reliability & Performance Engineer
The Site Reliability and Performance Engineer designs implement and maintains monitoring and observability solutions and supports the performance and scalability of IT infrastructure and applications. This role partners closely with infrastructure, operations, development, and security teams to help ensure IT services and systems remain available, reliable, and high performing.
This position is based at our Corporate Headquarters in Wood Dale, IL, with a planned relocation to the Merchandise Mart (Chicago) in early 2027.
What you will be responsible for\:
- Develop and maintain the organization’s monitoring and observability strategy, standards, and best practices.
- Design, deploy, and manage monitoring platforms and related tools.
- Collect, analyze, and visualize metrics, logs, traces, and events to deliver end-to-end observability across infrastructure and applications.
- Build and maintain dashboards, alerts, and reports for infrastructure, operations, development, and security teams.
- Improve performance, reliability, scalability, and operational maturity of systems and applications.
- Capacity Planning\: Plan for future growth and ensure that systems can handle increased loads without performance degradation.
- Partner with development teams to improve services through rigorous testing and release procedures
- Troubleshoot and resolve issues involving app performance, monitoring and observability tools.
- Provide guidance and support to IT teams on effective use of monitoring and observability solutions and standardize their usage.
- Research and assess emerging technologies and trends in app performance, monitoring and observability.