Senior Systems Engineer
Key Responsibilities
System Installation & Configuration – Software Administration:
- Monitors and maintains operating systems to ensure effective installation and performance.
Supports the deployment, maintenance, and operation of internal applications as needed.
- Performs regular administration and conducts standard performance trend analyses and manages server capacity to ensure service performance meets standards.
- Leverages working knowledge of application monitoring tools to ensure efficiency.
System Installation & Configuration – Installation and Configuration:
- Installs and configures servers, cloud infrastructure, and all software and environments.
- Performs system configurations and backups independently to ensure optimization of server infrastructure.
- Supports hardware maintenance, auditing, installation, and provisioning as necessary.
- Collaborates with internal technical experts and third-party vendors to resolve integration challenges.
Service Lifecycle Management – Security Maintenance:
- Follows existing procedures to provide assurance that compute and storage devices are secure.
- Maintains privileged accounts/secrets integrity of systems and compute and file system security for the compute and storage environment.
- Monitors and evaluates high-level service and infrastructure dashboards and takes action to address identified anomalies.
- Implements monthly, quarterly, or hotfix patches to address security vulnerabilities or bugs across the technical stack.
Service Lifecycle Management – System & Security Improvements:
- Deploys enhancements to improve the performance, reliability, and security of systems and environments.
Incident Management & Support – Incident Management:
- Supports the end-to-end incident management lifecycle to ensure systems are stable, secure, and performing accurately.
- Collates incident-based data for team metrics and key performance indicators (KPIs) by assisting with system and network incidents to identify patterns, root causes, and solutions.
- Participates in incident review meetings to provide feedback for operational performance and solution implementation.
- Partners with third party vendors and cross-functional teams (e.g., Development, Cloud Engineering, Product Engineering, other IT teams) to drive collaboration for implementation and/or resolution for high-severity incidents, risks, or migrations.
- Investigates standard system issues and triages high-severity incidents by implementing Corrective and Preventative Action plans (CAPA) as instructed.
Incident Management & Support – Escalation Cases:
- Coordinates escalated support cases by collaborating with internal technical teams and third party vendors to drive issue resolution for a wide range of production environment problems (e.g., immense growth, scaling, leveraging the cloud, extremely high performance, high availability requirements).
Incident Management & Support – Technical Support:
- Proactively monitors the production environment by checking system error logs, monitoring ticket queues, and consulting with other teams involved in maintaining the environments.
- Adheres to team schedule to contribute to ongoing technical support and service objectives.
- Resolves complex, critical customer system issues and implements and documents technical solutions.
Incident Management & Support – Backups and Disaster Recovery:
- Leverages working knowledge of systems to perform backup, restore, and disaster recovery processes.
- Participates in disaster recovery drills to ensure preparedness and compliance.
- Implements disaster recovery solutions to ensure preparedness and regulatory compliance.
Communication & Documentation – Technical Communication:
- Communicates technical information to both technical and nontechnical personnel.
- Serves as a technical liaison and provides domain-specific expertise to cross-organization projects, programs, and activities.
Communication & Documentation – Documentation & Reporting:
- Maintains documentation on ticket updates, code contributions, infrastructure, configurations, processes, and procedures (e.g., Disaster Recovery plans, Standard Operating Procedures, Corrective and Preventative Action Plans).
- Generates weekly and monthly reports on system performance and incident progress to support operational and management outcomes, and builds an awareness of the business impacts.
- Researches, proofs, and authors technical documentation in the area of standards and best practices for internal use.
Additional Responsibilities (as needed)
Cloud Infrastructure Support:
- Collaborates with DevOps and Site Reliability Engineer (SRE) teams to provide support for large-scale infrastructure.
- Implements continuous integration and continuous deployment (CI/CD) pipelines independently.
- Performs patching and version upgrades independently to support cloud infrastructure.
Automation:
- Supports Workload Automation tools through design support, administration, and optimization efforts.
- Troubleshoots issues with automation tools, agents, and other connectivity to 3rd party applications.
- Maintains and supports cloud technologies.
Core Responsibilities
Planning & Execution:
- Independently manages work, monitoring timelines and deliverables to ensure projects or initiatives stay on track and meet requirements.
- Proactively prioritizes work and adapts to resource or timeline shifts, suggesting adjustments to maintain project efficiency.
Collaboration & Partnership:
- Collaborates across teams to align on expectations and achieve shared objectives.
- Builds and maintains a comprehensive understanding of business, stakeholder, and/or customer needs to build and support effective partnerships.
- Actively listens to diverse perspectives and asks questions to ensure understanding of others.
Problem Solving:
- Independently identifies and addresses standard and non-standard issues in accordance with standard practices, escalating more complex issues as appropriate.
- Analyzes data and/or information from multiple sources to troubleshoot standard and non-standard errors.
- Contributes to knowledge sharing and best practices.
Continuous Learning:
- Embraces continuous learning by actively seeking to build knowledge and new skills and/or tools and staying current with industry trends and best practices.
- Seeks out and leverages feedback and training to improve skills.
- Contributes to a culture of continuous learning and knowledge sharing with team members.
Continuous Improvement:
- -Develops ideas and recommends updates to increase the efficiency and effectiveness of processes, protocols, and workflows within a team.
- Seeks input from team members on alternative approaches and methods for improving work.