Production Services Specialist II
Job Description:
At Bank of America, we are guided by a common purpose to help make financial lives better through the power of every connection. We do this by driving Responsible Growth and delivering for our clients, teammates, communities and shareholders every day.
Being a Great Place to Work is core to how we drive Responsible Growth. This includes our commitment to being an inclusive workplace, attracting and developing exceptional talent, supporting our teammates’ physical, emotional, and financial wellness, recognizing and rewarding performance, and how we make an impact in the communities we serve.
Bank of America is committed to an in-office culture with specific requirements for office-based attendance and which allows for an appropriate level of flexibility for our teammates and businesses based on role-specific considerations.
At Bank of America, you can build a successful career with opportunities to learn, grow, and make an impact. Join us!
LOB Description:
The technology areas of focus for the Data Center and DMZ Network Technology Operations Specialist II includes our Global Data Center and DMZ infrastructure. Technology Operations Specialist II are expected to be well versed in numerous networking protocols, technologies and troubleshooting methodology, including the use of proactive and reactive tools.
LOB Responsibilities:
- Operational Support Tier 2 engineer for our Internal Data Center and DMZ infrastructure support team.
- Cisco ASR and ISR routers, Cisco Nexus and Catalyst switches, Arista DCS switches / CloudVision are the main platforms. Technical areas include but are not limited to internet routing, aggregation, distribution and access layer switch and routing.
- Lead production support triage efforts for network infrastructure incidents, manage bridge line troubleshooting and appropriate team engagement, engage in technical research and troubleshooting, and escalate to next level of leadership as needed. Identify service impact, interpret monitors, dashboards, and logs.
- Provide status updates and technical detail for awareness communications, ensure accuracy of all communications sent, and ensure any necessary follow-ups are scheduled.
- Identify possible production failure scenarios, vulnerabilities, and opportunities for improvement, and take ownership of escalation.
- Participate in the documentation of application flows, upstream/downstream impacts during outages, the customer experience in failure scenarios, contacts for various support needs and ensures appropriate documents and wikis are up to date and available for use during triage.
- Proactive network reviews including, Routine testing of disaster recovery scenarios, identification of vulnerabilities and opportunities for improvement in observability across the network stack.
- Mentorship of junior engineers and technical leadership within the team. Work with senior team members to validate impacts and communicate to all stakeholder’s technical status updates.
- Participate in the documentation of application flows, upstream/downstream impacts during outages, the customer experience in failure scenarios, contacts for various support needs and ensures appropriate runbooks and wikis are up to date and available for use during triage.
- Work ad-hoc reports and offline incidents at the direction of the senior team members or leadership.
- Promote and enforce production governance during triage/testing and fix efforts, exercises judgment within defined procedures and practices to determine appropriate action.
- Adhere to design standards and global design authority processes and procedures.
- Assemble professional documents based on existing templates and ability to provide accurate work descriptions with assumptions, and caveats.
Job Description:
This job is responsible for providing front-line support to end users, responding to issues related to incidents and problem management governance for multiple applications, and leading triage activities on all business impacting incidents. Key responsibilities include ensuring compliance with incident management and problem management policies and procedures, serving as a focal point for the customer, client, and associate experience, restoring complex production incidents under tight Service Level Agreements, and pursuing root cause and problem resolution follow ups.
- Incident Leadership (Command Control)
- Lead major incident bridge calls and take command of triage activities
- Own engagement strategy, ensuring the right teams are mobilized quickly
- Direct troubleshooting efforts across multiple network domains
- Make real-time decisions on escalation, prioritization, and recovery actions
- Maintain clear control of incident flow, ensuring focused and efficient resolution
- Technical Execution
- Drive coordinated troubleshooting across technologies including routing, switching, firewalls, load balancing, and network security
- Identify service impact and validate findings with technical teams
- Anticipate failure scenarios and guide mitigation strategies
- Communication Business Alignment
- Translate technical issues into clear business impact statements
- Provide accurate, timely updates to stakeholders and leadership
- Ensure consistency and clarity in all incident communications
- Maintain alignment between technical actions and business priorities
- Governance, Quality Continuous Improvement
- Ensure all incident records are complete, accurate, and meet enterprise standards
- Enforce adherence to incident management processes and controls
- Identify patterns, recurring issues, and systemic risks
- Drive follow-ups that improve network stability and prevent repeat incidents
- Maintain and enhance documentation, playbooks, and knowledge artifacts
Responsibilities:
- Leads production support triage efforts, manages bridge line troubleshooting, engages in technical research, and escalates issues to leadership as needed
- Ensures all impacts are accurately recorded and documented in the system of record, oversees that documents and wikis are updated and available for use during triage, and supports the documentation of application flows, upstream/downstream impacts during outages, the customer experience, and contacts for support needs
- Identifies and/or validates business impacts through interpretation of monitors, dashboards, and logs to communicate with leadership and vendors
- Manages activities to identify incident root cause, resolution, preventative actions, and change requests, and reports on incident data quality
- Promotes and enforces production governance during triage/testing and identifies production failure scenarios, vulnerabilities, and opportunities for improvement
- Serves as a subject matter expert for applications within a portfolio, leveraging extensive knowledge of application functionalities and application flows
- Assesses and prioritizes research requests, ad hoc reports, and offline incidents at the direction of senior team members and delegates work as needed to team members and peers
- The TRS operates at the center of incident response, leading high-severity network events where speed, clarity, and decisive leadership are essential. Acting as the single point of technical authority during incidents, this role directs cross-functional teams, determines escalation paths, and ensures all actions are aligned to business impact.
- This role requires the ability to lead under pressure, make decisions with incomplete data, and communicate clearly to both technical teams and senior leadership.
- This role operates within a 24x7 follow-the-sun Global Network Operations environment and requires flexibility to support continuous technical and operational coverage.
- Work schedules and shift patterns will be aligned to regional business and operational needs and may include weekends, public holidays.
- The role is expected to provide technical leadership coverage during assigned shift hours, lead or support major network incident triage and escalation when needed, and partner closely with peer leaders across regions to ensure effective technical handoff, restoration continuity, and sustained service stability.
Required Qualifications:
- Expert experience with Network technologies: TCP/IP, IPv4, Layer 2 protocols, Multicast, BGP EIGRP, QoS, UDP, OSPF, Carrier Circuits, Leased Line, Broadband, Direct Internet Access, Tunneling protocols (MACSEC, IPSEC, SSL/TLS, GRE), Routers, Switches, HSRP, ACL, VPN.
- Experience with troubleshooting complex networking problems.
- Experience with JIRA, Confluence, Agile framework & SCRUM ceremonies.
- Understand configuration management with tools such as Forward Networks and HPNA
- Experience using (both proactive and reactive) advanced tooling; Inclusive of but not limited to SevOne, Splunk, Netscout, Wireshark, NDC, HPNA, NNMI, OBM, IBM Watson, NSO, etc.
- General experience in Network Automation tools and processes
- Working knowledge of Python scripting and basic REST API / JSON-based data exchange; exposure to backend frameworks such as Django or Flask is a plus.
- Fundamental enterprise networking knowledge across routing, switching, wireless, and basic SD-WAN concepts, with familiarity using monitoring, alerting, or telemetry tools.
- Basic Linux environment troubleshooting skills and ability to follow established network designs and operational processes.
- Strong communication and team collaboration skills, with a willingness to continue building technical depth in automation and networking.
- Self-starter/self-directed, organized and detail oriented.
- Strong technical acumen and analytical skills
- Excellent client interfacing skills
- Strong verbal and written communication skills and ability to work with all levels of management.
- Experience aligning actions to business impact and service restoral.
- Demonstrates ownership: Is accountable and can hold others accountable (professionally)
- Experience operating with colleagues across different time zones with a flexible approach to working hours (ability to work varied hours) to successfully interact and communicate on a global level.
- This role requires Weekend work
Desired Qualifications:
Desired Qualifications:
- Experience in Networking-related disciplines within a design, implementation, or operations role.
- Relevant Industry certifications in Network Technologies.
- Cloud or SDN knowledge and experience
- Experience with SDN; Cisco ACI, VMware NSX, Arista CloudVision.
- Experience with SDWAN, preferred if on CloudGenix
- Experience with automation tools such as Python, Ansible, YAML, REST or Django.
- Experience working in an Agile environment.
- Experience of working within Financial Services (Insurance, Banking, Investment banking).
- Experience with other network technologies Firewall, Proxy/Threat Prevention, DDI, Load Balancing, and AAA.
Skills:
- Adaptability
- Analytical Thinking
- Influence
- Production Support
- Risk Management
- Automation
- Collaboration
- Innovative Thinking
- Result Orientation
- Solution Design
- Other
Shift:
1st shift (United States of America)Hours Per Week:
40