Sr Mgr, Reliability Engineering Mgmt

ServiceNow · Dublin, ie

onsitefull-time6-10 years

posted 1d

Team :   Our Site Reliability Engineering (SRE) team consists of highly skilled engineers responsible for maintaining and enhancing the reliability, scalability, and performance of the ServiceNow infrastructure. Our SRE’s are empowered to resolve technical issues across the entire technology stack, from hardware to applications. Additionally, they work to improve the platform's operability, aiming to reduce the number of incidents and minimize Mean Time to Recovery (MTTR). To achieve this, the team combines software development, networking, database, and systems engineering skills to tackle complex problems, striving to maintain our platform operating for our customers.  Role  As a Sr Manager, Site Reliability Engineering at ServiceNow, you’ll lead a team of SRE leaders, managers, and engineers focused on ensuring the reliability and availability of critical enterprise platforms and applications, while helping drive our cloud modernization journey through automation, operational excellence, resilience, and continuous improvement.  Do you:  have a technical background in roles like systems engineering or devops or site reliability engineering?  know operating systems in various levels of troubleshooting and diagnostics?  have experience with cloud technologies and hyperscalers such as AWS, GCP, or Azure, and a passion for driving cloud modernization and cloud-native transformation?  have low tolerance to repetitive tasks and automate your way through work?  have experience in leading a team of engineers and exposure to people management?  If you Answered 'yes' to these questions, we want to hear from you. Hit the Apply button and let's have a chat about the role and your skills and experiences.  As a Sr Manager of the SRE team your responsibilities will be:  Lead and develop a global team of SRE leaders, managers, and engineers, with accountability for talent development, performance, prioritization, succession planning, and execution.  Define and drive the SRE strategy and operating model across reliability, observability, automation, incident response, production readiness, and continuous improvement.  Own and evolve observability capabilities across metrics, logs, traces, alerting, SLI/SLOs, error budgets, and service health practices to improve detection, diagnosis, and overall production reliability.  Lead the strategy and execution for production-like staging environments, ensuring critical services and releases can be validated in environments that closely represent production.  Partner with Engineering and Release teams to strengthen release pipelines and confidence gates for Now releases, including smoke, integration, resiliency, performance, rollback, and production-readiness validation.  Drive a culture of eliminating repetitive operational work through automation, orchestration, self-healing, and AI-assisted operations, shifting reliability practices earlier in the software-development lifecycle.  Lead cloud modernization initiatives, modernizing legacy infrastructure, tooling, and operational practices toward cloud-native architectures and scalable engineering patterns.  Drive adoption and operational excellence across AWS, GCP, and Azure hyperscalers, including architecture, scaling, resiliency, availability, cost efficiency, and operational readiness.  Advance containerization and Kubernetes-based operating models, including workload resiliency, scalability, orchestration, deployment patterns, observability, and lifecycle management.  Partner with Product Engineering, Platform Engineering, Release Engineering, Security, and Infrastructure teams to improve the reliability and availability of critical enterprise platforms and services.  Establish reliability standards and measurable outcomes around SLIs/SLOs, error budgets, availability, MTTR, change failure rate, automation, and operational readiness.

Sr Mgr, Reliability Engineering Mgmt at ServiceNow — TalentDesi