Sr. Staff Software Engineer – SRE, Release & Test Platforms

ServiceNow · Dublin, ie

onsitefull-time6-10 years

posted 1d

Join us to build the next generation of cloud-native reliability, release, and test platforms that enable engineering excellence, developer productivity, and high-confidence ServiceNow releases through automation, observability, and AI-driven operations.  What you will do in this role:  Build and operate cloud-native engineering platforms for software validation, release qualification, and operational readiness.  Design production-like release and test environments that improve release confidence and deployment readiness.  Develop automated quality gates to assess release health, operational risk, and production readiness.  Integrate automated testing, observability, reliability signals, and deployment intelligence into CI/CD pipelines.  Build reusable test frameworks, self-service environments, test data, mock services, and developer productivity tooling.  Advance shift-left engineering through automated validation, continuous verification, and quality gates.  Automate failure detection, policy validation, deployment verification, security checks, and reliability assessments.  Lead Kubernetes-based platform evolution for scalable test infrastructure, release automation, and developer self-service.  Resolve recurring infrastructure issues through sustainable software, systems, and networking solutions.  Partner with engineering teams on design reviews, architecture standards, and automation-first reliability practices.  To be successful in this role you have:  12+ years of experience in software, systems, platform, or reliability engineering.  Deep Kubernetes expertise across architecture, operations, networking, storage, security, autoscaling, and multi-cluster environments.  Experience building and operating large-scale Kubernetes platforms for cloud-native, mission-critical services.  Experience integrating Kubernetes with CI/CD, GitOps, automated testing, and deployment validation.  Experience designing cloud-native platforms for ephemeral environments, release qualifications, and automated validation.  Proven ability to lead engineering excellence across developer productivity, platform engineering, release confidence, and modernization.  Experience with progressive delivery, including canary releases, feature flags, automated rollback, and deployment verification.  Experience with chaos engineering, resilience validation, disaster recovery, and reliability assessments.  Expertise designing, authoring, testing, and debugging code in a team setting using languages such as Python, Go, Java, or Ruby.  Experience using AI-assisted engineering for intelligent testing, release risk analysis, incident diagnostics, and operational automation.  Strong coding, observability, SLO, and cross-team collaboration skills to improve reliability, performance, and engineering standards.    Good to have:  Expertise in observability and monitoring applications, services, and networks at scale.  Experience with DevOps automation, CI/CD pipelines, and agile methodologies, including GitLab CI/CD or similar tools.  Experience building enterprise-scale test automation frameworks such as Playwright, Selenium, Cypress, REST Assured, PyTest, JUnit/TestNG, or equivalent technologies.  Experience with test orchestration, test impact analysis, flaky test detection, parallel execution, and intelligent regression testing.  Experience with service virtualization, contract testing, synthetic testing, and test data management.  Experience building engineering platforms that support developer self-service and release engineering.  Experience with infrastructure configuration management tools such as Ansible.  Expertise with Kubernetes ecosystem technologies such as Helm, Argo CD, Argo Workflows, Kustomize, Istio/Linkerd, Gateway API/Ingress, Prometheus, OpenTelemetry, and container runtimes.