Senior Staff Machine Learning Engineer

ServiceNow · Santa Clara, CALIFORNIA, us

onsitefull-time6-10 years

posted 20h

About the team  The Security and Risk Engineering organization builds scalable, AI-powered security solutions that reduce risk and protect ServiceNow and its customers. We value AI-first thinking, clean architecture, intuitive experiences, and a culture of continuous learning.  This is a zero-to-one incubation. We’re building a new class of exposure analysis that ranks security work by exploitability—where an attacker could realistically get in—rather than raw severity. The architecture is evolving, and this role helps define what good looks like.  The role  As a Senior Staff Engineer, you own the architecture of an security harness with novel exploitability engine end to end, and you’re accountable for the decisions that shape everything downstream. You set technical direction, make the hard calls defensible, and multiply the engineers around you.  What you’ll own  The end-to-end architecture of the exploitability engine—from evidence ingestion and entity resolution, through the attack-path probability core and choke-point ranking, to the validation loop that keeps predictions honest.  The decisions that cascade through the system: calibrated probability versus ordinal rank, identity as a first-class graph edge, assume-breach seeding, and how the most critical assets are defined. These are model-shaping calls, not implementation details.  The probabilistic ranking core: edge-traversal probability, guided path search with hop and likelihood limits, correlated-control-failure modeling, and honest uncertainty bands.  The calibration and validation loop—canaries, purple-team and incident replay, calibration measured by zone and vector—that turns modeled weights into evidence rather than opinion.  Make-or-break metrics as first-class engineering targets, starting with entity-resolution accuracy and calibration quality.  The build-on strategy—extending the existing portfolio rather than rebuilding it, and knowing precisely what to reuse and what must be net-new.  What you’ll do  Lead zero-to-one work at production scale: turn an ambiguous, novel problem into a reliable system other teams build on, and set the bar where no precedent exists.  Drive technical direction across architecture, design, and code reviews, and raise the engineering bar across the incubation.  Mentor senior engineers and lead by influence, not title.  Partner with product, security R&D, SecOps to turn customer problems into architecture, and translate that architecture into decisions leaders can act on.  Establish AI safety, security, governance, and guardrails for agentic systems running in production.  What you bring  A track record of owning architecture across a large system or multiple teams, with deep experience operating production-quality software.  Hands-on depth in both agentic and LLM systems and probabilistic or ML-driven scoring—graph modeling, calibration, search and optimization, or risk and probability engineering.  Proven zero-to-one at scale: you’ve taken an ambiguous problem to a reliable production system that others depend on.  The judgment to make consequential architecture decisions under uncertainty, and make them defensible to engineers and executives alike.  Command of distributed systems, APIs, cloud-native development, and data or graph systems.  Expert-level Python, and/or Java, Go, or TypeScript.  Technical leadership and mentorship that moves teams through influence.  Applied interest in security problems—attack-path analysis, vulnerability management, identity security, threat intelligence, or detection and response—is strongly preferred.  Experience with AI-assisted development tools and coding agents such as Claude Code, Codex, Cursor, or Windsurf is a plus.  Experience with AI evaluation, safety, governance, o