Lead Site Reliability Engineer
Mastercard · Mexico City, Mexico
onsitefull-time6-10 years
posted 1d
Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Lead Site Reliability Engineer Overview: The role of Business Operations Organization is to be the production readiness steward for Mastercard products. As a Business Operations we are responsible for ensuring that our platform is stable and healthy. We break down barriers to run our products by fostering developer run ownership and empowering developers to build resilient products. We support our developers during the application build phase in software run principals that includes operational design, automation, capacity planning, monitoring that leads to fault-tolerant, scalable products. We see the big picture and help create and enforce operations standards while facilitating an agile and learning culture. We accomplish this transformation through supporting daily operations with a hyper focus on triage and then root cause by understanding the business impact of our products. The goal of every biz ops team is to shift left to be more proactive and upfront in the development process, and to proactively manage production and change activities to maximize customer experience and increase the overall value of supported applications. Biz Ops teams also focus on risk management by tying all our activities together with an overarching responsibility for compliance and risk mitigation across all our environments. A biz ops focus is also on streamlining and standardizing traditional application specific support activities and centralizing points of interaction for both internal and external partners by communicating effectively with all key stakeholders. Ultimately, the role of biz ops is to align Product and Customer Focused priorities with Operational needs. We regularly review our run state not only from an internal perspective but also understanding and providing the feedback loop to our development partners on how we can improve the customer experience of our applications. Key Responsibilities • Lead and own the full lifecycle of services—from architecture and design through deployment, operations, and continuous optimization—ensuring scalability, reliability, and alignment with business objectives. • Analyze platform-level ITSM performance and proactively establish feedback loops with engineering teams, influencing roadmap prioritization to address systemic gaps and improve resiliency. • Define and drive production readiness standards, including operational design reviews, capacity planning, and launch governance, ensuring services meet reliability and scalability benchmarks before go-live. • Define and evolve monitoring frameworks for availability, latency, and system health, leveraging metrics and telemetry to proactively prevent incidents and improve service performance. • Champion automation-first principles to scale systems efficiently, reducing manual toil while improving deployment velocity and overall system reliability. • Lead the design and governance of CI/CD pipelines, implementing robust validation, operational gates, and best practices to drive consistency, quality, and speed across environments. • Drive best-in-class incident response practices, including rapid mitigation, stakeholder communication, and blameless postmortems, ensuring continuous improvement and resilience. • Take a holistic, system-wide approach during critical incidents, connect • Collaborate effectively across distributed, global teams, ensuring alignment, continuity, and high performance across time zones and technology