Lead Data Engineer
Mastercard · Pune, India
onsitefull-time6-10 years
posted 16h
Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Lead Data Engineer Who is Mastercard? Mastercard is a global technology company in the payments industry. Our mission is to connect and power an inclusive, digital economy that benefits everyone, everywhere by making transactions safe, simple, smart, and accessible. Using secure data and networks, partnerships, and passion, our innovations help individuals, financial institutions, governments, and businesses realize their greatest potential. The Mastercard Services organization is a key differentiator, delivering cutting-edge solutions used by some of the world’s largest organizations to make critical business decisions. Focused on innovation and scale, Services provides data-driven capabilities across consulting, analytics, experimentation, and risk management. Role Overview Data Platform & Orchestration is seeking a Lead Data Engineer to design and build next-generation, cloud-native data platforms supporting Mastercard’s global data ecosystem. In this role, you will lead the development of scalable batch and real-time data pipelines, enabling efficient data processing across Data Lakes and Data Warehouses. You will work at the intersection of data engineering, cloud platforms, and distributed systems, contributing to high-impact initiatives and driving engineering excellence. This role is ideal for someone who thrives in a fast-paced, collaborative environment, enjoys solving complex data challenges, and is passionate about building resilient, high-performance systems at scale. Key Responsibilities • Design and build scalable batch and real-time data pipelines using Spark, Kafka, and (preferred) Apache Flink • Develop robust ETL/ELT frameworks for structured and unstructured data • Build and optimize data ingestion and transformation pipelines for Data Lakes and Data Warehouses • Implement stream processing solutions for near real-time use cases • Ensure data quality, lineage, observability, and governance across pipelines • Optimize data jobs for performance, scalability, and cost efficiency • Design and operate cloud-native data platforms on AWS, Azure, or GCP • Leverage managed services such as S3/ADLS/GCS, EMR/Databricks, BigQuery/Redshift/Snowflake • Implement Infrastructure as Code (Terraform, CloudFormation, or equivalent) • Ensure high availability, fault tolerance, and disaster recovery • Drive cost optimization strategies for large-scale data workloads • Implement secure data access controls aligned with enterprise standards Platform & Engineering Excellence • Build reusable data frameworks, libraries, and pipeline templates • Drive adoption of CI/CD, automated testing, and observability • Develop and enhance developer tooling and platform capabilities • Contribute to cloud-agnostic platform architecture and automation Technical Leadership & Collaboration • Provide technical leadership, mentorship, and design guidance • Conduct code reviews, architecture reviews, and best practice enforcement • Collaborate with architects, product owners, and cross-functional teams • Act as a Subject Matter Expert (SME) for data platform initiatives • Promote engineering excellence through documentation, design standards, and innovation • Work effectively across globally distributed teams Required Skills & Qualifications • Strong proficiency in Object-Oriented Programming and Design (OOP/OOAD) Java (JDK 8+); Python and/or Go is a plus • Experience building data services and distributed systems • Strong understa