Engineer III - Data + ML Platform (Hybrid)
CrowdStrike · India - Bangalore
onsitefull-time3-6 years
posted 19h
As a global leader in cybersecurity, CrowdStrike protects the people, processes and technologies that drive modern organizations. Since 2011, our mission hasn’t changed — we’re here to stop breaches, and we’ve redefined modern security with the world’s most advanced AI-native platform. We work on large scale distributed systems, processing almost 3 trillion events per day and this traffic is growing daily. Our customers span all industries, and they count on CrowdStrike to keep their businesses running, their communities safe and their lives moving forward. We're proud to work for a mission-driven company leveraging AI to transform the way we work. CrowdStrikers drive their careers through flexibility and autonomy while also being expected to contribute to a culture of responsible AI adoption, experimentation, and innovation. We use an AI-first mindset as a force multiplier to proactively and continuously accelerate execution, build expertise, uncover insights, and solve complex problems. We’re always looking to add talented CrowdStrikers to the team who have limitless passion, a relentless focus on innovation and a fanatical commitment to our customers, our community and each other. Ready to join a mission that matters? The future of cybersecurity starts with you. About the Role: We're seeking a Sr. Engineer - ML Platform to maintain and optimize CrowdStrike's mission-critical ML infrastructure. You'll diagnose complex distributed systems issues and ensure platform reliability for infrastructure processing billions of events daily. What You'll Do: Platform Reliability & Debugging: Diagnose and resolve issues across Ray, Spark, Airflow, MLflow, JupyterHub, Kubeflow, and SLURM Perform root cause analysis on production incidents affecting training and inference pipelines Debug performance bottlenecks, resource contention, memory leaks, and scheduling conflicts Develop debugging tools and diagnostic frameworks System Optimization & Performance: Profile and optimize Ray clusters and Spark jobs on K8s and Cloud (EMR/Dataproc) Troubleshoot JupyterHub spawner issues, kernel crashes, and resource allocation Optimize SLURM job scheduling, GPU allocation, and HPC cluster utilization Infrastructure & Monitoring: Build observability solutions and automated health checks Develop runbooks, alerting workflows, and incident response procedures Maintain platform stability metrics (SLAs, error rates, latency) Collaboration: Partner with ML and ML Platform engineers to resolve workflow issues Conduct post-mortems and mentor on debugging techniques What You'll Need: 5+ years in distributed systems engineering 5+ years debugging ML platforms in production Deep expertise in 3+ one of: Ray, Spark, JupyterHub, SLURM, K8 Performance profiling, optimization, and capacity planning Proven experience utilizing AI technologies to enhance decision-making, streamline workflows and processes, improve efficiency and drive business outcomes. Technical Skills (Expertise in at least one): Distributed ML: Ray, Spark, SLURM, Jupyter Ecosystem (debugging failures, performance tuning) ML Platforms: Airflow, MLflow, JupyterHub (troubleshooting core components) Infrastructure: Kubernetes, Docker, AWS/GCP/Azure/OCI Observability: Profiling tools, distributed tracing, Prometheus, Grafana, log aggregation Programming: Expert Python debugging, multi-language proficiency, Linux/Unix What Sets You Apart: Open-source ML infrastructure contributions Experience with high-throughput inference systems and reducing MTTR Published debugging guides or tools Chaos engineering and GPU/CUDA debugging experience On-call and incident management experience #LI-DP1 Benefits of Working at CrowdStrike: Market leader in compensation and equity awards Comprehensive physical and mental wellness programs Competitive vacation and holidays for recharge Paid parental and adoption leaves Professional development opportunities for all employees regardless of level or role Employee Network