Senior Applied Scientist / Engineer, Training & Inference
Adobe · San Jose
onsitefull-time6-10 years
posted 1d
Sign in to applyThe Opportunity Adobe Applied Science & Machine Learning (ASML) is seeking a Senior Applied Scientist / Engineer, Training & Inference to play a critical role in closing the gap between research and production for Adobe's next-generation video and image foundation models. In this role, you will serve as a technical owner for the training-to-deployment pipeline for our video and multimodal generation models. Rather than focusing solely on model research or systems infrastructure in isolation, you will bridge both — bringing the hands-on training expertise and the inference and deployment depth needed to take large generative models from the research cluster to reliable, performant, and cost-efficient production. This role is ideal for those who excels at the full arc of model development — distributed training at scale, inference optimization, and the practical engineering required to deploy and operate models reliably in production. Job Responsibilities Training & Inference Ownership. Own key components of the training-to-deployment pipeline — from distributed training execution through inference optimization, serving, and production handoff — ensuring models are delivered reliably, performantly, and cost-efficiently. Large-Scale Distributed Training. Implement and operate distributed training strategies including PyTorch FSDP, Tensor Parallelism, and Pipeline Parallelism across multi-node GPU environments, ensuring correctness, stability, and scalability for large video and multimodal models. Inference & Serving. Design and optimize inference and serving systems for large generative models, with a focus on latency, throughput, and cost across deployment targets. Research-to-Production Bridge. Reduce the gap between trained model checkpoints and reliable production deployments — owning the practical work of hardening, validating, and operationalizing models at scale. Performance & Cost-Aware Engineering. Identify and address inefficiencies across the training and inference stack — memory, communication, scheduling, and execution orchestration — with a clear focus on GPU efficiency and cost targets. Collaboration with Research & Engineering Teams. Partner closely with applied researchers, ML engineers, and infrastructure teams to align training and inference systems with model architecture needs and product delivery timelines. What You'll Need to Succeed Education: Master's or PhD in Computer Science, Electrical Engineering, AI/ML, or a related field, or equivalent practical experience. Distributed Training Expertise: Hands-on experience with large-scale distributed training using PyTorch (FSDP, Tensor Parallelism, Pipeline Parallelism) across multi-node GPU environments. Inference & Deployment Experience: Proven experience optimizing and deploying large generative models for production — including serving infrastructure, latency/throughput tuning, and cost-aware deployment. Strong Systems & Engineering Skills: Proficiency in Python and PyTorch, with experience working in large shared codebases and contributing to production-critical ML systems. Research-to-Production Execution: Demonstrated ability to take models from training through deployment, navigating the practical engineering challenges of reliability, reproducibility, and operational scale. Senior-Level Ownership: Demonstrated ability to independently own end-to-end technical areas, drive cross-team execution, and deliver high-quality systems on which product teams depend. Preferred Experience Experience training and deploying video, image, or multimodal generative models (e.g., diffusion models, flow matching, video generation). Familiarity with inference serving frameworks such as TensorRT, vLLM, or equivalent. Experience with performance profiling and optimization for both training and inference workloads. Track record of shipping generative AI models to production at scale. Prior work in an applied research environment bridging ML and systems engineering. About Ado