AI Engineer, Internship - Summer 2026 - Applications Open Now

Postman · Berkeley, California, United States; San Francisco, California, United States

onsiteinternshipFresher / Intern

posted 17 Aug

Sign in to apply

<div class="content-intro"><h2><strong>Who Are We?</strong></h2> <p>Postman is the world’s leading API platform, used by more than 45 million+ developers and 500,000 organizations, including 98% of the Fortune 500. Postman is helping developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster.</p> <p>The company is headquartered in San Francisco and has offices in Boston, New York, Austin, Tokyo, London, and Bangalore - where Postman was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners. Learn more at postman.com or connect with Postman on X via @getpostman.</p> <p>P.S: We highly recommend reading <a href="https://api-first-world.com/">The "API-First World" graphic novel</a> to understand the bigger picture and our vision at Postman.</p></div><h2><strong>The Opportunity (Summer 2026 AI Internship - Applications Open Now)</strong></h2> <p>We're seeking an&nbsp;<strong>AI Engineer Intern </strong>to work alongside our AI team on large-scale AI and Agentic&nbsp; systems from data pipeline to production deployment. This role is scoped for someone with foundational experience who wants to deepen it: you'll own discrete pieces of real systems under the mentorship of senior engineers, not shadow work or isolated coursework-style projects.</p> <h2><strong>What You'll Do</strong></h2> <p>You’ll work directly with the AI team, taking responsibility for well-scoped pieces of real systems, with mentorship from senior engineers.</p> <p><strong>Benchmarks &amp; Evaluation</strong></p> <ul> <li>Contribute to <a href="https://blog.postman.com/apiflow-bench/">APIFlow-Bench</a>, our open-source benchmark for real API-development work: design and review benchmark tasks and their mock API environments, extend the evaluation harness and task-generation pipeline in Python, and help maintain the public multi-model leaderboard with statistical confidence intervals.</li> <li>Help build a new action-level AI safety benchmark: instead of grading what a model says, it scores what an agent actually does inside a simulated enterprise API environment. You’ll work on scenario design, threat modeling (prompt injection, data exfiltration, permission overreach), and auditable evaluation design.</li> </ul> <p><strong>Model Training &amp; Efficiency</strong></p> <ul> <li>Fine-tune open-weight models for tool calling and agentic tasks (SFT, distillation, and RL) using PyTorch and the open-source training ecosystem, on both managed training platforms and self-managed cloud GPUs.</li> <li>Design and run experiments with rigor: evaluate every training run on our benchmarks, support ablation studies and error analysis, track experiments, and report results honestly, including cost.</li> <li>Evaluate ultra-low-bit quantized models for on-device use: extend our quantized vs. full-precision benchmark comparisons and analyze where and why they diverge.</li> </ul> <p><strong>Agent Systems &amp; Engineering Practice</strong></p> <ul> <li>Help build the next generation of Postman’s in-product AI agent (Agent Mode): a deliberately minimal agent architecture that calls LLM APIs directly (tool loops, multi-step execution, checkpointing), primarily in TypeScript. No prior TypeScript is required; strong Python fundamentals transfer quickly.</li> <li>Read the source code of open-source agent harnesses and turn what you learn into design specs and prototypes.</li> <li>Document experiments, design decisions, and runbooks so your work is legible to the next person; flag safety, fairness, or privacy concerns you observe in model or agent behavior.</li> </ul> <h2><strong>About You</strong></h2> <ul> <li>Currently pursuing a <strong>BS, MS, or PhD </strong>in <strong>Computer Science, Data Science, or a related quantitative field.</strong></li> <li>Hands-on experience tr