About the role
Baseten provides inference infrastructure that powers production AI systems for leading companies including Cursor, Notion, and Abridge. The platform combines applied AI research, flexible GPU infrastructure, and developer tools to help frontier AI companies deploy models at scale. Recently raising a Series F at $1.5B valuation, Baseten is expanding rapidly and seeking technical leadership to grow its performance engineering organization.
About Baseten
Baseten operates at the intersection of AI infrastructure and systems performance, serving companies that push the boundaries of what's possible with large language models and agentic AI. The company has built a runtime platform specifically designed to maximize GPU utilization and inference throughput. With backing from top-tier investors including Altimeter Capital and Spark Capital, Baseten is positioned as a critical piece of infrastructure for the AI wave sweeping through enterprise and consumer software.
This role offers the opportunity to lead technical strategy and team growth during a period of explosive expansion in AI inference. You'll work on problems that directly determine how quickly and cost-effectively the industry's most demanding models run in production.
What you'll do
- Build and mentor a growing team of inference performance engineers, conducting regular one-on-ones, providing structured feedback, and developing career paths as the runtime team scales
- Hire top GPU and systems engineering talent, establishing a collaborative culture that attracts world-class engineers
- Own the technical roadmap for runtime performance, balancing immediate customer needs with long-term platform investments and new model architecture releases
- Engage hands-on with the technical work: review architecture decisions, guide profiling and optimization work, and help engineers reason from first principles about GPU memory and compute constraints
- Lead productionization efforts for techniques like quantization, speculative decoding, KV-cache management, chunked prefill and custom scheduling algorithms
- Translate performance improvements into measurable business metrics including tokens per GPU-hour, GPU utilization, latency and total cost of inference
- Work cross-functionally with Infrastructure, Kernels, Platform and customer-facing teams to coordinate launches and align priorities
What you'll bring
- Bachelor's degree or higher in Computer Science, Engineering, Mathematics or related field
- Proven experience managing and hiring engineering teams, including performance reviews and career development
- Technical background leading or deeply supporting GPU optimization efforts in training, inference or recommendation systems at scale
- Strong foundation in GPU architecture, performance characteristics and optimization tradeoffs
- Hands-on familiarity with machine learning frameworks such as PyTorch, TensorRT or TensorRT-LLM
- Track record of owning complex technical roadmaps and shipping sophisticated projects with distributed teams
- Strong communication skills across technical and non-technical audiences
Nice to have
- Direct experience with inference engines like vLLM, SGLang or TensorRT-LLM
- Production experience optimizing LLMs using quantization, speculative decoding or continuous batching
- Background writing GPU kernels in CUDA, Triton or CUTLASS
- Experience scaling engineering teams through rapid startup growth
- Prior hands-on work as a performance or systems engineer
What they offer
- Competitive compensation package with meaningful equity participation
- Comprehensive health coverage including medical, dental and vision insurance for employees and dependents
- Flexible paid time off plus company-wide winter break closure from Christmas Eve through New Year's Day
- Paid parental leave and fertility support stipend via Carrot
- Company-facilitated 401(k) retirement plan
Pay, location & hours
$240K–270K / year. Based in San Francisco (Hybrid).
About Baseten
baseten.co · 4 open roles in this building · Company page → · See it on the map