About the role
Baseten powers inference infrastructure for leading AI companies including Cursor, Notion, Abridge, and Writer. The company recently closed a $1.5B Series F funding round led by Altimeter Capital, Conviction Partners, and Spark Capital. Baseten combines applied AI research, flexible infrastructure, and developer tooling to help frontier AI companies deploy cutting-edge models into production at scale.
As an inference performance engineer at Baseten, you'll solve some of the hardest optimization problems in AI infrastructure. The company's customers run some of the world's most demanding LLM workloads, and your work directly determines how fast those models execute and how efficiently Baseten serves them. You'll join the Engines, Performance, and Distributed (EPD) team in this full-time hybrid role based in San Francisco.
What you'll do
- Implement and productionize advanced inference techniques deep in runtime internals, including quantization, speculative decoding, KV-cache reuse, chunked prefill, LoRA, guided generation, and custom scheduling and routing algorithms.
- Profile and optimize inference end-to-end across the entire stack—from kernel launch overhead and memory layout through request scheduling, prefill/decode disaggregation, and cache-aware routing. Trace performance regressions across layers to identify root causes.
- Convert performance gains into cost savings by improving tokens per GPU-hour, raising utilization rates, and clearly communicating latency, throughput, and cost tradeoffs to customers and internal teams.
- Rapidly integrate new model architectures and hardware configurations, often shipping optimizations within days of new releases.
- Design and build benchmarking frameworks that measure real-world performance across different model architectures, batch sizes, sequence lengths, and hardware setups.
- Contribute to open-source inference engines like vLLM, SGLang, and TensorRT-LLM while partnering closely with model, infrastructure, and customer teams.
What you'll bring
- Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or a related field.
- Strong proficiency in general-purpose programming languages such as Python or C++.
- Solid understanding of LLM optimization techniques including quantization, speculative decoding, and continuous batching.
- Deep familiarity with ML frameworks, particularly PyTorch, TensorRT, or TensorRT-LLM.
- Demonstrated hands-on experience and genuine interest in large language models.
- Strong grasp of GPU architecture and how it impacts software performance.
Nice to have
- Previous contributions to vLLM, SGLang, TensorRT-LLM, or similar inference engines.
- Experience with large-scale distributed serving including autoscaling, load balancing, and multi-region or multi-cloud deployments.
- GPU kernel optimization experience using CUDA, Triton, CUTLASS, or related tools.
- Production experience with quantization techniques (FP8, FP4, AWQ, GPTQ) or speculative decoding implementations.
What they offer
- Competitive salary with meaningful equity participation.
- Medical, dental, and vision insurance with 100% employee coverage plus dependents (U.S. only).
- Flexible paid time off with company-wide Winter Break closure from Christmas Eve through New Year's Day.
- Paid parental leave and fertility and family-building support through Carrot.
- 401(k) plan with company facilitation (U.S. only).
Pay, location & hours
$180K–360K / year. Based in San Francisco (Hybrid).
About Baseten
baseten.co · 4 open roles in this building · Company page → · See it on the map