About the role
Anthropic is an AI safety company building reliable, interpretable, and steerable AI systems that are safe and beneficial for society. The organization brings together researchers, engineers, policy experts, and business leaders in a rapidly growing collaborative environment focused on high-impact work. You'll join the Infrastructure team as a Staff+ Software Engineer specializing in Research Systems Engineering, where you'll design and operate the distributed systems that power the company's AI model training, serving, and security operations.
In this role, you'll partner closely with Anthropic's research teams to build the infrastructure that underpins their most critical work. You'll take on complex, multi-month infrastructure initiatives from initial scoping through production deployment, identifying and resolving the performance and scalability constraints that affect the organization's growth velocity. Your work will touch core systems that every other team at Anthropic relies on, requiring you to balance reliability, scalability, and security as these systems grow in usage and complexity.
What you'll do
- Independently scope, design, and lead large-scale infrastructure projects across multiple months, translating ambiguous requirements into production systems
- Establish deep partnerships with research teams to understand their evolving needs and deliver solutions that enable their work
- Lead architectural decisions that shape how other engineers and teams build on your systems
- Mentor engineers across the infrastructure organization and elevate the technical standards of the team
- Own the operational health, scalability, and security of the systems you build as they scale
- Drive improvements to infrastructure processes including incident response, postmortems, and on-call practices, ensuring the team learns from every operational event
What you'll bring
- Proven track record designing, building, and operating large-scale distributed systems or infrastructure at scale in production environments
- Demonstrated ability to independently scope and deliver complex, multi-month technical projects with ambiguous starting points
- Prior experience serving as a technical lead or mentor to other engineers
- Strong software engineering fundamentals and fluency in at least one programming language such as Python, Rust, Go, or Java
- Hands-on experience with modern cloud infrastructure including Kubernetes, infrastructure as code, AWS, or GCP
- Excellent written and verbal communication skills with a track record of building alignment across multiple teams and stakeholders
Nice to have
- 10 or more years of professional software engineering experience
- Background with machine learning infrastructure including GPUs, TPUs, or Trainium chips and their associated networking requirements like NCCL
- Low-level systems expertise such as Linux kernel tuning or eBPF
- Applied experience with security or privacy engineering best practices
What they offer
- Annual salary range of 320,000 to 485,000 USD
- Visa sponsorship available, with immigration legal support to facilitate successful sponsorship
- Flexible remote work with expectation of in-office presence at least 25% of the time at San Francisco, Seattle, or New York City locations
- Bachelor's degree or equivalent combination of education, training, and professional experience required
Pay, location & hours
Salary not listed. Fully remote, open to applicants in Friendly, Travel-Required, San Francisco, CA, Seattle, WA, New York City, NY.
About Anthropic
39 open roles in this building · Company page → · See it on the map