TopWeb3JobsTopWeb3Jobs
Protocol & Infra district · Plot OH-A2-16

Technical Lead - GPU Infrastructure (100% Remote - Worldwide)

Tether ✓ Verified · Remote (job) · Full-time
Salary not listed

About the role

Tether is a blockchain infrastructure company that powers global financial innovation through reserve-backed stablecoins, energy solutions, AI infrastructure, and educational technology. Operating across more than 180 countries, Tether serves hundreds of millions of users and processes trillions of dollars in transactions annually. The company has established itself as a trusted bridge between traditional finance and decentralized systems, enabling seamless digital asset transfers across blockchain networks.

Tether Data division is at the forefront of AI infrastructure development, building the next generation of GPU compute platforms that reduce operational costs while expanding access to machine learning capabilities worldwide. The Data team operates Cosmic AC, a sophisticated orchestration system that manages GPU resources, inference endpoints, and computational workloads at scale across global infrastructure.

You will join as Technical Lead for GPU Infrastructure in a hands-on leadership role overseeing the evolution of Cosmic AC from a managed-cluster platform to a full bare-metal GPU stack. You will lead approximately twelve distributed engineers across backend, frontend, DevOps, QA and documentation roles spanning Europe and India. This is a fixed-scope role with clear delivery targets in the first six months, balancing direct technical contributions with team leadership and vendor partnerships.

What you'll do

  • Design and maintain platform architecture end to end, from high-level strategy through detailed implementation, managing proposals and design reviews with a live baseline architecture document
  • Lead a geographically distributed engineering team, providing technical direction, code and design review, release governance, individual performance feedback, and career development guidance
  • Build and operate a managed Slurm scheduling layer for research workloads, handling controller setup, accounting systems, partition configuration, node onboarding, CUDA driver management, health detection, and autohealing procedures
  • Architect Kubernetes control plane deployment on bare-metal infrastructure, including cluster bootstrap, NVIDIA GPU integration, virtual machine-based GPU isolation using KubeVirt and VFIO, and day-two operations like upgrades and recovery
  • Design managed inference serving architecture supporting multi-GPU and multi-node parallel processing, autoscaling policies, request routing, and endpoint reliability for production workloads
  • Establish observability systems across all layers including metrics, logging, alerting, SLOs, incident response processes, and sustainable on-call rotation practices
  • Serve as primary technical contact with infrastructure partners and vendors, translating requirements into specifications, managing escalations, and informing capacity planning decisions

What you'll bring

  • Eight or more years of hands-on infrastructure engineering experience, with at least three years leading teams that build and operate platforms depended on by other engineering organizations
  • Demonstrated production experience running Slurm at scale: direct experience with slurmctld and slurmdbd, partition and QoS configuration, accounting systems, job control scripting, node health monitoring, and managing upgrades with live workloads
  • Deep knowledge of Kubernetes architecture, cluster operations, and the NVIDIA GPU Operator ecosystem
  • Strong capability with backend technologies including Node.js, or demonstrated ability to lead teams using these tools
  • Excellent written and verbal English communication skills, essential for coordinating across distributed teams and external partners
  • Bachelor's or Master's degree in computer science or engineering, or equivalent professional infrastructure engineering experience
  • Comfort with infrastructure-as-code approaches and vendor relationship management

Nice to have

  • Experience with confidential computing technologies and secure workload isolation
  • Familiarity with KubeVirt, VFIO, or other GPU virtualization approaches
  • Knowledge of AI and machine learning workflow requirements
  • Previous experience in blockchain or fintech infrastructure environments

Pay, location & hours

Salary not listed. Fully remote, open to applicants in job.

About Tether

19 open roles in this building · Company page → · See it on the map

Apply ↗

More roles to explore

Salary not listed
Tether
Apply ↗

☆ Save this job

We'll e-mail you this role so you can come back to it. No account needed.

Report this job

Reports go to the TopWeb3Jobs team. Scam reports are checked first.