TopWeb3JobsTopWeb3Jobs
Tech district · Plot OH-A2-16

Technical Lead - GPU Infrastructure

Tether · Remote (job) · Full-time
⌘ Tech🌐 Remote
Salary not listed

About the role

Tether is a global fintech pioneer behind USDT, the world's most trusted stablecoin, used by hundreds of millions of people across exchanges, wallets, and payment systems. The company operates across multiple divisions—from blockchain finance to sustainable energy solutions to cutting-edge data infrastructure—all centered on transparency and global accessibility. Tether Data, one of these divisions, is building the next generation of GPU compute and AI infrastructure.

You'll join the Data team as Technical Lead for Cosmic AC, Tether's GPU compute platform. Cosmic AC orchestrates containers, managed inference endpoints and platform observability on Kubernetes today. Now it's expanding to own the full bare-metal GPU infrastructure stack: a managed Slurm scheduling layer for internal research and training teams, followed by a custom Kubernetes control plane for inference workloads. You'll own the architecture end to end, lead a distributed team of roughly twelve engineers across backend (Node.js), frontend (React), DevOps, QA and documentation spanning Europe and India, and be the primary technical contact with infrastructure partners. This is hands-on leadership with a clear six-month delivery window—not pure research, not just Kubernetes SRE work, and not management-only.

Your responsibilities span

  • Designing and owning platform architecture: producing proposals, high-level and low-level designs, driving them through review and keeping them current as the baseline
  • Leading and line-managing your distributed team: setting engineering standards, conducting code and design review, defining release gates, running one-to-ones, and providing growth and performance feedback
  • Building a managed Slurm service: designing controller and accounting, managing partitions and login nodes, handling driver and CUDA baselines, detecting stalled jobs and node health issues, implementing drain and autohealing
  • Managing Kubernetes on bare metal: cluster bootstrap and lifecycle, NVIDIA GPU and Network Operators, VM-based isolation via KubeVirt and VFIO, day-2 operations including upgrades, backup, recovery and node replacement
  • Architecting inference at scale: multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability and confidential-compute-capable capacity for sensitive workloads
  • Establishing observability: metrics, logging, alerting and SLOs across control plane, GPU fleet and application layers; incident response and on-call sustainability
  • Engaging partners and vendors: translating requirements into specifications and acceptance tests, running escalations, informing capacity planning and hardware decisions
  • Working directly with research, training and product teams to understand workloads and broker capacity

You bring eight or more years of hands-on infrastructure engineering, with at least three years leading teams that build and operate platforms others depend on. You've run slurmctld and slurmdbd for real production users: you understand partitions, QoS and priority, accounting, prolog and epilog scripting, node health checks and upgrades with live jobs. You hold a Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience. You communicate clearly in English and thrive in a distributed, global environment.

This is a remote, full-time position.

Data

Pay, location & hours

Salary not listed. Fully remote, open to applicants in job.

About Tether

7 open roles in this building · Company page → · See it on the map

Apply ↗

More roles to explore

Salary not listed
Tether
Apply ↗

☆ Save this job

We'll e-mail you this role so you can come back to it. No account needed.

Report this job

Reports go to the TopWeb3Jobs team. Scam reports are checked first.