About the role
Tether is a global fintech leader building reserve-backed digital assets and blockchain infrastructure that power hundreds of millions of users worldwide. The company operates across multiple verticals including stablecoins, energy solutions, peer-to-peer communications technology, and digital education. You'll join their Data team as an AI Research Engineer focused on large-scale pre-training for language models and multi-modal systems.
Tether has established itself as a trusted name in the Web3 and fintech space by combining technological innovation with transparent operations. Their commitment to building lean, high-performing teams across remote locations has made them a destination for top technical talent globally. The company's expansion into AI reflects the broader 94% surge in AI/ML engineering vacancies and the growing demand for researchers who can scale models across distributed infrastructure.
What you'll do
- Lead foundational pre-training initiatives for large language models and multi-modal architectures (combining text, vision, audio, and other modalities) on distributed servers with thousands of NVIDIA GPUs across multiple nodes
- Design and prototype novel architectures, tokenizers, and cross-modal alignment layers that strengthen model intelligence and multi-modal reasoning capabilities
- Build comprehensive data pipelines to source, filter, and curate massive-scale textual and multi-modal datasets for efficient pre-training at scale
- Execute experiments independently and collaboratively, analyzing results to refine training methodologies and optimize token efficiency
- Debug and resolve performance bottlenecks in model efficiency, computational performance, and multi-modal alignment stability during extended training runs
- Advance distributed training systems to improve scalability and hardware utilization across target platforms
What you'll bring
- A degree in Computer Science or related field, ideally with a PhD in NLP, Machine Learning, or adjacent discipline and a strong publication record at top-tier conferences
- Hands-on experience running large-scale LLM or multi-modal pre-training on distributed GPU clusters with thousands of NVIDIA processors
- Practical expertise with distributed training frameworks, libraries, and tools at production scale
- Deep knowledge of transformer and non-transformer architectures designed to enhance intelligence, efficiency, and scalability
- Advanced proficiency in PyTorch and Hugging Face libraries with demonstrated experience in model development, continual pre-training, and deployment
- Strong English communication skills and a research-driven mindset with ability to explore novel techniques and algorithms
This is a fully remote position offering you the opportunity to work alongside a globally distributed team of elite technical researchers pushing the boundaries of AI capabilities.
Pay, location & hours
Salary not listed. Fully remote, open to applicants in job.
About Tether
19 open roles in this building · Company page → · See it on the map
Databricks · Remote