TopWeb3JobsTopWeb3Jobs
AI & ML district · Plot OH-A2-16

AI Research Engineer (Kernel & Inference Optimization)

Tether · Remote (job) · Full-time
◎ AI & ML🌐 Remote
Salary not listed

About the role

Tether is a global leader in digital finance infrastructure, best known for USDT, the world's most trusted stablecoin. Beyond stablecoins, the company operates across multiple domains: Tether Power manages sustainable energy solutions for Bitcoin mining, Tether Data develops AI and peer-to-peer technology including KEET (a secure communication platform), and Tether Education provides digital learning access worldwide. The organization combines blockchain expertise with emerging tech innovation, operating as a lean, fast-growing team distributed across the globe.

Tether's Data division is looking for an AI Research Engineer to optimize how AI models are deployed and served in production. You'll focus on kernel-level optimization and inference performance, working across everything from lightweight models running on constrained hardware to complex multi-modal systems handling text, images, and audio. This is a hands-on research role where you'll drive measurable improvements in real-world AI performance.

What you'll do

  • Design and implement model serving architectures that achieve high throughput and low latency while minimizing memory footprint. Your pipelines need to run efficiently on resource-constrained devices and edge platforms, with clear performance targets around reduced latency and token response times.
  • Run controlled inference tests in both simulated and production environments. You'll track latency, throughput, memory consumption, and error rates—paying special attention to how models perform on low-resource hardware—then document results against established benchmarks.
  • Build test datasets and simulation scenarios that mirror real-world challenges on edge devices. You'll define measurable criteria to evaluate model performance, latency, and memory usage under different operational conditions.
  • Debug and optimize the serving pipeline by analyzing computational efficiency and identifying bottlenecks. This includes fixing batch processing inefficiencies, network delays, and memory usage problems to improve scalability on resource-constrained systems.
  • Partner with cross-functional teams to integrate optimized inference frameworks into production pipelines for edge and on-device deployment. You'll define success metrics around real-world performance, error rates, and memory usage, then drive continuous refinements.

What you'll bring

  • A degree in Computer Science or related field, ideally a PhD in NLP, Machine Learning, or similar, with a strong publication record in top-tier conferences.
  • Deep expertise in Metal Shading Language (MSL) and the ability to write custom compute shaders from scratch.
  • Proven track record optimizing kernels and inference on mobile devices. Your work should demonstrate measurable improvements in latency, throughput, and memory footprint for domain-specific applications, particularly on edge and resource-constrained platforms.
  • Strong research foundation in advanced model architectures and model serving pipeline design.

Remote, full-time position.

Data

Pay, location & hours

Salary not listed. Fully remote, open to applicants in job.

About Tether

7 open roles in this building · Company page → · See it on the map

Apply ↗

More roles to explore

Salary not listed
Tether
Apply ↗

☆ Save this job

We'll e-mail you this role so you can come back to it. No account needed.

Report this job

Reports go to the TopWeb3Jobs team. Scam reports are checked first.