Lead Inference Engineer
- $270,000 - $330,000
- Work Location Type: Remote
- Timezone/Location: United States
- ID: 4804
- Posted: 28.08.26
Plexus has partnered with a rapidly growing AI company looking for a GPU Performance Engineering Lead to build and lead a specialist infrastructure team.
This is a hands-on leadership role focused on improving the speed, efficiency and reliability of large-scale AI workloads. You’ll set the technical direction, hire the team and work directly on complex performance challenges across distributed GPU infrastructure.
Responsibilities:
-
Own the technical strategy for LLM inference performance
-
Recruit and lead the Inference Optimisation team
-
Optimise latency, throughput and cost per token across GPU infrastructure
-
Benchmark inference engines, quantisation methods and parallelism strategies
-
Improve inference routing and load balancing
-
Evaluate new kernels, hardware and optimisation techniques
Requirements:
-
8+ years in performance optimisation, HPC or GPU engineering
-
5+ years leading engineering teams
-
Production experience with vLLM, SGLang or similar inference engines at scale
-
Strong knowledge of GPU architecture, distributed inference and quantisation
-
Experience with GPU profiling tools such as Nsight or PyTorch Profiler
-
Proficiency in Python, Rust or Go
-
C++, CUDA or Triton experience is highly desirable
Offer:
-
$270k–$330k base salary
-
Fully remote within the US
-
Opportunity to build and lead the inference function
-
Work on high-volume AI infrastructure and cutting-edge GPU technology
If this piques your interest, please apply now!
