Stone head

Lead Inference Engineer

  • $270,000 - $330,000
  • Work Location Type: Remote
  • Timezone/Location: United States
  • ID: 4804
  • Posted: 28.08.26

Plexus has partnered with a rapidly growing AI company looking for a GPU Performance Engineering Lead to build and lead a specialist infrastructure team.

This is a hands-on leadership role focused on improving the speed, efficiency and reliability of large-scale AI workloads. You’ll set the technical direction, hire the team and work directly on complex performance challenges across distributed GPU infrastructure.

Responsibilities:

  • Own the technical strategy for LLM inference performance

  • Recruit and lead the Inference Optimisation team

  • Optimise latency, throughput and cost per token across GPU infrastructure

  • Benchmark inference engines, quantisation methods and parallelism strategies

  • Improve inference routing and load balancing

  • Evaluate new kernels, hardware and optimisation techniques

Requirements:

  • 8+ years in performance optimisation, HPC or GPU engineering

  • 5+ years leading engineering teams

  • Production experience with vLLM, SGLang or similar inference engines at scale

  • Strong knowledge of GPU architecture, distributed inference and quantisation

  • Experience with GPU profiling tools such as Nsight or PyTorch Profiler

  • Proficiency in Python, Rust or Go

  • C++, CUDA or Triton experience is highly desirable

Offer:

  • $270k–$330k base salary

  • Fully remote within the US

  • Opportunity to build and lead the inference function

  • Work on high-volume AI infrastructure and cutting-edge GPU technology

If this piques your interest, please apply now!

Apply for this job: