Sarvam AI logo

Performance Engineer, Inference

Sarvam AI
Posted a month ago

LOCATION

Bengaluru · Onsite

EXPERIENCE

5+ Years

TYPE

FullTime

SKILLS REQUIRED

Inference ServingDistributed SystemsPerformance ProfilingTransformer Inference Optimization

Job description

Overview

Senior Performance Engineer (Inference) to own the production inference serving path for large distributed models, including latency/cost metrics and deep integration of serving and kernel layers.

What you'll do

  • Own the end-to-end production serving path for large distributed models.
  • Source-level read and modify an inference stack such as SGLang, vLLM, NVIDIA Dynamo, or TensorRT-LLM.
  • Operate and extend a distributed-serving stack with disaggregated prefill-decode and distributed KV/cache transfer across nodes.
  • Integrate artifacts from model and kernel teams into a multi-node, multi-tenant serving stack.
  • Build and train speculative/draft models and tune acceptance rate against the live serving distribution.
  • Produce and defend latency/throughput and related performance numbers, and collaborate on architecture with kernels, model, and SRE teams.
  • Track TTFT (p50/p95/p99), TPOT, throughput, GPU utilization, and cost per million tokens.

What you'll need

  • 5+ years in ML systems, including 2+ years on inference serving at production scale.
  • Experience serving 100B+ parameter models in production using multi-node tensor, pipeline, or expert parallelism.
  • Source-level fluency in at least one of SGLang, vLLM, Dynamo, or TensorRT-LLM, plus reading-level familiarity with the other three.
  • Distributed serving experience at operating-and-extending depth (disaggregated prefill-decode, distributed KV/cache transfer, routing/scheduling across nodes).
  • Competency training and tuning speculative decoding (draft models/speculators, distillation, acceptance-rate tuning, composition with the stack).
  • Deep understanding of KV cache internals (block tables, copy-on-write, prefix sharing, fragmentation).
  • Working command of TP/PP/EP and NCCL primitives, plus C++ and CUDA at a read-and-modify level.

Details

  • Location: Bengaluru.

Read the full description and apply on the company’s own careers page.

Stay safe

Hiring on Abekus is free for applicants

We never charge a fee, and employers are prohibited from doing so. If a recruiter asks for payment, please report them right away.

Performance Engineer, Inference