Overview
Bring models from reference implementations to efficient execution on AI accelerator hardware. Own model bring-up, MLIR-based lowering, numerical correctness validation, and performance optimization across models, compilers, kernels, and runtime.
What you'll do
- Bring up new models by understanding model architectures, loading and converting weights, implementing supported execution paths, and establishing correctness against reference implementations.
- Lower models to hardware using MLIR by developing and extending dialects, graph transformations, lowering passes, and hardware-specific mappings.
- Support model operations including attention, matrix multiplication, normalization, positional embeddings, and other operators through compiler and kernel changes.
- Optimize execution through operator fusion, tensor layouts, tiling, memory allocation, data movement, and parallel execution.
- Optimize inference by tuning prefill and decode performance, KV cache management, batching, and quantization to improve latency, throughput, and memory efficiency.
- Validate model quality by investigating numerical differences and measuring the accuracy impact of precision changes and compiler optimizations.
- Diagnose bottlenecks using profiling, execution traces, and hardware counters to identify compute, memory, communication, and runtime limitations.
- Work closely with hardware, compiler, kernel, and runtime teams to deliver reliable model support and repeatable performance benchmarks.
What you'll need
- Strong programming skills in C++ and Python.
- Hands-on experience bringing up and debugging ML models in PyTorch or a comparable framework.
- Practical experience with MLIR, including dialects, rewrite patterns, transformation passes, and lowering pipelines.
- Understanding of compiler fundamentals, including intermediate representations, dataflow analysis, and code generation.
- Understanding of transformer architectures, attention mechanisms, tensor operations, and numerical precision.
- Experience profiling and optimizing workloads on GPUs or other AI accelerators.
- Ability to debug correctness and performance issues across model code, compiler-generated code, kernels, and runtime execution.
Nice to have
- Experience with LLM inference, including GQA, sliding-window attention, MoE, KV caching, and speculative decoding.
- Experience with FP16, BF16, FP8, or low-bit quantization and their accuracy and performance tradeoffs.
- Experience developing accelerator kernels or hardware-specific compiler backends.
- Familiarity with distributed execution, model parallelism, and accelerator memory hierarchies.
- Contributions to MLIR, LLVM, inference frameworks, or related open-source projects.
Details
- Location: Bengaluru, India.
Read the full description and apply on the company’s own careers page.