Overview
Lead end-to-end AI inference architecture, optimization, and production deployment on Qualcomm's datacenter AI platform. Serve as the technical lead for customer engagements and drive complex programs across engineering teams.
What you'll do
- Lead architecture and end-to-end design of AI inference solutions on Qualcomm AI hardware.
- Translate customer requirements into serving architectures and validate production readiness.
- Drive development and deployment of GenAI and LLM applications.
- Define benchmarking programs and investigate complex system-level issues.
- Mentor engineers and provide technical leadership on project execution.
- Collaborate across compiler, runtime, serving, model engineering, and customer engineering teams.
- Define reusable platform capabilities, tooling, and documentation.
What you'll need
- 5+ years of systems engineering or related experience with a bachelor's degree, or equivalent experience through a master's or PhD.
- 4+ years of software engineering or related experience with a bachelor's degree, or equivalent experience through a master's or PhD.
- 2+ years of experience with programming languages such as C, C++, Java, or Python.
- Strong proficiency in Python and ML frameworks including PyTorch and TensorFlow.
- Deep understanding of ML model development, deployment, and production inference.
- Deep understanding of system performance profiling and parallel computing.
- Bachelor's degree in Engineering, Information Systems, Computer Science, or a related field.
Nice to have
- Advanced understanding of GenAI architectures including transformers, diffusion models, LLMs, vision-language models, and embedding models.
- Experience with inference optimization techniques such as speculative decoding, KV cache management, continuous batching, and tensor or pipeline parallelism.
- Experience designing and operating large-scale distributed AI systems.
- Experience with MLOps, containerization, automation, and ML lifecycle tooling.
- Experience leading complex technical programs across large, matrixed organizations.
- Experience with rack-level orchestration and datacenter automation.
Details
- Location: Bangalore, India.
Read the full description and apply on the company’s own careers page.