Overview
Provide technical leadership and strategic guidance across AI Accelerator engagements with AI Native organizations, NeoCloud Providers, and ISVs. Advise on architecture and integration, define what good looks like, and bring learnings back to inform the DSX product roadmap when standard product capabilities are not enough.
What you'll do
- Provide architectural direction across strategic engagements requiring advanced implementation, optimization, or integration customization.
- Help customers integrate the right components to deliver on their outcomes and advise on adopting DSX software where it fits.
- Help customers succeed with the right alternative where DSX software does not fit, and bring product gaps back to product and engineering.
- Dive into complex technical challenges hands-on to solve critical problems, validate architectures, or prove out solutions.
- Lead technically demanding programs end to end, including third-party performance benchmarking across hardware and workloads.
- Identify common challenges and solution patterns across engagements.
- Share findings with internal teams and the broader AI community.
- Develop standardized approaches, reference architectures, and structured guidance rooted in successful engagement patterns.
- Partner with product, engineering, and other customer-facing NVIDIA teams so field learnings inform internal strategy and capabilities.
- Design technical strategies for distributed training, large-scale inference, model and pipeline optimization, and MLOps across multiple customers and partners.
- Help develop infrastructure patterns and playbooks for the latest NVIDIA hardware as it lands with customers and partners.
What you'll need
- Bachelor's degree or equivalent experience.
- 10+ years in technical roles such as solutions architecture, ML engineering, technical product management, or technical consulting across multiple customers or projects, or 5+ years of specialist-level experience working at the frontier of AI infrastructure.
- Strong technical leadership with the ability to guide teams and influence technical decisions without direct authority.
- Systems thinking with the ability to understand customer outcomes and translate them into clear technical requirements and architectures.
- Willingness to prototype, implement, validate, and troubleshoot hands-on when needed to solve critical problems or prove out approaches.
- A solid technical foundation in the technologies AI infrastructure is built on, especially Linux systems administration.
- Ability to ramp on brand new technologies and unfamiliar technical domains independently as a self-directed learner.
- Strong communication skills with the ability to engage technical teams, executives, and multi-functional collaborators.
Nice to have
- Solutions architecture or technical consulting background across multiple customer engagements simultaneously, with experience bringing novel AI hardware or frameworks to production with frontier AI Native organizations, hyperscalers, NeoClouds, or ISVs.
- A foundational cloud or distributed systems background built at hyperscaler scale.
- A public technical voice demonstrated through blog posts, talks, open-source contributions, or reference work that shows depth and opinion.
- Hands-on technical expertise in one or more of the NVIDIA Stack: CUDA, NeMo, Triton, TensorRT, NIM, DGX Cloud, and the broader DSX software portfolio.
- Hands-on technical expertise in inference systems, including large-scale inference with frameworks like vLLM and SGLang, prefill-decode disaggregation, and performance optimization across hardware.
- Hands-on technical expertise in training systems, including distributed training, model and pipeline optimization, and open-source generative AI frameworks.
- Hands-on technical expertise in infrastructure, including SLURM, Kubernetes, GPU scheduling, distributed computing frameworks, rack-scale systems, and multiple CSP or NCP cloud environments.
- Hands-on technical expertise in observability and automation, including CI/CD, infrastructure as code, and GPU performance monitoring.
Details
- Location: Bengaluru, India.
Read the full description and apply on the company’s own careers page.