Overview
Architect and deliver generative AI solutions focused on Speech AI, Large Language Models (LLMs), Agentic AI and retrieval-augmented generation (RAG) workflows. Collaborate with customers, sales, business development and engineering teams to design, deploy and optimize solutions using NVIDIA hardware and software platforms.
What you'll do
- Architect end-to-end generative AI solutions focused on Speech AI, LLMs, Agentic AI and RAG workflows.
- Collaborate with customers to understand language-related business challenges and design tailored solutions.
- Support pre-sales activities with sales and business development teams, including technical presentations and demonstrations of Speech AI, LLM and RAG capabilities.
- Provide feedback to NVIDIA engineering teams and contribute to the evolution of generative AI technologies.
- Engage directly with customers to understand their language-related requirements and challenges.
- Lead workshops and design sessions to define and refine generative AI solutions focused on LLMs and RAG workflows.
- Lead the training and optimization of LLMs using NVIDIA hardware and software platforms.
- Implement strategies for efficient and effective LLM training to achieve optimal performance.
- Design and implement RAG-based workflows to enhance content generation and information retrieval.
- Integrate RAG workflows into customers' applications and systems.
- Stay abreast of developments in language models and generative AI technologies.
- Provide technical leadership and guidance on best practices for training LLMs and implementing RAG-based solutions.
What you'll need
- B.Tech, Master's or Ph.D. in Computer Science, Artificial Intelligence, or equivalent experience.
- 8+ years of hands-on experience in a technical role, specifically focusing on generative AI, with a strong emphasis on training Large Language Models (LLMs) and Speech AI.
- Proven track record of successfully deploying and optimizing LLM and Speech AI models for inference in production environments.
- In-depth understanding of state-of-the-art language models, including but not limited to GPT-3, BERT, or similar architectures.
- Expertise in training and fine-tuning LLMs using popular frameworks such as TensorFlow, PyTorch, or Hugging Face Transformers.
- Proficiency in model deployment and optimization techniques for efficient inference on various hardware platforms, with a focus on GPUs.
- Strong knowledge of GPU cluster architecture and the ability to leverage parallel processing for accelerated model training and inference.
- Excellent communication and collaboration skills with the ability to articulate complex technical concepts to both technical and non-technical stakeholders.
- Experience leading workshops, training sessions, and presenting technical solutions to diverse audiences.
Nice to have
- Proven ability to optimize LLM models and Speech models for inference speed, memory efficiency, and resource utilization.
- Familiarity with containerization technologies, such as Docker, and orchestration tools, such as Kubernetes, for scalable and efficient model deployment.
- Deep understanding of GPU cluster architecture, parallel computing, and distributed computing concepts.
- Hands-on experience with NVIDIA GPU technologies and GPU cluster management.
- Ability to design and implement scalable and efficient workflows for LLM training and inference on GPU clusters.
Details
- Location: Bengaluru, India.
Read the full description and apply on the company’s own careers page.