Overview
Senior Systems Software Engineer (Local AI) to build and optimize on-device/local inference software for RTX and DGX-class systems.
What you'll do
- Partner with NVIDIA software, research, architecture, and product teams to align systems strategies and technical needs.
- Build and optimize a local AI inference stack for RTX, RTX Pro, and DGX GPUs for performance, stability, and scalability.
- Design inference runtimes and execution stacks across frameworks including Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT-RTX.
- Conduct end-to-end optimization of AI models, data pipelines, and inference runtimes across GPU architectures.
- Apply model optimization techniques such as quantization, pruning, sparsity, and distillation for efficient local deployment.
- Perform system-level debugging and performance–accuracy trade-off analysis, including infrastructure for performance/accuracy sweeps.
- Establish engineering guidelines to accelerate bring-up and support production readiness of new models and backends.
What you'll need
- 5+ years of experience with a BS/MS/PhD in Computer Science, Software Engineering, Mathematics, or related field, or equivalent experience.
- Excellent C++ programming and debugging skills, including data structures, algorithms, and machine learning.
- Proven experience with AI inferencing pipelines and applications using ML/DL frameworks such as Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT.
- Deep interest in inference backends and runtime internals (e.g., scheduling, memory management, KV-cache, graph execution, quantization).
- Strong analytical and problem-solving abilities for multitasking in a dynamic environment.
- Outstanding written and oral communication skills for collaboration with management and engineering teams.
Details
Read the full description and apply on the company’s own careers page.