Overview
System Software Engineer for NVIDIA’s LocalAI team to build and optimize on-device AI software for RTX and DGX-class systems, focusing on local inference performance and efficient resource use.
What you'll do
- Partner with NVIDIA teams to align strategies and technical needs for local AI on RTX and DGX.
- Build and optimize a local AI inference stack for RTX, RTX Pro, and DGX GPUs.
- Design inference runtimes and execution stacks using frameworks and tooling listed in the posting.
- Perform end-to-end optimization of AI models, data pipelines, and inference runtimes.
- Apply model optimization techniques including quantization, pruning, sparsity, and distillation.
- Conduct system-level debugging and performance–accuracy trade-off analysis.
- Establish engineering guidelines for bring-up and production readiness of new models and backends.
What you'll need
- 2+ years of experience with a Bachelor’s, Master’s, or PhD in CS/Software Engineering/Math or related field (or equivalent experience).
- Excellent C++ programming and debugging skills.
- Strong understanding of data structures, algorithms, and machine learning.
- Proven experience with AI inferencing pipelines using frameworks such as Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT.
- Knowledge of inference backends/runtime internals including scheduling, memory management, KV-cache behavior, and quantization.
- Strong analytical and problem-solving abilities.
- Outstanding written and oral communication skills for collaboration.
Details
Read the full description and apply on the company’s own careers page.