Overview
Performance Engineering Architect responsible for characterizing and optimizing AI solution performance across GPUs, servers, storage, networking, system software, and AI frameworks. The role combines benchmarking, system analysis, workload optimization, automation, and cross-functional collaboration.
What you'll do
- Design and execute performance studies for generative AI, machine learning, and other AI workloads.
- Benchmark inference and training across models, GPUs, precision formats, deployment profiles, and scale configurations.
- Analyze GPU, CPU, memory, PCIe, storage, networking, operating-system, and application performance.
- Develop benchmark suites and automate execution, telemetry collection, analysis, and reporting.
- Identify bottlenecks and recommend hardware, firmware, operating-system, driver, runtime, framework, or application changes.
- Establish repeatable methodologies, baselines, acceptance criteria, and release-performance gates.
- Produce performance reports, sizing evidence, engineering recommendations, and field guidance.
What you'll need
- Strong understanding of generative AI and machine-learning performance, especially inference.
- Hands-on experience deploying and benchmarking AI or GenAI workloads on GPU-accelerated systems.
- Experience with LLM serving or inference runtimes such as NVIDIA NIM, vLLM, Triton, TensorRT-LLM, Hugging Face, or PyTorch.
- Knowledge of precision formats, quantization, parallelism, distributed execution, and inference metrics.
- Experience with containers, Kubernetes or OpenShift, Linux administration, Python, and shell scripting.
- Strong understanding of modern server architecture, enterprise hardware, BIOS and firmware performance settings.
- Ability to document reproducible tests, analyze telemetry, and communicate performance findings clearly.
Nice to have
- Experience with VMware, KVM, OpenShift Virtualization, or SUSE Harvester.
- Experience with industry benchmarks, databases, RAG pipelines, AI storage, or performance-regression automation.
- Experience supporting benchmark audits or externally published performance results.
Details
- Onsite role with an expectation of primarily working from an HPE office.
- Location: Bengaluru, Karnataka, India.
Read the full description and apply on the company’s own careers page.