Overview
Specialist, HPC Cloud Engineering responsible for maintaining and supporting HPC solutions in the cloud, using infrastructure-as-code and automation to improve researcher workflows.
What you'll do
- Deploy, maintain, and support high-performance computing clusters, platforms, and AWS managed services.
- Monitor and maintain HPC cluster performance on premises and in AWS for availability, reliability, and user experience.
- Optimize compute and data workflows for performance and cost-efficiency in cloud-based HPC.
- Create scripts and tools to automate cluster deployment, monitoring, and management (e.g., Terraform, Ansible, AWS CLI).
- Analyze bottlenecks in HPC cloud environments, including network performance and parallel file systems.
- Diagnose and resolve issues involving job scheduling, software, and cloud infrastructure.
- Document processes, configurations, and best practices for HPC workloads in AWS and Linux environments.
What you'll need
- Bachelor’s degree in Information Technology, Computer Science, or any Technology stream.
- 3-7 years of experience in HPC infrastructure engineering, including network principles and Linux/Unix system administration.
- Hands-on proficiency in Linux, shell scripting, HPC architecture, and job schedulers in a DevOps environment.
- Experience with job schedulers/orchestrators such as GridEngine, PBS Pro, and Altair NavOps.
- Knowledge of network architectures and high-speed interconnects.
- Understanding of relevant regulations and compliance standards.
- Experience with disaster recovery, data backup solutions, and virtualization platforms.
Details
- Location: Hyderabad (Hitec City Raidurg), Telangana, India.
- Work mode: Hybrid.
- Employment status: Regular.
- Travel requirements: No Travel Required.
- VISA sponsorship: No.
- Relocation: Domestic.
Read the full description and apply on the company’s own careers page.