Overview
Senior Engineer- DevOps on the Hyperscale (HSX) Engineering team, focused on automation, infrastructure, system-level troubleshooting, and reliability across distributed environments spanning compute, storage, networking, virtualization, containers, and cloud platforms.
What you'll do
- Design, develop, and maintain automation using Python, Shell scripting, and Ansible for server provisioning, configuration, deployment, upgrades, and patching of Hyperscale environments.
- Develop reusable Ansible roles, modules, and playbooks to standardize and simplify product deployment and infrastructure configuration.
- Build and enhance automation frameworks for deploying and configuring distributed system components.
- Automate deployment and configuration of HAProxy, Corosync, and other clustering, high-availability, and monitoring components.
- Develop and maintain automation for S3 deployment, configuration, and validation/testing.
- Deploy and validate software across multiple operating systems and hardware configurations to ensure compatibility and reliability.
- Work with modern bare-metal provisioning frameworks such as RackN and extend existing frameworks when required.
- Develop tools and scripts that improve engineering productivity, diagnostics, deployment, testing, and operational efficiency.
- Navigate and maintain existing C++ and XML-based codebases, making targeted enhancements and fixes when required.
- Effectively use AI-assisted software development tools, including GitHub Copilot and similar technologies, to accelerate development, code understanding, debugging, test creation, and modernization of legacy components.
- Design and maintain Infrastructure as Code (IaC) solutions for provisioning and managing development and test environments.
- Automate creation, modification, and deletion of virtual machines and infrastructure resources on VMware and Hyper-V.
- Deploy, configure, and troubleshoot containerized workloads using Docker, Kubernetes, and OpenShift.
- Develop infrastructure automation using Python libraries and frameworks such as Boto3, Paramiko, Fabric, and PyYAML.
- Configure and troubleshoot compute, storage, and networking resources across cloud, virtualized, and on-premises environments.
- Automate bare-metal operating system and HSX ISO installation using PXE/network boot technologies, including RackN and Cobbler.
- Use Git-based development workflows for source control, collaboration, code reviews, and release management.
- Analyze system logs, application logs, crash information, error reports, network behavior, storage behavior, and performance metrics to identify root causes.
- Build and maintain engineering environments for functional, integration, regression, upgrade, and compatibility testing across major and minor product releases.
- Validate product functionality across supported hardware platforms, operating systems, hypervisors, cloud environments, and infrastructure configurations.
What you'll need
- Bachelor's degree.
- 4 or more Years of experience.
- Strong software development experience with Python and Shell scripting.
- Hands-on experience with configuration management and automation technologies, particularly Ansible.
- Strong understanding of Linux operating systems, system administration, networking, storage, and distributed system concepts.
- Experience with virtualization platforms such as VMware and/or Hyper-V.
- Working knowledge of at least one major cloud platform: AWS, Azure, or GCP.
- Experience with Docker, Kubernetes, and/or OpenShift.
- Experience developing infrastructure and system automation using Python libraries such as Boto3, Paramiko, Fabric, or PyYAML.
- Familiarity with bare-metal provisioning and network boot technologies such as PXE, RackN, or Cobbler.
- Working knowledge of Git and modern software development workflows.
- Strong debugging and problem-solving skills with the ability to analyze complex issues across application and infrastructure layers.
- Ability to read and understand existing C++ and XML code and make targeted changes when required.
- Experience troubleshooting production or customer environments and performing detailed root-cause analysis (RCA).
Nice to have
- Enterprise storage, backup, recovery, or data-management platforms preferred.
- Security and SaaS experience strongly preferred.
Details
- Locations: Bangalore, Hyderabad, and Pune, India.
Read the full description and apply on the company’s own careers page.