Overview
The Senior Principal Infrastructure Engineer - Technical Major Incident Manager serves as the senior technical authority for leading the response, coordination and resolution of critical technology incidents across infrastructure, cloud and application ecosystems. Operating as an Individual Contributor within Infrastructure Operations, the role drives service restoration, incident command, technical escalation management, root cause elimination, operational resilience and service reliability improvements.
What you'll do
- Lead and coordinate P1/P2/P3 major incidents across infrastructure and application domains.
- Act as Technical Incident Commander during critical business-impacting outages.
- Drive technical triage, escalation management, service restoration and executive communications.
- Facilitate post-incident reviews and corrective action tracking.
- Review incident trends and identify systemic risks and reliability gaps.
- Drive root cause analysis and permanent resolution of recurring issues.
- Partner with Operations and Engineering teams to implement resilient and scalable solutions.
- Govern RCA quality and effectiveness of corrective actions.
- Establish reliability scorecards and operational health measures.
- Drive reliability maturity assessments and define operational reliability standards across multiple technology domains.
- Identify operational risks and implement preventive measures to minimize outages.
- Support observability, monitoring, event management and alert optimization initiatives.
- Ensure adherence to operational procedures, standards and service management processes.
- Identify opportunities to improve service reliability, operational efficiency and infrastructure performance.
- Develop and maintain operational runbooks, standard operating procedures and technical documentation.
- Promote automation and process optimization to reduce manual effort, operational risks and repetitive tasks.
- Analyze operational trends, recurring incidents and service issues to recommend preventive and corrective actions.
- Monitor and improve service availability, incident reduction, SLA compliance and operational efficiency.
- Lead operational automation programs and drive adoption of AIOps capabilities.
- Expand self-healing operational practices and reduce operational toil through workflow automation.
- Mentor and upskill team members and drive technical communities of practice.
- Lead strategic programs through influence, align stakeholders toward reliability outcomes and coordinate initiatives across Global Teams.
- Drive major incident command and rapid service restoration during critical technology outages.
- Advance Reliability Engineering and SRE practices to improve service performance and resilience.
- Establish and govern observability, monitoring and alert optimization capabilities.
- Drive automation and self-healing initiatives to reduce operational risk and manual effort.
- Partner with engineering and support teams to deliver sustainable solutions that enhance service reliability and prevent future disruptions.
What you'll need
- 16+ years of experience in Infrastructure Operations, Production Support, Technical Operations, Enterprise Technology Services or Operations Management within large-scale 24x7 enterprise environments.
- Bachelor’s degree in Engineering, Computer Science, Information Technology, MCA or a related discipline.
- Proven experience leading and supporting enterprise infrastructure environments across Windows Server, Linux/AIX, VMware Virtualization, Network Services, Databases, Storage, Backup, Container Platform and application support.
- Strong technical knowledge of Windows Server Administration, VMware vSphere/ESXi/vCenter/Hyper-V, Linux & AIX, Routing & Switching, DDI (Infoblox, Efficient IP), Load Balancers (F5, A10), Cisco ISE/NAC, Wireless Technologies (Aruba, Cisco) and Automation Platforms (Ansible, Gluware).
- Experience supporting and managing Oracle, Microsoft SQL Server, PostgreSQL, SAN, NAS and Commvault environments.
- Strong understanding of Infrastructure Operations, Service Delivery, Incident Management, Problem Management, Change Management and Service Request Fulfillment processes.
- Demonstrated experience driving operational excellence, infrastructure stability, service availability, process standardization, automation and continuous improvement initiatives.
- Strong knowledge of ITIL frameworks and enterprise operational governance practices.
- Experience working with cross-functional teams, technology vendors and business stakeholders to ensure delivery of reliable and secure infrastructure services.
- Excellent stakeholder management, communication, collaboration and problem-solving skills.
- Hands-on experience in at least three of the following technology domains: Windows Server Administration; Virtualization Platforms (VMware vSphere, ESXi, vCenter, Hyper-V); Unix Systems (Linux & AIX); Network Services & Platforms; Database; Storage & Backup Technologies; Container platforms.
- Working knowledge of Ansible, Python, Shell Scripting and PowerShell.
- Ability to identify and implement automation opportunities that improve incident response, operational efficiency and service reliability.
- Experience using Generative AI (GenAI) tools to enhance incident analysis, reporting, knowledge management and operational effectiveness.
Nice to have
- Experience within Financial Services, Banking or other highly regulated enterprise environments.
- Experience with automation and operational tooling.
- ITIL Foundation / ITIL 4 Managing Professional certification.
- Site Reliability Engineering (SRE) Certification.
- Major Incident Management or IT Service Management (ITSM) Certification.
- DevOps Certifications (DevOps Foundation, Azure DevOps, etc.).
- Cloud Certifications (AWS, Microsoft Azure, Google Cloud Platform).
- Six Sigma Green Belt.
- VMware Certified Professional (VCP).
- Red Hat Certified Engineer (RHCE) / RHEL Certification.
- Automation Certifications (Ansible, Python, PowerShell).
Details
- Location: Bangalore, India.
- Work mode: Hybrid.
- Career level: P5.
- Job category: Vice President.
- Role type: Individual Contributor.
- Reports to the Director – Technology.
- Partners with Technology Leaders, Engineering Leaders, Infrastructure Leaders, Architecture Teams, Risk & Governance Organizations and Vendor Partners.
- The role operates across large-scale 24x7 enterprise environments and requires availability during critical incidents and high-impact business events.
Read the full description and apply on the company’s own careers page.