Overview
The Storage and Backup Engineer protects and recovers data across on-premises and cloud environments and administers the associated storage, compute and virtualization platforms. The role operates within a shift-based infrastructure team in Navi Mumbai and is accountable for backup success, restore performance, disaster recovery capability and recovery objectives.
What you'll do
- Administer Rubrik, Veeam and NetBackup data protection platforms, including policy configuration, protection groups, SLA domains, scheduling and agent or connector health.
- Administer tape infrastructure, including libraries, drives, media pools, barcode and slot management, and offsite media rotation and vaulting.
- Monitor backup and replication outcomes daily, investigate failures through root cause and remediate them.
- Maintain immutable and air-gapped protection copies and verify that immutability is enforced.
- Maintain platform currency through software and firmware upgrades and manage vendor cases through closure.
- Ensure protected workloads have appropriate policies and identify and onboard unprotected workloads across on-premises and cloud estates.
- Administer AWS backup and archive using S3 and Glacier storage classes, bucket and lifecycle policies, versioning, retention, expiry and tier transitions.
- Configure and maintain S3 Object Lock and AWS Backup Vault Lock and verify WORM retention enforcement.
- Administer AWS Backup, EBS and EC2 snapshots, RDS snapshots, and EFS or FSx backups.
- Manage replication from on-premises Rubrik clusters to Rubrik Cloud, including archival locations, replication and retention policies, seeding, bandwidth and reconciliation.
- Maintain IAM roles, policies and KMS keys for backup and restore, applying least privilege and cross-account separation.
- Maintain backup connectivity and throughput, including VPC endpoints, Direct Connect or VPN capacity, and shared bandwidth impacts.
- Execute and test cloud recovery, including Glacier restores, recovery of on-premises workloads into AWS and recovery of cloud-native workloads.
- Monitor cloud protection costs, including storage class placement, retrieval and egress charges, orphaned snapshots and unattached volumes, and report budget variance.
- Execute restores and recoveries across file, virtual machine, database and application workloads.
- Provide second- and third-level support for complex data protection, storage and compute recovery issues through root cause across backup agent, array, hypervisor, database and network layers.
- Perform data copy and migration between arrays, sites, tiers and cloud targets, including cutover planning, integrity validation and rollback.
- Support legal hold, eDiscovery and audit requests requiring retained or archived data retrieval.
- Participate in priority and major incidents involving data loss, corruption or platform failure and provide recovery estimates.
- Operate restoration tests across file, virtual machine, database and application-level workloads on a defined cadence.
- Agree restoration test schedules, scope and acceptance criteria with infrastructure teams and service and application owners.
- Document test outcomes, failures, corrective actions and revised recovery estimates, and retest after remediation.
- Publish restoration test results and maintain visible outstanding findings and an audit-ready evidence pack.
- Maintain and improve disaster recovery capability, including replication topology, failover design and runbooks.
- Maintain per-workload RTO and RPO positions against the Institute's global standards and formally escalate shortfalls.
- Plan and execute disaster recovery drills covering failover, failback and validation, and track findings to verified closure.
- Author and maintain runbooks for backup, restore, failover and recovery.
- Follow and enforce change control, including rollback planning and coordination of maintenance, drill and test windows.
- Administer enterprise SAN and NAS storage, including volume and LUN provisioning, snapshots, array-based replication, zoning support and performance troubleshooting.
- Administer VMware and Nutanix hosts and clusters, Windows and Linux workloads, and SQL Server and Oracle database backup interfaces.
- Maintain the physical storage, tape and server estate, including racking, cabling, firmware levels and OEM support status.
- Manage capacity planning and forecasting across storage, compute and protection estates and raise procurement requirements.
- Optimize retention, deduplication, compression, tiering, archive placement, cloud storage class placement, orphaned snapshots and unattached volumes.
- Manage storage, compute and backup hardware, media and licensing lifecycles, including refresh planning, media retirement and secure disposal.
- Produce recurring backup, restore, restoration-test, capacity and compliance reports for IT leadership, information security, audit and service owners.
- Develop PowerShell, Python or Bash scripting and platform API automation for administration and protection reporting.
- Maintain runbooks, procedures, configuration and retention documentation and provide technical guidance and mentoring to less experienced engineers.
- Complete documented shift handovers covering running jobs, failed backups, open restores, in-flight migrations, active incidents and pending actions, and confirm ownership of open items.
- Coordinate across shifts and infrastructure teams so protection and recovery activity continues without loss of context across regions.
What you'll need
- Bachelor’s degree in computer science, Information Technology, Engineering or a closely related field, or equivalent professional experience.
- 7+ years of experience in systems engineering, storage administration or data protection, including substantial hands-on responsibility for a production backup estate.
- Hands-on administration of at least two enterprise data protection platforms from Rubrik, Veeam, NetBackup or comparable products, including policy design and failure remediation.
- Hands-on administration of enterprise SAN and NAS storage, including volume and LUN provisioning, snapshots, replication and capacity management.
- Hands-on experience with compute and virtualization platforms as they relate to protection and recovery, including VMware or Nutanix, Windows and Linux workloads, and database backup interfaces such as SQL Server or Oracle.
- Experience with tape infrastructure, including libraries, media management and offsite rotation.
- Experience executing restores and recoveries across files, virtual machine, database and application workloads.
- Experience operating a scheduled restoration testing program, including test design, execution with service owners, evidence capture and remediation of findings.
- Experience maintaining per-workload recovery time and recovery point objectives against a defined organizational standard, including measurement and escalation of shortfalls.
- Active participation in disaster recovery drills as an executing engineer, including failover, failback and validation.
- Experience producing recurring backup, recovery, capacity and compliance reporting for IT leadership, security or audit.
- Hands-on AWS experience as a backup and archive target, including S3 and Glacier storage classes, bucket and lifecycle policy, versioning, retention and cross-region or cross-account copies.
- Hands-on cloud immutability experience, including S3 Object Lock or AWS Backup Vault Lock, and IAM roles, policies and KMS encryption for backup access and recovery.
- Hands-on experience with Rubrik replication and archival to Rubrik Cloud, or equivalent vendor cloud replication and archive from an on-premises backup platform.
- Experience protecting AWS-native workloads, including AWS Backup, EBS and EC2 snapshots, and RDS, EFS or FSx backups.
- Experience executing cloud recovery, including archive-tier restores with retrieval time and cost understood in advance.
- Experience managing cloud protection cost, including storage class placement, retrieval and egress charges, and removal of orphaned snapshots.
- Experience working within a shift-based infrastructure team, including structured handover of running jobs and open recovery work.
- Scripting proficiency in PowerShell, Python or Bash, with practical use of platform APIs for reporting and bulk administration.
- Professional proficiency in spoken and written English sufficient to coordinate with global infrastructure and service teams and communicate with senior stakeholders.
Nice to have
- Certification in a relevant platform such as Rubrik, Veeam Certified Engineer (VMCE), NetBackup, or a storage vendor credential from NetApp, Dell, Pure Storage or equivalent.
- AWS certification such as AWS Certified Solutions Architect – Associate or AWS Certified SysOps Administrator.
- ITIL 4 Foundation, or equivalent demonstrated knowledge of incident, change and problem practice.
Details
- Location: Navi Mumbai.
- The role operates within a shift-based infrastructure team and reports to the Manager IT Infrastructure for the assigned shift.
- The role requires documented shift handover and coordination with other shifts and infrastructure teams.
Read the full description and apply on the company’s own careers page.