Overview
The Systems Engineer (Linux) is the senior technical authority for the Institute's Linux estate, responsible for configuration, hardening, patching, automation, cloud infrastructure, application platform support, resilience and day-to-day operations. The role works a swing shift within a multidisciplinary infrastructure team, supports a globally distributed estate and provides technical leadership to less experienced engineers.
What you'll do
- Provision, configure, harden and maintain Linux servers across Red Hat Enterprise Linux, Rocky, Oracle Linux, SUSE and Ubuntu on premises and in cloud through decommissioning.
- Maintain standardized automated build pipelines and golden images with parity between on-premises and cloud builds.
- Administer filesystems and LVM, storage presentation and multipath, networking, system services, kernel parameters, performance tuning, user and Sudo access, SSH and certificate management.
- Administer Linux workloads on VMware, Nutanix and physical server hardware, including firmware currency and OEM support status.
- Integrate Linux hosts with Active Directory or LDAP through SSSD, Kerberos, centralized authentication and privileged access controls.
- Manage Linux subscriptions, repositories, packages, repository mirroring and content lifecycle promotion across development, test and production.
- Provide operational support for Linux instances during the assigned shift to users and application teams in every region.
- Build, maintain and support Linux-hosted web and application servers, middleware, containers and runtime dependencies with application owners.
- Support database hosting on Linux at the platform layer, including filesystem layout, kernel and resource tuning and backup interfaces.
- Act as second- and third-level escalation for complex Linux, application platform and performance issues, troubleshooting to root cause across operating system, storage, network and application layers and liaising with vendors.
- Monitor system health and capacity, tune alerting and respond to degradation before it becomes an incident.
- Support application deployment and release activity, including environment preparation, pre- and post-deployment validation and rollback.
- Execute recurring patching cycles and remediate enterprise vulnerability-management findings through closure, with validation, exception governance and compliance reporting.
- Maintain CIS or equivalent hardening baselines and use Ansible to enforce and remediate configuration drift.
- Manage kernel updates, live patching, reboot orchestration and cluster-aware patching.
- Follow change control, rollback planning and maintenance-window coordination, and support audits with documented evidence.
- Build and maintain Ansible playbooks, roles, collections and inventories for provisioning, hardening, patching, application deployment and drift remediation.
- Operate Ansible Automation Platform or AWX where deployed, including job templates, workflows, credentials, scheduling, inventory sources and role-based access.
- Maintain automation code in version control with peer review, testing and a documented promotion path across environments.
- Extend automation into cloud provisioning and infrastructure-as-code, and integrate with CI/CD pipelines where used by application teams.
- Develop Bash and Python scripting where Ansible is not the right tool, and quantify manual effort removed by each automation delivery.
- Administer Linux instances in AWS and Microsoft Azure, including compute, storage, images, virtual networking, security groups, resource tagging, IAM roles, instance profiles, secrets and bastion or session-manager access.
- Maintain build, hardening, patching and monitoring parity between cloud and on-premises Linux instances.
- Support assessment, migration and rollback of Linux workloads between on premises and cloud.
- Monitor and optimize cloud consumption through right-sizing, commitment planning, removal of orphaned volumes and snapshots, and budget variance reporting.
- Configure backup protection at build, execute and validate restores, and participate in restoration testing.
- Maintain high availability and clustering, including Pacemaker and Corosync, load balancing and replication.
- Participate in disaster recovery drills and maintain recovery runbooks and recovery time and recovery point position for in-scope Linux platforms.
- Provide technical leadership and mentoring to junior and mid-level engineers.
- Contribute to infrastructure standards, reference builds and technology evaluation, and lead root cause analysis for critical Linux incidents through permanent fix.
- Coordinate with network, security, database and application teams on infrastructure initiatives and cross-domain troubleshooting.
- Complete and verify documented shift handovers covering open incidents, in-flight changes, maintenance and pending actions.
- Maintain runbooks and technical documentation standards for the Linux estate.
What you'll need
- Bachelor’s degree in computer science, Information Technology, Engineering or a closely related field, or equivalent professional experience.
- 7+ years of experience in Linux systems engineering or administration, with substantial hands-on responsibility for a production Linux estate.
- Advanced administration of Red Hat Enterprise Linux or a comparable enterprise distribution, including build, hardening, filesystems and LVM, systemd, networking, kernel tuning and performance troubleshooting.
- Hands-on experience building and maintaining Ansible playbooks, roles and inventories for provisioning, hardening, patching and drift remediation at estate scale.
- Demonstrated experience holding automation code in version control with peer review and a promotion path across environments.
- Hands-on administration of Linux instances in AWS or Microsoft Azure, including compute, storage, virtual networking, IAM roles and instance profiles, and secret management.
- Experience maintaining build and patching parity between on-premises and cloud Linux instances.
- Ownership of recurring Linux patching cycles, including validation, reboot orchestration, exception governance and compliance reporting against defined standards.
- Experience remediating findings from an enterprise vulnerability management program, including tracking to closure and documented risk acceptance.
- Experience with CIS benchmarks or equivalent hardening baselines, and detection and remediation of configuration drift.
- Experience supporting Linux-hosted application services such as Apache, NGINX, Tomcat, JBoss or comparable middleware, including build and troubleshooting.
- Experience integrating Linux into Active Directory or LDAP, including SSSD, Kerberos and centralized authentication.
- Experience administering Linux workloads on VMware and Nutanix, and on physical server hardware.
- Experience with backup and restore of Linux systems and application data, and participation in disaster recovery drills and restoration testing.
- Strong Bash and Python scripting ability.
- Experience operating within a formal change management process, including Change Advisory Board submission, rollback planning and post-implementation review.
- Experience working within a shift-based infrastructure team supporting a globally distributed estate, including structured handover.
- Professional proficiency in spoken and written English, sufficient to coordinate with global infrastructure and application teams and communicate with senior stakeholders.
Nice to have
- Experience operating Ansible Automation Platform or AWX, including job templates, workflows, credentials and scheduling.
- Experience with high availability and clustering for Linux workloads, such as Pacemaker and Corosync, load balancing or replication.
- Experience with containers and container platforms such as Docker, Podman, Kubernetes or OpenShift.
- Experience with Red Hat Satellite or an equivalent subscription, repository and content lifecycle management tool.
- Certification such as Red Hat Certified Engineer (RHCE) or Red Hat Certified System Administrator (RHCSA).
- Certification such as an AWS or Microsoft Azure associate-level credential, or Red Hat Certified Specialist in Ansible Automation.
- ITIL 4 Foundation, or equivalent demonstrated knowledge of incident, change and problem practice.
Details
- Location: Navi, Mumbai.
- Work a swing shift within a multidisciplinary infrastructure team.
- Support users and application teams across multiple regions and time zones.
- Complete structured handovers at shift boundaries.
Read the full description and apply on the company’s own careers page.