Overview
The Systems Engineer (Linux) is the senior technical authority for the Institute's Linux estate, responsible for configuration, hardening, patching and day-to-day operation of Linux servers and cloud across on-premises infrastructure and AWS and Microsoft Azure. The role works a shift within a multidisciplinary infrastructure team, supports global users and application teams, provides technical leadership and mentoring, and acts as a second- and third-level escalation point.
What you'll do
- Provision, configure, harden and maintain Linux servers across Red Hat Enterprise Linux, Rocky, Oracle Linux, SUSE and Ubuntu throughout the lifecycle from build through decommission.
- Maintain standardized, automated build pipelines and golden images with parity between on-premises and cloud builds.
- Administer filesystems and LVM, storage presentation and multipath, networking, system services, kernel parameters and performance tuning, user and Sudo access, and SSH and certificate management.
- Administer Linux workloads on VMware, Nutanix and physical server hardware, including firmware currency and OEM support status for physical hosts.
- Integrate Linux hosts with enterprise directory and access management, including Active Directory or LDAP integration through SSSD, Kerberos, centralized authentication and privileged access controls.
- Manage Linux subscription, repository and package estates, including Red Hat Satellite or equivalent, repository mirroring and content lifecycle promotion across development, test and production.
- Provide operational support for Linux instances during the assigned shift and support users and application teams in every region.
- Build, maintain and support application services hosted on Linux, including Apache, NGINX, Tomcat, JBoss, middleware, containers and runtime dependencies, in coordination with application owners.
- Support database hosting on Linux at the platform layer, including filesystem layout, kernel and resource tuning, and backup interfaces, in coordination with database owners.
- Act as second- and third-level escalation for complex Linux, application platform and performance issues, troubleshooting to root cause across operating system, storage, network and application layers and liaising with vendors.
- Monitor system health and capacity through monitoring tooling and manual checks, and tune alerting so engineers respond to degradation before it becomes an incident.
- Support application deployment and release activity on Linux, including environment preparation, pre- and post-deployment validation and rollback.
- Execute recurring patching cycles for the global Linux estate and remediate findings from the enterprise vulnerability management program through closure.
- Maintain hardening baselines such as CIS benchmarks and use Ansible to enforce and remediate configuration drift.
- Manage kernel update and live-patching strategy, reboot orchestration and cluster-aware patching.
- Perform structured pre-patch and post-patch validation and document the outcome of every cycle.
- Follow and enforce change control, rollback planning and maintenance-window coordination, and support internal and external audits with documented evidence.
- Build and maintain Ansible playbooks, roles, collections and inventories for provisioning, hardening, patching, application deployment and configuration-drift remediation.
- Operate Ansible Automation Platform or AWX where deployed, including job templates, workflows, credentials, scheduling, inventory sources and role-based access.
- Hold automation code in version control with peer review, testing and a documented promotion path across environments.
- Extend automation into cloud provisioning and infrastructure-as-code, and integrate with CI/CD pipelines where used by application teams.
- Develop Bash and Python scripting where Ansible is not the right tool, and quantify the manual effort removed by each automation delivered.
- Administer Linux instances in AWS and Microsoft Azure, including compute, block and object storage, images, virtual networking, security groups and resource tagging.
- Maintain build, hardening, patching and monitoring parity between cloud and on-premises Linux instances.
- Manage cloud identity and access for Linux hosts, including IAM roles and instance profiles, key and secret management, and bastion or session-manager access patterns.
- Support assessment, migration and rollback of Linux workloads between on-premises and cloud, including identity, network, storage and protection consequences.
- Monitor and optimize cloud consumption for Linux infrastructure, including right-sizing, commitment planning, removal of orphaned volumes and snapshots, and reporting variance against budget.
- Ensure Linux workloads are protected at build with correctly configured backup agents, policies and application-consistent handling.
- Execute and validate restores of Linux systems and applications and participate in scheduled restoration testing with the Storage and Backup Engineer.
- Maintain high availability and clustering for Linux workloads where required, including Pacemaker and Corosync, load balancing and replication.
- Participate in disaster recovery drills as an executing engineer, maintain recovery runbooks, and maintain recovery time and recovery point position for Linux platforms in scope.
- Provide technical leadership and mentoring to junior and mid-level engineers on Linux administration, automation practice and operational discipline.
- Contribute to infrastructure standards, reference builds and technology evaluation, and lead root cause analysis for critical Linux incidents through post-incident review to permanent fix.
- Coordinate with network, security, database and application teams on infrastructure initiatives and complex cross-domain troubleshooting.
- Complete a documented handover at the end of each shift covering open incidents, in-flight changes, running maintenance and pending actions, and verify the handover received at shift start.
- Maintain runbooks, technical documentation and documentation standards for the Linux estate.
What you'll need
- A Bachelor's degree in computer science, Information Technology, Engineering or a closely related field, or equivalent professional experience.
- 7+ years of experience in Linux systems engineering or administration, with substantial hands-on responsibility for a production Linux estate.
- Demonstrated advanced administration of Red Hat Enterprise Linux or a comparable enterprise distribution, including build, hardening, filesystems and LVM, systemd, networking, kernel tuning and performance troubleshooting.
- Demonstrated hands-on experience building and maintaining Ansible playbooks, roles and inventories for provisioning, hardening, patching and drift remediation at estate scale.
- Demonstrated experience holding automation code in version control with peer review and a promotion path across environments.
- Demonstrated hands-on administration of Linux instances in AWS or Microsoft Azure, including compute, storage, virtual networking, IAM roles and instance profiles, and secret management.
- Demonstrated experience maintaining build and patching parity between on-premises and cloud Linux instances.
- Demonstrated ownership of recurring Linux patching cycles, including validation, reboot orchestration, exception governance and compliance reporting against defined standards.
- Demonstrated experience remediating findings from an enterprise vulnerability management program, including tracking to closure and documented risk acceptance.
- Experience with CIS benchmarks or equivalent hardening baselines, and with detection and remediation of configuration drift.
- Demonstrated experience supporting application services hosted on Linux, such as Apache, NGINX, Tomcat, JBoss or comparable middleware, including build and troubleshooting.
- Demonstrated experience with Linux integration into Active Directory or LDAP, including SSSD, Kerberos and centralized authentication.
- Experience administering Linux workloads on VMware and Nutanix, and on physical server hardware.
- Experience with backup and restore of Linux systems and application data, and participation in disaster recovery drills and restoration testing.
- Strong Bash and Python scripting ability.
- Experience operating within a formal change management process, including Change Advisory Board submission, rollback planning and post-implementation review.
- Experience working within a shift-based infrastructure team supporting a globally distributed estate, including structured handover.
- Professional proficiency in spoken and written English, sufficient to coordinate with global infrastructure and application teams and communicate with senior stakeholders.
Nice to have
- Experience operating Ansible Automation Platform or AWX, including job templates, workflows, credentials and scheduling.
- Experience with high availability and clustering for Linux workloads, such as Pacemaker and Corosync, load balancing or replication.
- Experience with containers and container platforms such as Docker, Podman, Kubernetes or OpenShift.
- Experience with Red Hat Satellite or an equivalent subscription, repository and content lifecycle management tool.
- Certification such as Red Hat Certified Engineer (RHCE) or Red Hat Certified System Administrator (RHCSA).
- Certification such as an AWS or Microsoft Azure associate-level credential, or
Read the full description and apply on the company’s own careers page.