Overview
Develop and operate custom software solutions across application and infrastructure layers, owning customer environments end-to-end. Focus on reliability, scalability, operational efficiency, automation and AI-assisted troubleshooting while partnering with SRE on platform reliability.
What you'll do
- Own customer environments across application and infrastructure layers, including configurations, deployments, patching and scaling.
- Execute and automate runbook-driven tasks, including SQL scripts, queue management, data fixes and restores.
- Build automation using PowerShell, Python and GitHub Actions to convert manual processes into self-service workflows.
- Use AI tools such as Copilot and Claude AI for code generation, log analysis and faster root cause analysis.
- Monitor system health using Datadog APM, logs, dashboards and monitors.
- Respond to alerts using PagerDuty, Opsgenie and ServiceNow and ensure timely incident resolution.
- Troubleshoot and resolve incidents across application and infrastructure layers.
- Partner with SRE to ensure overall platform reliability.
- Participate in a 24x7 on-call rotation.
- Document processes and mentor junior engineers.
- Continuously identify and eliminate operational toil through automation.
What you'll need
- BTECH.
- Minimum 5 year(s) of experience.
- Experience in systems, operations or platform engineering roles.
- Hands-on experience across both application and infrastructure domains.
- Proven experience in production support, monitoring and incident management.
- An automation mindset and the ability to automate repetitive tasks and improve processes.
- Windows Server experience, including RDP, IIS, Services and Task Scheduler.
- Linux experience, including SSH, systemctl, bash and cron.
- Strong SQL skills, including SSMS, backups, restores, troubleshooting and data fixes.
- Hands-on experience with Datadog, including APM, logs, dashboards and monitors.
- Experience with on-call or incident tools such as PagerDuty, Opsgenie or ServiceNow.
- Strong alert triaging and incident handling skills.
- Experience with XML, JSON and property files.
- Experience with PowerShell, Python or Bash.
- Comfort using tools such as Copilot to write and debug code and accelerate troubleshooting.
- Cloud fundamentals, including AWS/Azure basics.
- Experience with server resizing and patching.
- Networking basics, including TLS/certificates and load balancing.
- Experience with web servers such as IIS and Apache Tomcat.
- Git and GitHub workflows.
- Ability to troubleshoot under pressure across application and infrastructure layers.
- Strong attention to detail and the ability to follow and improve runbooks.
- Strong interest in using AI to reduce operational effort and improve system reliability.
Nice to have
- Experience with infrastructure automation tools such as Terraform and Ansible.
- Familiarity with AI/ML-based observability or predictive alerting.
- Experience in teams that blend engineering and operations responsibilities.
Details
- Location: Hyderabad.
- Participate in a 24x7 on-call rotation.
Read the full description and apply on the company’s own careers page.