Overview
Site Reliability Engineer for the Network Assurance Data Platform team, operating and scaling AWS infrastructure that supports large-scale data pipelines, ML workloads, cloud-native services, and reliability engineering practices.
What you'll do
- Own and operate scalable, reliable, and cost-efficient infrastructure for the Network Assurance Data Platform.
- Manage, optimize, and improve large-scale data workflows using Apache Airflow, AWS EMR, Spark, and Hadoop-based processing platforms.
- Operate and improve Amazon EKS environments supporting containerized services, ML workloads, and production-grade platform components.
- Build and maintain infrastructure automation using Terraform and other infrastructure-as-code practices.
- Develop Python-based automation, integrations, operational tooling, reporting, and reliability improvements.
- Drive FinOps practices including cost visibility, cost allocation, forecasting, anomaly detection, optimization, and governance.
- Partner with data engineering, ML, platform, finance, and product teams to improve reliability, performance, scalability, and cost efficiency.
- Identify infrastructure bottlenecks, performance issues, inefficient workloads, and cost optimization opportunities.
- Improve observability, alerting, incident response, and operational readiness across data and ML platforms.
- Support capacity planning, right-sizing, autoscaling, storage optimization, and workload efficiency across AWS services.
- Lead technical discussions, influence design decisions, and guide teams toward reliable and cost-conscious architecture.
- Provide senior-level technical leadership, mentorship, and operational guidance to engineers across the team.
- Drive continuous improvement in platform reliability, automation, deployment practices, and operational excellence.
- Operate infrastructure supporting reliable, scalable, performant, and cost-efficient NADP data and ML workloads.
- Improve AWS cost visibility, forecasting, governance, cloud-wastage reduction, and stakeholder reporting.
What you'll need
- Bachelor’s degree or higher in Engineering, Computer Science, or equivalent practical experience.
- 8–10 years of relevant experience in Site Reliability Engineering, DevOps, Cloud Infrastructure, Platform Engineering, Data Infrastructure, or Production Engineering.
- Strong hands-on experience operating production infrastructure on AWS.
- Strong experience with Apache Airflow for workflow orchestration, pipeline operations, scheduling, monitoring, and troubleshooting.
- Hands-on experience with AWS EMR, Spark, Hadoop, or similar large-scale data processing platforms.
- Strong experience operating Amazon EKS or Kubernetes-based environments in production.
- Experience supporting containerized workloads, preferably including ML workloads or data platform services.
- Strong hands-on experience with Terraform and infrastructure-as-code practices.
- Strong Python programming or scripting experience for automation, integrations, operational tooling, and infrastructure workflows.
- Practical experience with cloud cost optimization, FinOps, AWS cost analysis, tagging, budgeting, forecasting, and cost governance.
- Strong understanding of Linux systems, networking, distributed systems, and production troubleshooting.
- Experience with observability tools such as CloudWatch, Prometheus, Grafana, Splunk, OpenSearch, Datadog, or similar platforms.
- Experience with incident management, production support, root cause analysis, reliability improvements, and operational excellence.
- Ability to analyze infrastructure, performance, and cost data and convert findings into clear technical recommendations.
- Strong communication skills with the ability to collaborate across data, ML, platform, finance, and engineering teams.
- Ability to operate independently, drive initiatives end to end, and provide technical leadership in a fast-paced environment.
Nice to have
- Experience working with large-scale SaaS platforms or high-volume data infrastructure.
- Experience with ML infrastructure, model execution platforms, batch processing, or data pipeline reliability.
- Experience optimizing EMR, Spark, Airflow, EKS, storage, and compute workloads for performance and cost.
- Experience with AWS services such as EC2, S3, RDS, IAM, VPC, CloudWatch, OpenSearch, Lambda, ElastiCache, and related cloud-native services.
- Experience with AWS Savings Plans, Reserved Instances, Spot adoption, Graviton migration, storage lifecycle management, and workload right-sizing.
- Experience with cloud cost management tools such as AWS Cost Explorer, AWS CUR, Cloudability, CloudHealth, Kubecost, or similar platforms.
- Experience driving FinOps programs, cost reviews, stakeholder reporting, OKR tracking, and executive-level updates.
- Experience with CI/CD systems, GitHub workflows, Atlantis, or similar deployment automation platforms.
- Experience with Puppet, Ansible, Helm, Argo CD, or other configuration and deployment management tools.
- FinOps certification or equivalent hands-on cloud financial management experience is a plus.
- Ability to influence engineering teams toward cost-aware, scalable, and reliable design patterns.
Details
- Location: Bangalore, India.
- Work mode: Hybrid.
Read the full description and apply on the company’s own careers page.