C

Site Reliability Engineer

Cisco
NewPosted today

LOCATION

Bangalore · Hybrid

EXPERIENCE

8 - 10 Years

TYPE

FullTime

SALARY

Negotiable

SKILLS REQUIRED

Reliability EngineeringObservabilityInfrastructure as CodeCloud InfrastructureIncident response

Job description

Overview

Site Reliability Engineer for the Network Assurance Data Platform team, operating and scaling AWS infrastructure that supports large-scale data pipelines, ML workloads, cloud-native services, and reliability engineering practices.

What you'll do

  • Own and operate scalable, reliable, and cost-efficient infrastructure for the Network Assurance Data Platform.
  • Manage, optimize, and improve large-scale data workflows using Apache Airflow, AWS EMR, Spark, and Hadoop-based processing platforms.
  • Operate and improve Amazon EKS environments supporting containerized services, ML workloads, and production-grade platform components.
  • Build and maintain infrastructure automation using Terraform and other infrastructure-as-code practices.
  • Develop Python-based automation, integrations, operational tooling, reporting, and reliability improvements.
  • Drive FinOps practices including cost visibility, cost allocation, forecasting, anomaly detection, optimization, and governance.
  • Partner with data engineering, ML, platform, finance, and product teams to improve reliability, performance, scalability, and cost efficiency.
  • Identify infrastructure bottlenecks, performance issues, inefficient workloads, and cost optimization opportunities.
  • Improve observability, alerting, incident response, and operational readiness across data and ML platforms.
  • Support capacity planning, right-sizing, autoscaling, storage optimization, and workload efficiency across AWS services.
  • Lead technical discussions, influence design decisions, and guide teams toward reliable and cost-conscious architecture.
  • Provide senior-level technical leadership, mentorship, and operational guidance to engineers across the team.
  • Drive continuous improvement in platform reliability, automation, deployment practices, and operational excellence.
  • Operate infrastructure supporting reliable, scalable, performant, and cost-efficient NADP data and ML workloads.
  • Improve AWS cost visibility, forecasting, governance, cloud-wastage reduction, and stakeholder reporting.

What you'll need

  • Bachelor’s degree or higher in Engineering, Computer Science, or equivalent practical experience.
  • 8–10 years of relevant experience in Site Reliability Engineering, DevOps, Cloud Infrastructure, Platform Engineering, Data Infrastructure, or Production Engineering.
  • Strong hands-on experience operating production infrastructure on AWS.
  • Strong experience with Apache Airflow for workflow orchestration, pipeline operations, scheduling, monitoring, and troubleshooting.
  • Hands-on experience with AWS EMR, Spark, Hadoop, or similar large-scale data processing platforms.
  • Strong experience operating Amazon EKS or Kubernetes-based environments in production.
  • Experience supporting containerized workloads, preferably including ML workloads or data platform services.
  • Strong hands-on experience with Terraform and infrastructure-as-code practices.
  • Strong Python programming or scripting experience for automation, integrations, operational tooling, and infrastructure workflows.
  • Practical experience with cloud cost optimization, FinOps, AWS cost analysis, tagging, budgeting, forecasting, and cost governance.
  • Strong understanding of Linux systems, networking, distributed systems, and production troubleshooting.
  • Experience with observability tools such as CloudWatch, Prometheus, Grafana, Splunk, OpenSearch, Datadog, or similar platforms.
  • Experience with incident management, production support, root cause analysis, reliability improvements, and operational excellence.
  • Ability to analyze infrastructure, performance, and cost data and convert findings into clear technical recommendations.
  • Strong communication skills with the ability to collaborate across data, ML, platform, finance, and engineering teams.
  • Ability to operate independently, drive initiatives end to end, and provide technical leadership in a fast-paced environment.

Nice to have

  • Experience working with large-scale SaaS platforms or high-volume data infrastructure.
  • Experience with ML infrastructure, model execution platforms, batch processing, or data pipeline reliability.
  • Experience optimizing EMR, Spark, Airflow, EKS, storage, and compute workloads for performance and cost.
  • Experience with AWS services such as EC2, S3, RDS, IAM, VPC, CloudWatch, OpenSearch, Lambda, ElastiCache, and related cloud-native services.
  • Experience with AWS Savings Plans, Reserved Instances, Spot adoption, Graviton migration, storage lifecycle management, and workload right-sizing.
  • Experience with cloud cost management tools such as AWS Cost Explorer, AWS CUR, Cloudability, CloudHealth, Kubecost, or similar platforms.
  • Experience driving FinOps programs, cost reviews, stakeholder reporting, OKR tracking, and executive-level updates.
  • Experience with CI/CD systems, GitHub workflows, Atlantis, or similar deployment automation platforms.
  • Experience with Puppet, Ansible, Helm, Argo CD, or other configuration and deployment management tools.
  • FinOps certification or equivalent hands-on cloud financial management experience is a plus.
  • Ability to influence engineering teams toward cost-aware, scalable, and reliable design patterns.

Details

  • Location: Bangalore, India.
  • Work mode: Hybrid.

Read the full description and apply on the company’s own careers page.

Stay safe

Hiring on Abekus is free for applicants

We never charge a fee, and employers are prohibited from doing so. If a recruiter asks for payment, please report them right away.

Site Reliability Engineer