Overview
Lead Software Engineer responsible for the strategy, architecture, and execution of next-generation AIOps capabilities across the enterprise. Drive intelligent, autonomous operations using AI/ML, observability, automation, and event-driven architectures to minimize manual intervention, improve resilience, and enable self-healing systems.
What you'll do
- Lead complex technology initiatives, including companywide initiatives with broad impact.
- Develop standards and companywide best practices for engineering complex and large-scale technology solutions.
- Design, code, test, debug, and document projects and programs.
- Review and analyze complex, large-scale technology solutions for tactical and strategic business objectives, the enterprise technological environment, and technical challenges requiring in-depth evaluation of multiple factors.
- Make decisions in developing standard and companywide best practices for engineering and technology solutions.
- Influence and lead technology teams to meet deliverables and drive new initiatives.
- Collaborate and consult with key technical experts, senior technology teams, and external industry groups to resolve complex technical issues and achieve goals.
- Lead projects and teams, or serve as a peer mentor.
- Define and deliver intelligent, autonomous operations by leveraging AI/ML, observability, automation, and event-driven architectures.
- Partner with senior engineering, platform, SRE, and business leaders to accelerate AIOps adoption and embed intelligence into production ecosystems.
- Deliver measurable improvements in availability, efficiency, and operational risk reduction.
What you'll need
- 5+ years of Software Engineering experience, or equivalent demonstrated through work experience, training, military experience, education, or a combination of these.
Nice to have
- Deep expertise in AIOps platforms and tools, such as Prometheus, AppDynamics, Splunk, ITRS Geneos, BigPanda, and OpenTelemetry ecosystems.
- Strong experience with AI/ML for IT operations, including anomaly detection, event correlation, forecasting, and intelligent alerting.
- Hands-on experience with automation frameworks, such as Ansible, Terraform, or similar, and event-driven architectures.
- Strong understanding of SRE principles, SLIs/SLOs, error budgets, and reliability engineering practices.
- Experience building self-healing systems and closed-loop remediation workflows.
- Proficiency in cloud platforms and cloud-native architectures, including Kubernetes and microservices.
- Knowledge of data pipelines, streaming platforms such as Kafka, and telemetry ingestion and processing.
- Familiarity with GenAI/LLM-assisted operations, including incident summarization, knowledge mining, and automated runbook generation.
- Ability to operate across complex organizational structures with strong stakeholder management and communication skills.
- Proven ability to define target-state architecture, operating models, and actionable roadmaps.
Details
- Location: Bengaluru, India.
- Ability to manage multiple high-complexity engineering initiatives with significant enterprise impact.
- Strong analytical, problem-solving, and architectural design skills.
- Excellent communication and documentation skills, including Confluence, Git, and architecture diagrams.
- Comfortable driving transformation and influencing senior leadership in a fast-paced, evolving environment.
- This position is not open to visa initiation or transfer.
Read the full description and apply on the company’s own careers page.