Overview
Support Engineer focused on incident resolution, root-cause analysis, and performance optimization for an enterprise platform and autonomous AI agent workflows.
What you'll do
- Drive incident resolution within strict SLAs for critical outages.
- Perform root-cause analysis and performance troubleshooting using monitoring/telemetry tools.
- Debug cloud-hosted AI/ML service pipelines and LLM orchestration layers.
- Carry out hands-on AI evaluation of LLM outputs, conversation logs, and system traces.
- Review and update prompts, workflows, and system logic based on failure-mode findings.
- Manage high-severity (P1) support queues and prioritize critical incidents.
What you'll need
- Minimum 3 years in technical support engineering, platform operations, or escalation management for enterprise SaaS.
- Minimum 2 years reviewing/analyzing/troubleshooting LLM pipelines, prompt/tool-calling structures, or conversation traces/logs.
- Minimum 2 years monitoring/debugging/troubleshooting services on public cloud infrastructure (AWS, GCP, or Azure).
- Minimum 2 years using monitoring/observability tools such as Grafana, Kibana, Datadog, or Prometheus.
- Minimum 2 years using programming languages and writing/executing SQL queries for data analysis and tuning.
- Minimum 2 years analyzing API structures and data formats (JSON or XML).
- Minimum 2 years troubleshooting operating systems (Linux or Windows) and cloud networking components.
Details
- Location: Pune, India.
- Includes weekend on-call rotation.
Read the full description and apply on the company’s own careers page.