Overview
Lead complex machine-learning platform work from business context through production outcomes, combining hands-on engineering, architectural judgment, production ownership, and mentorship. Build software foundations that enable data scientists and machine learning engineers to develop, validate, deploy, and operate models safely and efficiently in production.
What you'll do
- Lead the design, implementation, deployment, and operation of complex platform capabilities that support the machine-learning lifecycle.
- Design scalable systems for model deployment and serving, feature computation, data validation, workflow orchestration, monitoring, and production support.
- Build reusable platform components, standards, and automation while improving reliability, scalability, observability, and operational readiness across ML systems.
- Partner with Data Science, Product, Risk, Fraud, and Engineering teams to translate ambiguous business needs into practical technical solutions.
- Own significant technical initiatives end to end and lead complex production investigations.
- Clarify requirements, lead design, deliver safely, measure results, and implement durable fixes.
- Raise engineering quality through thoughtful code and design reviews.
- Mentor engineers through pairing and technical guidance.
- Contribute to platform technical strategy.
- Use AI tools thoughtfully to work smarter, reduce manual work, and make better decisions faster.
What you'll need
- 5+ years of professional software-engineering experience building and operating production systems.
- Strong programming skills in Python.
- Deep knowledge of API design, distributed systems, service architecture, testing, and software-design fundamentals.
- Experience designing and operating reliable, scalable cloud-native services or data-intensive systems.
- Experience in ML platforms, model serving, feature engineering, data platforms, MLOps, or related systems.
- Practical experience with SQL, databases, caching, and tracing data or requests across multiple systems.
- Familiarity with the ML lifecycle from feature engineering and model training through evaluation, deployment, serving, monitoring, and governance.
- Experience with Docker, Kubernetes, CI/CD, Git-based workflows, and production observability.
- Ability to evaluate technical alternatives and clearly communicate recommendations, trade-offs, risks, and dependencies.
- Ability to mentor engineers, improve engineering practices, lead through technical influence, and connect engineering work to member, business, product, and operational outcomes.
Nice to have
- Experience with AWS, Databricks, MLflow, SageMaker, Spark/PySpark, Kafka, or streaming systems.
- Experience building feature platforms, model-serving systems, workflow-orchestration platforms, or internal developer platforms.
- Familiarity with monitoring model quality, data quality, feature coverage, latency, reliability, or cost efficiency.
- Experience working with sensitive data in a regulated or financial-services environment.
- Experience with Java or another backend language in addition to Python.
- Familiarity with AI tools such as ChatGPT or Copilot to improve personal productivity.
Details
- Remote position that must be performed from within India.
Read the full description and apply on the company’s own careers page.