Overview
Work across Generative AI, model fine-tuning, optimization, and on-device deployment for iOS and Android. Take models beyond experimentation and make them work on consumer devices under constraints such as latency, memory, battery, thermal limits, and model size.
What you'll do
- Work on image generation, VTON, identity-preserving and reference-based generation.
- Perform LoRA training, fine-tuning, personalization, and adapter-based techniques.
- Conduct PyTorch-based training and experimentation.
- Work with diffusion and Transformer-based image generation architectures.
- Prepare datasets and build training pipelines and evaluation processes.
- Convert and deploy models using ONNX and ONNX Runtime.
- Optimize inference using TensorRT.
- Apply quantization using FP16, INT8, INT4, and other optimization techniques.
- Deploy on-device inference on iOS using Core ML and Android using TFLite/LiteRT, ONNX Runtime, or similar runtimes.
- Profile and debug performance across CPU, GPU, NPU, and ANE.
- Reduce inference latency, peak memory usage, and model footprint.
- Build production-ready model pipelines from training to device deployment.
What you'll need
- 2-5 Years of Experience.
- Strong experience with Python, PyTorch, and deep learning.
- Hands-on experience with Computer Vision and Generative AI.
- Experience training or fine-tuning image-generation models.
- Good understanding of LoRA, model optimization, and quantization.
- Strong knowledge of inference runtimes such as ONNX, TensorRT, Core ML, TFLite/LiteRT, or ONNX Runtime.
- Strong debugging, profiling, and software engineering fundamentals.
Nice to have
- Experience with FLUX, Stable Diffusion, SDXL, ControlNet, IP-Adapter, VAE, CLIP, or similar architectures.
- CUDA / GPU optimization.
- Metal / Apple Neural Engine.
- Android NNAPI / GPU delegates.
- C++ / Swift / Kotlin.
- Distributed or multi-GPU training.
- Model compression, pruning, and knowledge distillation.
Details
- Location: Bangalore.
- Work across iOS and Android on-device deployment.
- Work across the end-to-end flow from Data to Training, Fine-tuning, Evaluation, Quantization, Model Conversion, Runtime Optimization, On-Device Deployment, and Production Monitoring.
Read the full description and apply on the company’s own careers page.