L

Senior AI production Engineer

L&T Finance · Mumbai, Maharashtra, India

full_timePosted 2 days ago
Apply now →

Job description

**Position: Senior AI Production Engineer** **Function: AI Scaling Platform** **Location: Mumbai – Santacruz (Kalina) / Mahape (Navi Mumbai)** **Role Overview** As a Senior AI Production Engineer – you will own the end-to-end implementation of scaling GenAI and Agentic AI use cases from PoC to large-scale production. You’ll bring hands-on experience on LLM’s , Agentic AI solutions, microservices , cloud , DevOps and system build experience. You will work closely with business leadership, AI research teams, engineering, infosec, compliance, and vendors to deliver secure, scalable, and production-grade AI systems across onboarding, copilots, servicing, collections, and customer engagement. **Key Responsibilities :** - Own the AI Scaling Program, driving GenAI and Agentic AI use cases from PoC to production. - Leading conversation with business leadership, and AI/engineering, infosec, Architecture Team , DC/Infrastructure Team to finalise the solutions, approach . - Design, build, and deploy production-grade GenAI applications using Python (FastAPI), Java backend services, and React.js frontends. - Design and Build event-driven, asynchronous microservices supporting high concurrency and low latency. - Develop LLM-powered systems including chatbots, voicebots, copilots, and agentic workflows. - Fine-tuning (SFT, LoRA) - Prompt engineering (ReAct, CoT, tool calling) - Hands-on experience with multiple LLMs: GPT-4.x, Gemini Pro, Claude, LLaMA (open & closed models). - Configuring AI projects into modular and deployable assets. - Build multi-agent systems using frameworks such as LangGraph, ADK, Pydantic AI, supervisor-agent and hierarchical patterns. - Enable stateful workflows, long-term memory, retries, conditional execution, and orchestration (LangGraph, n8n). - Building project plans , aligning resources and leading the Jira user stories to accomplish the projects. - Ensures all approvals related to project go live are taken into account. - Dealing and seeking approvals with stakeholders like Audit, compliance and regulatory before going live. - Support post–go-live operations, incident triage, and continuous improvement. - Implement CI/CD pipelines, containerization (Docker), artifact registries, and automated rollbacks. - Ensure observability through logging, metrics, tracing, health checks, and drift monitoring. - Optimize inference pipelines using GPU acceleration (multi-GPU setups) for performance and cost efficiency. - Work closely with Infosec, Infrastructure, Data Ops, and Cloud Ops teams on BAU Operations. - Lead partner onboarding, technical documentation, and closure of InfoSec and risk processes. **Key Qualifications:** - Experience: 10-15 years of professional experience, with at least 3 years focused on MLOps, ML Engineering, or a related field in a production environment. - Cloud & Infrastructure: Strong hands-on experience with at least one major cloud provider (AWS, GCP, Azure). Cloud certification is a plus. - Containerization & Orchestration: Expertise in Docker and Kubernetes. **Technical Expertise** - Strong hands-on experience with Python (FastAPI), Java (backend services), and React.js. - Proven expertise in LLMs, RAG architectures, and vector databases including FAISS, Pinecone, Chroma, Weaviate, and GCP Vector Search. - Hands-on experience with LangChain, LangGraph, LlamaIndex, Pydantic AI, and ADK (Agent Development Kit). - Solid understanding of API design, data flows, and microservices architecture. - Strong experience with cloud-native system design on Google Cloud and Azure Platform . - Working knowledge of databases including BigQuery, Redis, SQL/NoSQL, and ORMs such as SQLAlchemy. **Leadership Skills:** - Excellent communication and stakeholder management skills, with the ability to engage business, technology, and leadership teams. - Strong capability to translate complex AI and technical concepts into measurable business outcomes. - Experience working in Agile / Scrum delivery models, using tools such as JIRA, Confluence, MS Visio, and Lucidchart. - High ownership mindset with strong problem-solving, decision-making, and execution skills. **Educations:** - Bachelor’s or Master’s degree in Computer Science, PhD, Software Engineering, Data Science, Artificial Intelligence, or MCA . - Additional certifications in GenAI, Agentic AI, or Cloud Platforms (GCP/Azure) will be an added advantage. **substitute for hands-on experience.** **Why should you work with us?** At LTF, our footprint spans both rural and urban India, serving customers from remote villages to bustling cities. Here, you’ll solve real-world problems that impact millions—helping empower rural communities while shaping the financial future of Gen Z. We are truly building an AI Scaling program that will work in Production for Voice,Agentic ,Workflow orchestration to solve many Banking and NBFC problems using Modern Technology. Join the AI-Driven Transformation Be part of a technology transformation powered by advanced machine learning and AI. You will leverage Many AI Platforms,Toolings LLM,SLMS,RAG from GCP AI and Agentic System,Azure AI Foundry, Many open source like Hugging Face etc. and also modern machine learning tools including Predictive and Prescriptive Modeling, Time Series Modeling, Transformers and Computer Vision - to drive innovation across credit underwriting, portfolio monitoring, financial fraud detection,marketing models, collection and settlement models.Recognition and Leadership Artificial Intelligence is redefining the financial landscape - driving smarter decisions, deeper insights, and inclusive growth. At LTF, we’re at the forefront of this change. Through our flagship RAISE (2024, 2025 editions) initiative, we celebrate AI’s transformative role in financial services alongside industry peers. We’re proud to be recognized as a Great Place to Work®, where innovation, impact, and purpose come together.