T

Sr Software AI Engineer

Tredence · Bengaluru, Karnataka, India

full_timePosted Today
Apply now →

Job description

Senior AI Engineer - Agentic AI Platform Location:Bangalore Experience:5-8 Years Hybrid - 4 days to office Employment Type:Full-Time About the Role We are building a next-generation Enterprise AI Platform that enables organizations to create, deploy, govern, and scale AI Agents, Multi-Agent Systems, AI Workflows, Knowledge Systems, Data Pipelines, and Enterprise Integrations. We are looking for exceptional engineers who are passionate about building real AI products, not demos. This role is for individuals who can combine strong software engineering fundamentals with modern AI technologies to build scalable, production-grade systems. You will work on core platform capabilities including Agent Servers, Workflow Engines, RAG Infrastructure, Model Orchestration, Tool Execution Frameworks, Enterprise Integrations, and AI Governance. This is a high-ownership role with significant growth opportunities for engineers who aspire to become Technical Leads, Principal Engineers, and Architects. What You Will Build - Production-grade AI Agents and Multi-Agent Systems. - Agent Servers capable of orchestrating complex enterprise workflows. - AI Workflow Engines supporting state management, branching, approvals, retries, checkpoints, and human-in-the-loop execution. - Enterprise-scale RAG platforms with advanced retrieval, grounding, and knowledge management capabilities. - AI Copilots for software engineering, data engineering, analytics, migration, automation, and enterprise operations. - Model orchestration layers supporting OpenAI, Claude, Gemini, and open-source models. - Tool-calling frameworks integrating databases, APIs, SaaS platforms, and enterprise systems. - Highly scalable backend services, APIs, and distributed platform components. Core Skills & Experience Software Engineering Excellence - Exceptional Python programming skills with a strong focus on performance, maintainability, and scalability. - Deep production experience with FastAPI and modern backend development. - Strong understanding of AsyncIO, concurrency, multithreading, multiprocessing, queues, event-driven systems, and high-performance architectures. - Expertise in software design patterns, clean architecture, SOLID principles, API design, testing strategies, and system design. - Proven ability to build and operate production systems handling real customer workloads. Agentic AI & Generative AI - Hands-on experience building applications using LangGraph, LangChain, OpenAI, Azure OpenAI, Anthropic Claude, Gemini, LlamaIndex, CrewAI, AutoGen, DSPy, or similar frameworks. - Strong experience designing AI Agents, Multi-Agent Systems, Workflow Agents, Tool-Calling Systems, Memory Architectures, and Autonomous Workflows. - Deep understanding of LLM application architecture, structured outputs, prompt engineering, evaluations, guardrails, model routing, and reasoning workflows. - Experience taking AI applications from proof-of-concept to enterprise production deployments. RAG, Search & Knowledge Systems - Strong production experience with Retrieval-Augmented Generation (RAG). - Expertise in embeddings, chunking strategies, hybrid retrieval, semantic search, reranking, metadata filtering, context optimization, and grounding techniques. - Experience with Pinecone, Qdrant, Weaviate, Chroma, pgvector, Azure AI Search, Elasticsearch, OpenSearch, or similar technologies. - Understanding of enterprise knowledge systems and large-scale document processing pipelines. Data, Cloud & Infrastructure - PostgreSQL expertise is mandatory, including schema design, indexing, query optimization, and performance tuning. - Experience with Redis, MongoDB, Cosmos DB, DynamoDB, BigQuery, Snowflake, Databricks, or similar platforms. - Strong hands-on experience with Docker, Kubernetes, Helm, CI/CD, and cloud-native architectures. - Experience deploying and operating solutions on Azure, AWS, or GCP. - Familiarity with Redis Pub/Sub, Kafka, RabbitMQ, Azure Service Bus, or distributed messaging systems. Production Operations - Experience with OpenTelemetry, Grafana, Prometheus, ELK, Datadog, Application Insights, or similar observability platforms. - Strong understanding of monitoring, tracing, performance tuning, fault tolerance, scalability, reliability, and operational excellence. - Ability to diagnose and resolve complex production issues across application, infrastructure, and AI layers. Highly Desirable - MCP (Model Context Protocol) ecosystem. - AI Workflow Platforms and Orchestration Engines. - Low-Code/No-Code Platform Development. - Enterprise Integration Platforms. - Knowledge Graphs and Neo4j. - MLOps Platforms. - Data Engineering Platforms. - Multi-tenant SaaS architectures. - AI Governance and Evaluation Frameworks. Who We Are Looking For - Engineers who love building and shipping products. - Individuals who can quickly understand complex systems and contribute meaningful solutions. - Developers who take ownership and can independently drive features from design to production. - Strong problem solvers who are equally comfortable discussing architecture, writing code, optimizing performance, and troubleshooting production systems. - People who are hungry to learn, grow rapidly, and take on increasing technical leadership responsibilities. What Success Looks Like Within your first year, you will own critical platform capabilities such as Agent Servers, Workflow Engines, RAG Infrastructure, Model Gateways, Enterprise Integrations, or AI Runtime Services. You will play a key role in shaping the architecture of a modern enterprise AI platform while working on some of the most challenging problems in Agentic AI, workflow orchestration, and large-scale AI systems. If you are the kind of engineer who enjoys building complex systems, writing excellent code, solving hard problems, and turning cutting-edge AI capabilities into real enterprise products, we would love to hear from you. **Requir