C

Applied AI Research Engineer

Chargebee · Chennai, Tamil Nadu, India

2–8 yrs experiencefull_timePosted 2 days ago
Apply now →

Job description

**About Chargebee** Chargebee is building AI-native billing and monetization infrastructure for modern, high-growth businesses. We serve transformative, category-defining customers such as Lambda, Conde Nast, Gorgias, HeyGen, Zapier, and CodeRabbit. **About the team** Chargebee AI Labs is a Chennai-based team of AI engineers, model behaviour researchers, and deployed-AI architects. This role sits within that effort, closing the gap between what general-purpose language models can do and what billing actually requires: accurate answers, traceable reasoning, and auditable calculations. We're already running AI systems in production across real-world billing and revenue workflows, which gives us direct visibility into both the capabilities and limitations of current models. **About the role** We are looking for an Applied AI Research Engineer who is interested in solving open-ended problems at the intersection of large language models and complex billing and revenue workflows. You will investigate limitations observed in production AI agents, study existing research and techniques, design experiments, build prototypes, and help convert successful approaches into production-ready systems. This role is suitable for an engineer who enjoys both experimentation and engineering: someone who can read a paper, reproduce or adapt an idea, evaluate whether it works, and then build the surrounding system needed to use it reliably in production. **What You Will Work On** *Domain-Specialized Billing Intelligence* General-purpose language models don't understand Chargebee's billing concepts, entities, workflows, and business rules natively — that context has to be reconstructed for every request. Today, we improve their performance using techniques such as retrieval-augmented generation, knowledge-base traversal, domain-specific prompting, and tool-assisted retrieval. These approaches improve accuracy but also add latency, token cost, and orchestration complexity. You will explore ways to reduce how much of that domain context must be reconstructed for every request. This may include: - Fine-tuning and parameter-efficient model adaptation - Model distillation - Domain-specific embeddings and representations - Synthetic training-data generation - Context compression - Improved retrieval and knowledge-representation techniques - Smaller specialized models - Structured domain models and ontologies - Hybrid model, retrieval, and deterministic approaches The goal is not to use a specific technique. The goal is to determine which approach produces the best balance of accuracy, latency, reliability, and cost. *Verifiable Reasoning over Financial Data* For financial questions, producing a plausible answer is not enough. The answer must use the right records, apply the correct business meaning, calculate the result accurately, and provide a traceable explanation. You will help build systems where language models interpret user requests and generate structured execution plans, while deterministic components perform data retrieval, filtering, transformations, and calculations. You may work on: - Understanding and classifying financial and billing questions - Generating structured query or execution plans - Natural-language-to-SQL or natural-language-to-code systems - Planner–executor architectures - Deterministic calculation engines - Validation of generated queries and execution plans - Data and calculation lineage - Evidence-backed answers - Semantic and numerical correctness evaluations - Detection of ambiguity and unsupported assumptions *Production AI Reliability* You will study actual failures from production AI systems, including: - Incorrect interpretation of billing terminology - Hallucinated product behaviour or business rules - Incorrect tool selection - Incorrect filters, joins, and aggregations - Inconsistent answers - Retrieval failures - Excessive latency or token consumption - Model regressions - Weaknesses in current evaluation methods You will convert these observations into measurable research questions and experiments. **What we are looking for** - Approximately 3–5 years of software engineering, machine learning, or applied AI experience. - Experience building production AI, ML, NLP, search, or data-intensive systems. - Good understanding of machine-learning and deep-learning fundamentals. - Hands-on experience with large language models, retrieval systems, tool calling, structured generation, or AI agents. - Strong Python programming skills. - Experience with at least one ML framework such as PyTorch, TensorFlow, JAX, or Hugging Face. - Ability to design experiments and evaluate results objectively. - Ability to read technical papers and translate ideas into working prototypes. - Strong software-engineering fundamentals, including system design, APIs, testing, and debugging. - Comfort working on problems where the solution is not yet known. - Curiosity about how and why AI systems fail in real-world environments. **Helpful experience** - Fine-tuning or adapting open-source language models - Building evaluation frameworks for LLMs or AI agents - Natural-language-to-SQL or code-generation systems - Knowledge graphs, ontologies, semantic layers, or domain-specific languages - Model serving, inference optimisation, distillation, or quantisation - Payments, billing, accounting, ERP, fintech, or financial systems - Meaningful open-source contributions, technical writing, research projects, or publications **Education** A bachelor's or master's degree in computer science, machine learning, data science, mathematics, or a related discipline is helpful but not mandatory. A PhD is not required. We value demonstrated engineering ability, strong fundamentals, experimental thinking, and evidence that you can learn and apply new techniques.