Search 100,000+ live jobs across India

Free to search · AI fit score against your CV · tailor your résumé in one click

I

AI / Senior AI Engineer

Innovaccer · Noida, Uttar Pradesh, India

6–15 yrs experiencevolunteerPosted 4 days ago

Job description

The role Most companies that say they "build AI" are wiring API calls to somebody else's model. We do that too, where it makes sense. But the interesting problems at Innovaccer are the ones a general-purpose model cannot solve: reasoning over a patient record that spans eleven systems and thirty years, deciding whether a prior authorization meets a payer policy that changed last Tuesday, extracting a diagnosis code with an accuracy bar where being wrong has consequences. Those problems need models we train, evaluate, and own. This team builds them. You will work on post-training and domain adaptation of language models from a few billion parameters up to the hundreds of billions, agent architectures that hold up over long horizons, evaluation systems that tell the truth, and the inference infrastructure that gets all of it running under real latency and cost budgets. You will read papers on Monday and have something running by Friday. You will also be the person who takes it to production, because we do not have a research group that throws prototypes over a wall. Innovaccer is committing $250M over three years to its agentic AI platform. This team is where a meaningful share of that goes. What You Will Work On Nobody does all of this. You go deep on one or two areas and stay conversant in the rest. Building the model systems behind entire products Our products are not one model with a prompt in front of it. Flow, our revenue cycle platform, runs document understanding, policy retrieval and reasoning, code assignment, denial prediction, and drafting, each with different accuracy, latency, and cost requirements. Somebody has to design the whole thing and make the pieces work together. That is a core part of this job. You will own the model portfolio behind a product surface, which means: • Fine-tuning across the full size range. Small language models in the 1B to 8B range for high-volume, latency-sensitive, cost-constrained tasks. Medium models in the 8B to 70B range where the reasoning gets harder. Large open-weight models from 100B into the hundreds of billions where the task genuinely demands frontier capability. Knowing which tier a task actually needs is a judgment call worth a lot of money, and getting it wrong in either direction is expensive • Composing them into a system. Routing and cascades, small models handling the common case with escalation to larger ones, retrieval and tool layers, structured decoding, verifier models checking generator output, and the fallback paths for when a component fails • Owning it in production. Versioning across a model fleet, shadow deployment, staged rollout, drift detection, retraining triggers, and the regression suite that runs before anything ships • Holding a product-level accuracy bar, not a per-model benchmark score. A pipeline of individually good models can still produce a bad product, and finding out why is your problem Post-training at scale Domain adaptation of open-weight models to clinical, claims, and payer-policy data, across the full size range described above. • Supervised fine-tuning, preference optimization (DPO, GRPO, and what replaces them), and RL with verifiable rewards on tasks where correctness is machine-checkable • Continued pretraining and domain-adaptive pretraining where the vocabulary and distribution shift enough to justify it • Parameter-efficient methods where they suffice, full fine-tuning where they do not, and the experimental discipline to know which case you are in • Distillation from large teachers into small models that hold the accuracy bar at a fraction of the serving cost • Reward modeling, and the harder problem of specifying reward on tasks where clinical correctness is contested Running these jobs is its own discipline. You will work with multi-node training on hundreds of GPUs: FSDP and DeepSpeed, tensor and pipeline and sequence parallelism, activation checkpointing, mixed precision and its failure modes, checkpoint and resume strategy, throughput and MFU tuning, and diagnosing the loss spike at hour 40 of a run that cost real money. Experience keeping a large distributed run healthy is a specific skill and we are hiring for it explicitly. Data: raw to training-ready This is where most of the actual gain comes from, and it is the part most candidates undersell. Healthcare data does not arrive as a dataset. It arrives as HL7 feeds, FHIR bundles, claims files, PDFs, scanned faxes, free-text notes with inconsistent structure, and payer policy documents that change without notice. Turning that into a training corpus is a research problem in its own right. • Building the pipelines that take raw, messy, multi-format healthcare data to deduplicated, decontaminated, quality-filtered, PHI-safe training data • Data mixture design and the experiments that justify it. Ablations on what to include, at what ratio, and at what stage of training • Synthetic data generation, self-instruct and teacher-model pipelines, rejection sampling, and the quality controls that keep synthetic data from quietly poisoning a run • Annotation strategy with clinical and coding experts on staff: what to label, how to measure inter-annotator agreement, and when expert disagreement means your task definition is wrong rather than your labelers • Decontamination against your own eval sets, done properly, before someone else finds the leak Agents and reasoning systems Multi-step agents that complete real operational work rather than producing suggestions: submitting an authorization, closing a care gap, resolving a denial. Tool use, planning, memory, and recovery from failure. Research questions here are open: how to train agentic behavior rather than prompt it, how to do credit assignment over long trajectories, how to make an agent recognize it is failing and stop. The hard part is reliability over long horizons when every intermediate ste

More jobs at Innovaccer

All Innovaccer jobs (67)

Data jobs in Noida

Data jobs in Noida (373)

Other Data jobs in India

All Data jobs in India (11,234)