Ai Ml Engineer
Care Health Insurance · Gurugram, Haryana, India
Free to search · AI fit score against your CV · tailor your résumé in one click
Care Health Insurance · Gurugram, Haryana, India
AI/ML Engineer (GenAI, OCR & ML) Experience: 1-5 Years Job Summary We are looking for a Backend AI Engineer to design and scale GenAI- and CV-powered production systems. You will own end-to-end backend architecturefrom raw data in S3 through OCR and object detection, to high-concurrency FastAPI services that call LLMs and return reliable, structured outputs for business workflows. Key Stack: FastAPI, Django (optional), AWS (S3/EC2/ECS), LLMs, Vector DBs, YOLO / CV, OCR, Redis/Celery Key Responsibilities • Production APIs: Build and maintain scalable async backends with FastAPI (high-throughput AI inference) and, where needed, Django for application logic and state. • LLM systems: Integrate and optimize LLMs (API or self-hosted)prompt engineering, structured output parsing, fallbacks, and cost/latency trade-offs. • RAG: Design Retrieval-Augmented Generation using vector DBs (Pinecone, Milvus, or pgvector) for context-aware responses over policies, tickets, and documents. • OCR at scale: Architect OCR pipelines for multi-page PDFs and noisy scans; combine OCR with layout understanding and LLM post-processing for field extraction and validation. • Computer vision (YOLO & beyond): Lead detection/classification models for document regions, seals, signatures, ID cards, damage/claims images, or quality gates; manage training data, evaluation metrics (mAP, precision/recall), and model versioning in S3. • AWS & data orchestration: Deploy on EC2/ECS/Fargate; manage large datasets and weights in S3; build Boto3 pipelines for ingest, cleaning, and model-ready datasets. • Performance: Optimize latency with Redis caching, Celery async jobs, batching, and efficient GPU/CPU inference paths. • Production hardening: Monitoring, retries, idempotency, PII-aware handling, and clear SLAs for AI services. Required Skills • 1-5 years Python with FastAPI and/or Django; strong async programming and ORM/SQL optimization (PostgreSQL). • Proven GenAI delivery: LangChain or LlamaIndex (or equivalent), embeddings, tokenization awareness, production LLM integration. • Solid OCR experience in production (Textract, Tesseract, PaddleOCR, or similar) including accuracy/error analysis. • Hands-on YOLO / CV detection or document AI (PyTorch/Ultralytics/OpenCV); ability to improve models with data, not only run pretrained weights. • AWS: S3 (versioning/lifecycle), EC2/ECS, IAM; Docker for reproducible AI environments. • Vector DB experience and Redis/Celery (or equivalent) for async workloads. Nice-to-Have • WebSockets / SSE for real-time AI streaming • Model serving (TorchServe, Triton, vLLM, or similar) • MLOps basics: experiment tracking, dataset versioning, CI for models • Domain experience: insurance documents, claims, KYC, email/ticketing automate