Senior Data & AI Engineer
Bristol Myers Squibb · State of Telangāna, India
Bristol Myers Squibb · State of Telangāna, India
**Working with Us** Challenging. Meaningful. Life-changing. Those aren’t words that are usually associated with a job. But working at Bristol Myers Squibb is anything but usual. Here, uniquely interesting work happens every day, in every department. From optimizing a production line to the latest breakthroughs in cell therapy, this is work that transforms the lives of patients, and the careers of those who do it. You’ll get the chance to grow and thrive through opportunities uncommon in scale and scope, alongside high-achieving teams. Take your career farther than you thought possible. Bristol Myers Squibb recognizes the importance of balance and flexibility in our work environment. We offer a wide variety of competitive benefits, services and programs that provide our employees with the resources to pursue their goals, both at work and in their personal lives. Read more: careers.bms.com/working-with-us . **Position Summary** We are looking for a passionate Senior Data & AI Engineer to design, build, and scale modern data and AI-powered solutions that enable advanced analytics, reporting, and intelligent business applications. The ideal candidate combines deep expertise in cloud data engineering, Databricks, and AWS with hands-on experience in Generative AI, Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), Vector Databases, and Agentic AI architectures. This role will be responsible for delivering trusted, scalable, and analytics-ready data products while driving innovation through modern engineering practices, AI-assisted development, and cloud-native technologies. **Key Responsibilities** - Design, develop, and support scalable batch, streaming, and real-time data pipelines. - Build and maintain analytics-ready data products, data lakehouse solutions, and curated datasets. - Develop ETL/ELT frameworks using Databricks, PySpark, SQL, and AWS data services. - Create and optimize data models, data marts, and enterprise data products. - Ensure data quality through validation, monitoring, reconciliation, and automated testing. - Implement data governance, lineage, security, access controls, and compliance standards. - Design and develop APIs, reusable data services, and engineering accelerators. - Leverage LLMs and AI-assisted development tools (Claude Code, GitHub Copilot, Microsoft Copilot) to improve engineering productivity and operational efficiency. - Build and support RAG pipelines, vector embedding frameworks, semantic search solutions, and AI-powered applications. - Develop and deploy AI Agents using frameworks such as LangChain, LangGraph, CrewAI, AutoGen, or similar technologies. - Implement CI/CD pipelines, GitHub workflows, Infrastructure as Code (IaC), and automated deployment processes. - Collaborate with Data Architects, Data Scientists, Analysts, and Product Owners to deliver business-driven solutions. - Mentor team members and drive engineering standards, best practices, and innovation initiatives. **Skills & Competencies** - Expert-level proficiency in Python, PySpark, SQL, and distributed data processing frameworks. - Strong expertise in Databricks, including Delta Lake, Unity Catalog, Workflows, and Lakehouse architecture. - Deep knowledge of the AWS Data & Analytics ecosystem, including Glue, Lambda, S3, Athena, Redshift, Lake Formation, CloudFormation, and related services. - Hands-on experience with Generative AI, Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), Vector Databases, and AI Agent development. - Experience leveraging AI-assisted engineering tools such as Claude Code, GitHub Copilot, Microsoft Copilot, or similar platforms to accelerate development and operational efficiency. - Experience with Agentic AI orchestration frameworks such as LangChain, LangGraph, CrewAI, AutoGen, or equivalent technologies. - Strong expertise in Git, GitHub Actions, CI/CD pipelines, DevOps practices, Infrastructure as Code (IaC), and AWS DevOps automation. - Experience designing and developing APIs, microservices, and cloud-native data solutions. - Solid understanding of data governance, metadata management, data lineage, security, privacy, and enterprise data platforms. - Experience implementing data quality frameworks, monitoring, observability, and operational excellence practices. - Exposure to BI and visualization platforms such as Tableau, Power BI, or equivalent tools for downstream analytics consumption. - Strong analytical, problem-solving, communication, and stakeholder management capabilities. - Proven ability to collaborate effectively with both technical and business stakeholders in a cross-functional, agile environment. - Self-driven with an ownership mindset and the ability to thrive in a fast-paced, product-oriented organization. **Qualifications & Experience** - Bachelor's or Master's degree in Computer Science, Engineering, Data Science, Information Systems, or a related field. - 5-8 years of hands-on experience in Data Engineering, Software Engineering, or Cloud Data Platforms. - SME-level expertise in Databricks and modern Lakehouse architectures. - Proven experience building enterprise-scale data pipelines and data products on AWS and/or Azure. - Strong experience with real-time and batch data ingestion frameworks. - Experience developing APIs and cloud-native data services. - Hands-on experience in one or more of the following areas: AI/ML, Generative AI, Agentic AI, Retrieval-Augmented Generation (RAG), Vector Search, and LLM-powered applications. - Experience working in Agile, product-oriented environments with end-to-end ownership. - Ability to lead complex technical initiatives and mentor engineering teams. - Prior experience in Life Sciences, Healthcare, or other regulated industries is a plus. - Certifications in AWS, Databricks, Azure, Google Cloud, or related cloud and data technologies are a strong plus. **Good to Have** - Experience with Vector Databases such as