Job description

**Working with Us** Challenging. Meaningful. Life-changing. Those aren’t words that are usually associated with a job. But working at Bristol Myers Squibb is anything but usual. Here, uniquely interesting work happens every day, in every department. From optimizing a production line to the latest breakthroughs in cell therapy, this is work that transforms the lives of patients, and the careers of those who do it. You’ll get the chance to grow and thrive through opportunities uncommon in scale and scope, alongside high-achieving teams. Take your career farther than you thought possible. Bristol Myers Squibb recognizes the importance of balance and flexibility in our work environment. We offer a wide variety of competitive benefits, services and programs that provide our employees with the resources to pursue their goals, both at work and in their personal lives. Read more: careers.bms.com/working-with-us . **Position: Data Engineer II, Data & Analytics** **Location: Hyderabad, India** At Bristol Myers Squibb, we are inspired by a single vision – transforming patients’ lives through science. In oncology, hematology, immunology, and cardiovascular disease – and one of the most diverse and promising pipelines in the industry – each of our passionate colleagues contribute to innovations that drive meaningful change. We bring a human touch to every treatment we pioneer. Join us and make a difference. **Position Summary** The ideal candidate will have a strong background in data engineering, and data systems and data governance and will be comfortable working with both structured and unstructured data. Prior domain experience with enterprise functions (HR/ Finance/ Compliance/ Procurement) is a big plus. Experience In building production-grade applications and APIs that expose data and AI capabilities to enterprise users. Exposure to Generative AI and LLM-powered solutions in an enterprise or cloud environment is highly desirable If you want an exciting and rewarding career that is meaningful, consider joining our diverse team! **Key Responsibilities** - Accountable for enhancing and delivering high quality, data products and analytic ready data solutions for Enabling Functions - Working with stakeholders to define the overall strategy for the organization's data products and new features. This involves understanding business goals, identifying data sources, and determining the appropriate technology stack and making recommendations/ building enhancements. - Developing and implementing data engineering solutions that support the organization's business needs. This may involve working with various technologies, including data warehouses, data lakes, data marts, and data APIs. - Understand and collaborate with data architect to maintain data models to support our reporting and analysis needs. - Optimize data storage and retrieval to ensure efficient performance and scalability. - Implement standard data governance policies and procedures to ensure that data is accurate, consistent, and secure. This includes defining data quality standards data security protocols, and data privacy policies. - Close partnering with the Enterprise Data and Analytics Platform team, other functional data pods and Data Community lead to shape and adopt data and technology strategy. - Staying up to date with industry trends: Keeping up to date with the latest trends and advancements in data architecture/ data engineering and technology and applying this knowledge to enhance the organization's data systems and processes. - Comfortable working in a fast-paced environment with minimal oversight - Prior experience working in an Agile/Product based environment. - Design and develop RESTful APIs to expose data products and analytics capabilities to business applications - Build reusable API frameworks to accelerate development across multiple enterprise use cases - Integrate LLM APIs (OpenAI, AWS Bedrock, Anthropic Claude) into enterprise workflows for use cases like natural language querying, document summarization, and conversational AI - Build RAG (Retrieval-Augmented Generation) pipelines using vector databases (pgvector, Pinecone, OpenSearch) combined with LLM APIs - Leverage orchestration frameworks such as LangChain or LlamaIndex to build multi-step AI agents and automated workflows - Proficiency in SQL and Python for data transformation, validation, and pipeline development. - Experience building ETL/ELT pipelines and working with lakehouse concepts (e.g., medallion architecture; bronze/silver/gold layers). - Hands-on experience with Databricks (notebooks/jobs/workflows) and Delta Lake concepts (ACID tables, incremental processing, upserts/merge). **Qualifications & Experience** - 3-5 years of experience in information technology field in developing AWS cloud native data lakes and ecosystems including some production support experience - Hands on experience with programming and analytics capabilities using some AWS. - (Amazon Web Services) Native Services (Glue Studio, Athena, Redshift, Postgres DB etc.) is required. Cloudera Data Platform (CDP) is a plus. - Programming skills in languages such as Python, Pyspark, etc. - At least 1-2 years of experience in an onshore offshore delivery model - Solid programming skills in Python and Spark and strong proficiency in Cloud – AWS - Knowledge of data security and privacy best practices - Experience with SQL and database technologies such as MySQL, PostgreSQL etc. - Comfortable working in a dynamic global environment - Excellent communication and collaboration skills. Functional knowledge or prior experience in Enterprise Functions is a plus. - Demonstrates a focus on improving processes, structures, and knowledge within the team. - Familiarity with LLM APIs and prompt engineering best practices - Experience with vector databases (pgvector, Pinecone, or OpenSearch) is a plus - Hands-on experience building and consuming RESTful APIs using Python frameworks (FastAP