Data Engineering Manager
Kpmg India Services · Bengaluru/Bangalore, Karnataka
Kpmg India Services · Bengaluru/Bangalore, Karnataka
Manager- Data Engineer Key responsibilities Design, implement, and optimize end-to-end ETL pipelines in Microsoft Fabric, from ingestion through multi-stage transformations to data loading and delivery. Build pipelines and notebooks using Python and PySpark; implement data validation, error handling, and quality controls. Collaborate with business analysts and stakeholders to translate requirements and accounting logic into transformation rules and data solutions. Work with data architects to design schemas and data models (star/snowflake) and, where needed, OLAP cubes aligned to application requirements. Ensure efficient, accurate processing from source systems into Fabric's data layers; optimize performance and scalability (partitioning, indexing, resource tuning). Leverage modern tooling and practices (e.g., Azure DevOps for boards, repos, CI/CD); uphold consistent development standards across the global team. Conduct unit, integration, and end-to-end testing; troubleshoot and continuously improve ETL processes. Maintain comprehensive, up-to-date documentation (processes, sources, data flows, models) accessible to stakeholders; ensure compliance with policies and standards. Required skills and experience High proficiency with Microsoft Fabric ETL; experience with related tools such as Azure Data Factory. Strong SQL for extraction, transformation, and querying; hands-on with SQL Server, Azure SQL Database, and Synapse Analytics. Data engineering fundamentals: data modeling and schema design (star/snowflake), transformation, and optimization for warehousing/analytics; experience with OLAP where applicable. Proficiency in Python and PySpark for ETL development within Fabric notebooks and pipelines. Experience loading and optimizing data at scale in Fabric and prior exposure to Azure Synapse and Azure Data Lake. Familiarity with Azure DevOps workflows (work tracking, version control, pipelines) and modern development practices. Rigorous testing approach (unit, integration, E2E), with robust data validation and error-handling procedures. Strong collaboration and communication in globally distributed teams; ability to share best practices and evolve standards based on feedback and industry trends. Key responsibilities Design, implement, and optimize end-to-end ETL pipelines in Microsoft Fabric, from ingestion through multi-stage transformations to data loading and delivery. Build pipelines and notebooks using Python and PySpark; implement data validation, error handling, and quality controls. Collaborate with business analysts and stakeholders to translate requirements and accounting logic into transformation rules and data solutions. Work with data architects to design schemas and data models (star/snowflake) and, where needed, OLAP cubes aligned to application requirements. Ensure efficient, accurate processing from source systems into Fabric's data layers; optimize performance and scalability (partitioning, indexing, resource tuning). Leverage modern tooling and practices (e.g., Azure DevOps for boards, repos, CI/CD); uphold consistent development standards across the global team. Conduct unit, integration, and end-to-end testing; troubleshoot and continuously improve ETL processes. Maintain comprehensive, up-to-date documentation (processes, sources, data flows, models) accessible to stakeholders; ensure compliance with policies and standards. Required skills and experience High proficiency with Microsoft Fabric ETL; experience with related tools such as Azure Data Factory. Strong SQL for extraction, transformation, and querying; hands-on with SQL Server, Azure SQL Database, and Synapse Analytics. Data engineering fundamentals: data modeling and schema design (star/snowflake), transformation, and optimization for warehousing/analytics; experience with OLAP where applicable. Proficiency in Python and PySpark for ETL development within Fabric notebooks and pipelines. Experience loading and optimizing data at scale in Fabric and prior exposure to Azure Synapse and Azure Data Lake. Familiarity with Azure DevOps workflows (work tracking, version control, pipelines) and modern development practices. Rigorous testing approach (unit, integration, E2E), with robust data validation and error-handling procedures. Strong collaboration and communication in globally distributed teams; ability to share best practices and evolve standards based on feedback and industry trends. Key responsibilities Design, implement, and optimize end-to-end ETL pipelines in Microsoft Fabric, from ingestion through multi-stage transformations to data loading and delivery. Build pipelines and notebooks using Python and PySpark; implement data validation, error handling, and quality controls. Collaborate with business analysts and stakeholders to translate requirements and accounting logic into transformation rules and data solutions. Work with data architects to design schemas and data models (star/snowflake) and, where needed, OLAP cubes aligned to application requirements. Ensure efficient, accurate processing from source systems into Fabric's data layers; optimize performance and scalability (partitioning, indexing, resource tuning). Leverage modern tooling and practices (e.g., Azure DevOps for boards, repos, CI/CD); uphold consistent development standards across the global team. Conduct unit, integration, and end-to-end testing; troubleshoot and continuously improve ETL processes. Maintain comprehensive, up-to-date documentation (processes, sources, data flows, models) accessible to stakeholders; ensure compliance with policies and standards. Required skills and experience High proficiency with Microsoft Fabric ETL; experience with related tools such as Azure Data Factory. Strong SQL for extraction, transformation, and querying; hands-on with SQL Server, Azure SQL Database, and Synapse Analytics. Data engineering fundamentals: data modeling and schema design (star/snowflake), transformation, and optimization for wareh