Associate Director - Data Engineering
Kpmg India Services · Bengaluru/Bangalore, Karnataka
Kpmg India Services · Bengaluru/Bangalore, Karnataka
Data Engineer - Associate Director KPMG Global Services KPMG Global Services (KGS) was set up in India in 2008. It is a strategic global delivery organization, which works with more than 50 KPMG member firms to provide a progressive, scalable and customized approach to business requirements The KGS journey has been one of consistent growth, with a current employee count of nearly 10,000 operating from four locations in India --- Bengaluru, Gurugram, Kochi and Pune, providing a range of Advisory and Tax-related services to member firms within the KPMG network. As part of KPMG in India, we were ranked among the top companies to work for in the country for four years in a row by LinkedIn, and recognized as one of the top three employers in the region for women, as well as for policies on Inclusion \& Diversity by ASSOCHAM (The Associated Chambers of Commerce \& Industry of India). Furthermore, as KPMG in India, we were recognized as one of the 'Best Companies for Millennials' at The Millennial Max Conference 2019 presented by The LNOD Roundtable as well as 'the Great Indian Workplace' at the Culture Summit and Great Indian Workplace Awards 2019. Team Overview KPMG's network of Data \& Analytics professionals recognizes that analytics has the power to create great value. That is why they take a business-first perspective, helping solve complex business challenges using analytics that clients can trust. Data \& Analytics professionals focus on solving complex business issues across all the key drivers of organizational value, including growth, risk and performance. And they work to deliver an unwavering commitment to precision and quality in everything they do. Roles and Responsibilities Designation Associate Director Reporting to Jagadish Doki Role type Senior Manager - Data Employment type Full-time Job Requirements Mandatory Skills 12 years of experience in data engineering, with at least 3 years leading large technical teams. Industry vendor certifications are desired (e.g. AWS, Azure, GCP, CNCF/Kubernetes or Databricks certifications); although not essential if you have demonstrable ability. Strong hands ‑ on expertise with Python, PySpark, and Databricks (including Lakehouse \& Databricks Workflows). Experience designing and deploying data solutions on Azure, AWS, or GCP. Strong understanding of data modelling techniques, data lifecycle management, and enterprise data principles. Expertise in designing, developing, and managing scalable, end-to-end data pipelines (ADF, Airflow or dbt,). Proficient in Big Data Platforms (Hadoop, Databricks, Hive, Kafka, Apache Iceberg or Microsoft Fabric), Data Warehouses (Teradata, Snowflake, BigQuery etc.) and lakehouses (Delta Lake, Apache Hudi) Proficient in programming languages such as SQL, Python and Pyspark with strong skills in writing scalable, readable and maintainable code using object-oriented programming concept. Implement DevOps practices, including Git workflows and CI/CD pipelines (Azure DevOps, Jenkins, GitHub Actions) to enhance automation and streamline deployments. Experience in project management frameworks such as Waterfall or Agile. Solid experience with SQL and handling large datasets. Familiarity with data governance frameworks (e.g., data quality controls, cataloging, lineage, roles \& policies). Excellent communication skills---able to translate complex data concepts to both technical and non ‑ technical stakeholders. Primary Roles and Responsibilities Lead and mentor a technical team (8--15 engineers) in designing and delivering data ‑ engineering solutions across cloud platforms (Azure/AWS/GCP). Architect, develop, and optimize ETL/ELT pipelines using Python, PySpark, Databricks, and distributed processing frameworks. Design and maintain scalable data architectures, including data lakes, lakehouses, warehouses, Delta tables, and semantic layers. Establish best practices in data modelling (dimensional, wide ‑ table, data vault), data quality, metadata management, and data governance. Collaborate with business owners, architects, product managers, and analytics teams to ensure end ‑ to ‑ end project delivery. Define standards for data operations, monitoring, lineage, and CI/CD for data workloads. Drive cloud platform adoption and modernization, ensuring solutions meet performance, cost ‑ efficiency, and compliance requirements. Promote innovation by evaluating new data engineering tools, frameworks, and practices. Mentor junior team members, providing guidelines to ensure high-quality deliverables. Communicate complex technical solutions to senior management and diverse stakeholders effectively. Contribute to sales activities through data and platform architecture expertise Stay updated on industry trends and contribute to internal initiatives, R\&D, and business development projects. Preferred Skills Experience with Azure Data Factory, AWS Glue, or GCP Dataflow. Experience with Microsoft Fabric solutions Knowledge of Lakehouse architecture, Delta Live Tables, and Medallion design patterns. Familiarity with CI/CD, DevOps, containerization (Docker/Kubernetes). Exposure to data management tools (Purview, Collibra, Informatica, BigID, etc.). Experience with different data execution paradigms, including low latency/streaming, batch, and micro-batch processing (Apache Kafka, Databricks Streaming processing, Apache StreamSets). Knowledge of data management frameworks (data governance, data quality, data security, data dictionary, metadata management) with tools like Databricks Unity Catalog, Apache Atlas, Informatica etc. Familiarity with containerization and orchestration tools like Docker and Kubernetes. Understanding of API gateway and service mesh architectures (e.g., Istio). Familiarity working with Linux-based operating systems. Familiarity with working with REST APIs. Experience in pre ‑ sales, solutioning, or client ‑ facing delivery leadership. Fluency in English (verbal and written). Other Information Number of interview r