Data Engineer - Senior
Cummins · Pune, Maharashtra, India - China
Free to search · AI fit score against your CV · tailor your résumé in one click
Cummins · Pune, Maharashtra, India - China
Leads the design, development, deployment, and maintenance of data and analytics platforms. Develops reliable, scalable, and efficient data pipelines and data processing solutions that enable data to be effectively processed, stored, governed, and made available to analysts and other data consumers. Collaborates with business stakeholders, IT experts, data scientists, architects, and subject-matter experts to deliver enterprise data and analytics solutions aligned with business and technical requirements. Key Responsibilities • Design, develop, and automate distributed data ingestion and transformation solutions using data from relational, event-based, semi-structured, and unstructured sources. • Build reliable, scalable, and efficient ETL/ELT data pipelines using appropriate tools, technologies, and scripting languages. • Design and implement data quality, validation, monitoring, and alerting frameworks to identify and resolve data integrity issues. • Implement data governance practices covering metadata, data access, retention, compliance, and security. • Design and implement physical data models, including database structures, indexing, and table relationships, to support performance and scalability. • Develop and operate large-scale data storage and processing solutions across cloud and distributed data platforms, including data lakes, warehouses, and lakehouse environments. • Optimize data pipelines, Spark workloads, databases, and cloud infrastructure for performance, reliability, scalability, and cost efficiency. • Integrate data from a variety of enterprise applications and source systems and support real-time and event-driven data processing. • Develop automation for common and repeatable data preparation, integration, deployment, and platform-management activities to minimize manual and error-prone processes. • Implement CI/CD and DevOps practices to support automated deployment, testing, and release management. • Participate in troubleshooting, testing, validation, and continuous improvement of data pipelines and platform solutions. • Ensure data platforms and solutions meet applicable quality, governance, security, compliance, and regulatory requirements. • Collaborate with data scientists, analysts, architects, IT teams, and business stakeholders to understand requirements and deliver effective data solutions. • Document data solutions, processes, designs, and technical information to support knowledge transfer and operational effectiveness. • Apply Agile development methodologies such as Scrum and Kanban to deliver data engineering initiatives. • Provide technical leadership and mentor less experienced team members, promoting engineering excellence and collaboration. Skills • Strong technical leadership and decision-making skills, with the ability to lead large-scale data engineering initiatives from concept through production deployment. • Strong problem-solving and analytical skills, including the ability to diagnose complex data pipeline, platform, and performance issues. • Excellent communication and collaboration skills, with the ability to work effectively with technical and business stakeholders. • Ability to translate complex technical concepts and stakeholder requirements into actionable data solutions. • Strong customer focus and ability to develop solutions aligned with business objectives. • Ability to balance strategic architecture decisions with hands-on technical execution. • Strong project management, prioritization, and organizational skills in a fast-paced environment. • Strong understanding of data quality, governance, security, compliance, and data management principles. • Passion for continuous improvement, emerging technologies, and modern data engineering practices. • Ability to mentor, guide, and develop technical talent while fostering a collaborative engineering culture. Technical Skills Data Engineering & Processing • Advanced proficiency in Azure Databricks, Apache Spark, and distributed data processing frameworks. • Strong expertise in enterprise-scale ETL/ELT and data ingestion pipeline design and development. • Experience processing structured, semi-structured, streaming, and large-scale datasets. • Strong understanding of Big Data technologies and scalable cloud-native data architectures. Programming & Development • Advanced proficiency in Python and Scala for large-scale data processing and engineering solutions. • Strong SQL skills for querying, transformation, optimization, and analysis of large datasets. • Experience with scripting, automation, version control, testing, and build processes. Cloud & Platform Technologies • Strong experience with Azure data services, including: • Azure Databricks • Azure Data Lake Storage (ADLS) • Azure Blob Storage • Azure Synapse Analytics • Azure SQL Data Warehouse • Strong understanding of cloud architecture principles, scalability, reliability, and security best practices. Data Storage & Lakehouse • Experience with modern data storage and lakehouse technologies, including: • Delta Lake • Apache Iceberg • Parquet • ORC • Strong understanding of data lake, data warehouse, and lakehouse architectures. Streaming & Data Integration • Experience with Kafka or similar real-time streaming and event-processing technologies. • Experience integrating data from a wide variety of enterprise applications and source systems. • Familiarity with Qlik Replicate or similar data replication and ingestion technologies is preferred. DevOps & Engineering Excellence • Experience implementing CI/CD pipelines and automated deployment processes. • Strong knowledge of Git, Jenkins, testing methodologies, version control, and release management. • Experience establishing coding standards, technical documentation, and engineering governance practices. Data Quality, Governance & Security • Experience implementing data validation frameworks, monitoring solutions,