AWS Data Engineer - G Noida
Coforge · Delhi, Delhi, India - Hyderabad, Telangana, India - San Carlos, Rio San Juan, Nicaragua - Pune, Maharashtra, India
Free to search · AI fit score against your CV · tailor your résumé in one click
Coforge · Delhi, Delhi, India - Hyderabad, Telangana, India - San Carlos, Rio San Juan, Nicaragua - Pune, Maharashtra, India
2. Key Responsibilities • Design, develop and maintain robust, scalable ETL/ELT pipelines on Databricks using Python, PySpark and Spark SQL, following medallion (Bronze/Silver/Gold) architecture principles. • Independently analyze business requirements, translate them into technical designs and solution options, and drive them to completion with minimal handholding. • Design and implement conceptual, logical and physical data models (dimensional and 3NF) for analytics, reporting and master data use cases. • Build and manage data ingestion from enterprise sources (SAP S/4HANA, Salesforce, flat files, APIs, streaming) into the Lakehouse. • Develop and operate workloads on AWS (S3, Lambda, IAM, networking fundamentals) integrated with the Databricks platform. • Implement data quality checks, validation rules, reconciliation and error-handling frameworks; prepare test scenarios and perform thorough unit and business-level validation of own deliverables before handover. • Apply Unity Catalog-based governance: access controls, lineage, cataloguing and aligned handling of personal data. • Optimize Spark jobs and Delta tables for performance and cost (partitioning, Z-ordering, cluster sizing, job orchestration). • Contribute to CI/CD practices for data pipelines (Git-based version control, code reviews, automated deployment) and produce clear technical documentation. • Support BI and MDM workstreams by delivering curated, well-modelled datasets for Power BI and Informatica IDMC consumption. • Provide L2/L3 support for deployed pipelines, troubleshoot production issues and drive root-cause resolution. 3. Primary (Must-Have) Skills • Python: Strong, hands-on programming for data engineering: clean, modular, well-tested code. • SQL: Advanced SQL for complex transformations, analysis and performance optimization. • Apache Spark / PySpark: Deep understanding of Spark architecture, Data frame API, Spark SQL, performance tuning and debugging of distributed jobs. • Data Modelling: Solid experience in dimensional modelling (star/snowflake), normalized modelling, slowly changing dimensions and canonical/master data models. • Databricks: Proven project experience with Databricks workspaces, Delta Lake, Delta Live Tables/Jobs, Workflows and Unity Catalog. • AWS: Working proficiency with core AWS data services (S3, Glue, Lambda, IAM, CloudWatch) and integration with Databricks. 4. Other Required Skills • Experience with orchestration tools (Databricks Workflows or equivalent). • Git-based development workflow, code review discipline and CI/CD for data pipelines (Azure DevOps or similar). • Exposure to integrating with enterprise systems such as SAP S/4HANA and Salesforce (APIs, CDC, extractors) is highly desirable. • Familiarity with Power BI datasets/semantic models and with MDM concepts (e.g., Informatica IDMC) is an advantage. • Exposure to streaming technologies (Structured Streaming, Kafka/Kinesis) is a plus. 5. Ways of Working & Behavioral Expectations • Independent delivery: Works independently end-to-end: clarifies requirements early, proposes designs, and delivers complete, tested solutions without repeated review cycles. • Ownership: Takes accountability for quality and timelines; proactively prepares test scenarios and validates business logic before submitting work. • Speed of comprehension: Grasps new requirements and domain logic quickly and converts them into working solutions at the pace the project demands. • Communication: Communicates progress, risks and blockers clearly and early; documents solutions to a standard others can maintain. • Teamwork: Collaborates effectively with BI developers, MDM consultants, business SMEs and vendor teams in a multi-vendor environment. 6. Qualifications • Bachelor's degree in Computer Science, Engineering or a related field (Master's preferred). • Relevant certifications are an advantage: Databricks Certified Data Engineer (Associate/Professional), AWS Certified Data Engineer / Solutions Architect.