S

AWS Data Lead(Must have: Glue, Athena, Redshift, Lake formation)

Sonata Software · Bengaluru, Karnataka, India - Chennai, Tamil Nadu, India - Hyderabad, Telangana, India

8–15 yrs experiencefull_timePosted Yesterday

Job description

Role Summary The AWS Data Lead will lead the design and delivery of AWS data engineering and analytics solutions, guide data engineers and analytics engineers, and ensure that data platforms are scalable, secure, governed, reliable, and aligned to business outcomes. The role bridges architecture and hands-on delivery across data ingestion, transformation, storage, analytics, governance, operational readiness, and AI / GenAI data foundations. Required Primary Skills - AWS data platform delivery using Amazon S3, AWS Glue, Amazon Redshift, Amazon Athena and AWS Lake Formation. - Strong SQL, Python and PySpark skills for data ingestion, transformation, validation and optimization. - ETL / ELT pipeline design, data lake and lakehouse implementation, data warehousing and dimensional modeling. - Technical leadership for data engineering teams, solution reviews, sprint planning and delivery governance. - Data quality, metadata, lineage, cataloging, security, privacy and access-control implementation. Secondary Skills - Amazon EMR, Amazon Kinesis, AWS DMS, AWS DataSync and Amazon OpenSearch. - Streaming and event-driven data processing using Kinesis, Kafka or equivalent technologies. - Infrastructure as Code using Terraform or AWS CloudFormation and CI/CD for data workloads. - Amazon QuickSight, Power BI or Tableau integration for analytics consumption. - AI / GenAI data foundations including feature stores, embeddings, vector-ready datasets and RAG data pipelines. Technical Skills - AWS services: Amazon S3, AWS Glue, Amazon Redshift, Amazon Athena, Amazon EMR, AWS Lake Formation, Amazon Kinesis, AWS DMS, AWS DataSync and Amazon OpenSearch. - Programming and data processing: SQL, Python, PySpark, Apache Spark, shell scripting and Scala as an advantage. - Data architecture: data lakes, lakehouse patterns, enterprise data warehouses, batch and streaming pipelines, CDC, data marts and semantic layers. - Databases: Amazon RDS, Aurora, DynamoDB, PostgreSQL, MySQL and NoSQL platforms. - Engineering practices: Git, CI/CD, automated testing, monitoring, logging, performance tuning, cost optimization and operational support. - Governance and security: IAM, encryption, Lake Formation permissions, data classification, lineage, auditability and compliance controls. Key Responsibilities - Lead discovery sessions and translate business, analytics and AI requirements into an executable AWS data platform backlog. - Own technical delivery of data ingestion, ETL / ELT, data lake, warehouse, analytics and streaming workstreams. - Define solution standards for pipeline design, data modeling, coding, testing, orchestration, monitoring and deployment. - Guide and mentor data engineers and analytics engineers through design reviews, code reviews and production-readiness checkpoints. - Collaborate with the AWS Data Architect on target architecture, service selection, scalability, resilience, security and cost optimization. - Establish data quality rules, reconciliation controls, metadata, cataloging, lineage and governance practices. - Coordinate with AI / ML and GenAI teams to prepare trusted datasets, feature pipelines, knowledge repositories and RAG-ready data assets. - Manage technical risks, dependencies, defects, performance issues and production incidents through root-cause analysis and corrective actions. - Support estimations, proposals, proof-of-concepts, customer demonstrations and reusable AWS data accelerators when required. - Ensure documentation, knowledge transfer and operational handover are complete for each release. Preferred Certifications - AWS Certified Data Engineer - Associate. - AWS Certified Solutions Architect - Associate. - AWS Certified Solutions Architect - Professional, preferred for senior candidates. - Databricks Data Engineer certification or equivalent data-platform certification, desirable. Expected Deliverables - AWS data solution design and delivery plan. - Data ingestion, ETL / ELT, streaming and orchestration pipelines. - Curated data lake, lakehouse, warehouse and analytics-ready datasets. - Data models, interface specifications, mappings and transformation rules. - Data quality, reconciliation, lineage, security and governance documentation. - Code-review records, test evidence, performance tuning results and release-readiness checklist. - Monitoring dashboards, operational runbooks, incident RCA and handover documentation. - Reusable templates, reference patterns, accelerators and lessons learned for the AWS competency. Qualification - B.E. / B.Tech / M.Tech / MCA or equivalent degree in Computer Science, Information Technology, Data Engineering, Analytics or a related discipline. - 6 to 10 years of overall experience in data engineering or analytics, including hands-on AWS data-platform delivery and team leadership responsibilities. - Demonstrated experience delivering production-grade data pipelines, data lakes, warehouses or analytics platforms. Good to have - Experience with migration from on-premises, Azure, GCP, Hadoop, Snowflake or legacy warehouse platforms to AWS. - Exposure to AWS Well-Architected reviews, FinOps, data mesh, lakehouse and domain-oriented data-product practices. - Experience building data foundations for SageMaker, Bedrock, recommendation systems, predictive analytics or GenAI / RAG use cases. - Customer-facing consulting experience across discovery workshops, architecture reviews, estimations, proposals and technical presentations. - Experience creating competency assets such as reference architectures, playbooks, reusable frameworks and proof-of-concepts.