Job description

**Roles and Responsibilities** - Design, develop, test, deploy and maintain scalable data pipelines using BigQuery, Cloud Storage, Pub/Sub, and other GCP services. - Develop high-quality code in PySpark on AWS EMR clusters for big data processing tasks such as batch and streaming analytics. - Troubleshoot issues related to pipeline failures or errors by analyzing logs, debugging techniques, and collaborating with team members. **Desired Candidate Profile** - 5-10 years of experience in developing large-scale data pipelines using BigQuery, Cloud Storage, Pub/Sub, etc. - Strong expertise in working with PySpark on AWS EMR clusters for big data processing tasks like batch & streaming analytics. - Experience with GCP services including Dataflow (Astronaut), Cloud Functions (Cloud Run), Pub/Sub Messaging System.