S

Data Engineer

Sonata Software · Pune District, Maharashtra

2–7 yrs experiencefull_timePosted 2w ago
Apply now →

Job description

JUNIOR DATA ENGINEER — - 1–3 years of hands-on experience with PySpark (DataFrames, Spark SQL basics) — production or strong academic/project experience acceptable. - Working knowledge of Fivetran or a similar ELT tool (Airbyte, Stitch) — setting up connectors, monitoring syncs. - Solid SQL fundamentals and exposure to a cloud data warehouse (Snowflake / Redshift / BigQuery). - Basic Python scripting ability for data tasks. - Exposure to at least one cloud platform (AWS / Azure / GCP), internship or project-level acceptable. - Eager to learn, works well under guidance from senior engineers; not expected to design architecture independently. SENIOR DATA ENGINEER — REQUIREMENTS - 5+ years of hands-on experience with PySpark (DataFrames, Spark SQL, RDDs, performance tuning, job optimization) in production. - 3+ years configuring and managing Fivetran connectors at scale (custom connectors, schema drift handling, sync orchestration, troubleshooting). - Strong SQL and proven data modeling / warehouse design experience (Snowflake / Redshift / BigQuery). - Deep experience with at least one major cloud platform (AWS, Azure, or GCP), including cost and performance optimization. - Proven ability to design end-to-end pipeline architecture and lead technical decisions. - Experience mentoring junior engineers and reviewing code/pipeline designs. RESPONSIBILITIES - Design, build, and optimize PySpark ETL/ELT pipelines for large-scale batch and/or streaming data. - Own Fivetran connector strategy across [X] source systems, including custom connector development. - Define data architecture and standards; review junior engineers' pipeline designs. - Collaborate with analytics/BI teams and business stakeholders to define data requirements. - Implement data quality frameworks, validation, and monitoring across pipelines.