Senior Software Engineer- ETL Developer
CGI · Bengaluru, Karnataka, India - Chennai, Tamil Nadu, India
CGI · Bengaluru, Karnataka, India - Chennai, Tamil Nadu, India
Job Title: ETL Developer Position: Software Engineer Experience: 4- 7 Years Category: Software Development/ Engineering Shift: 1-10PM (Hybrid) Main location: Bangalore / Chennai Position ID: J0726-1184 **Employment Type:** Full Time Education Qualification: Bachelor's degree in Computer Science or related field or higher with minimum 3 years of relevant experience. **Position Description:** We are looking for an experienced ETL Developer to join our team. The ideal candidate should be passionate about coding and developing scalable and high-performance applications. The successful candidate will be responsible for designing, developing, optimizing, and supporting enterprise-grade ETL and ELT pipelines primarily using Azure Databricks, PySpark, SQL, and related Microsoft Azure services. This is a hands-on engineering role focused on building scalable, reliable, and high-performing data pipelines across development, testing, and production environments. The engineer will work closely with data architects, business analysts, application teams, quality assurance teams, and production support teams to deliver secure and well-governed data solutions in a regulated financial-services environment. #LI-MP14 **Your future duties and responsibilities:** • Design, develop, test, deploy, and support ETL and ELT pipelines using Azure Databricks. • Build scalable data transformation workflows using PySpark, Spark SQL, and Databricks notebooks. • Develop reusable Databricks frameworks, utilities, and common pipeline components. • Ingest and process data from relational databases, flat files, APIs, cloud storage, and enterprise applications. • Design and implement batch and near-real-time data pipelines. • Build and manage Databricks jobs, workflows, schedules, dependencies, retries, and alerts. • Develop data cleansing, validation, enrichment, reconciliation, and transformation logic. • Implement data quality controls, audit checks, exception handling, and restart capabilities. • Optimize Spark workloads, SQL queries, cluster configuration, partitioning, caching, and file formats. • Work with Delta Lake to support reliable, scalable, and transactional data processing. • Design and maintain bronze, silver, and gold data layers within a lakehouse architecture. • Support source-to-target mapping, data profiling, data lineage, and technical design activities. • Develop and maintain SQL queries, stored procedures, views, and data reconciliation scripts. • Integrate Databricks pipelines with Azure Data Lake Storage and other Azure services. • Build and maintain automated deployment processes using Azure DevOps CI/CD pipelines. • Promote code, notebooks, jobs, and configuration across development, testing, and production environments. • Monitor pipeline performance and troubleshoot failures, data issues, and production incidents. • Perform root-cause analysis and implement permanent fixes for recurring issues. • Create and maintain technical documentation, data mappings, deployment procedures, and support runbooks. • Ensure solutions comply with CIBC security, data governance, risk management, and regulatory requirements. **Must-Have Skills:** • Bachelors degree or diploma in Computer Science, Software Engineering, Data Engineering, Information Technology, or a related discipline. • Strong professional experience developing ETL or ELT solutions using Azure Databricks. • Strong hands-on experience with: o Databricks notebooks o Databricks Jobs and Workflows o Apache Spark o PySpark o Spark SQL o Delta Lake o Cluster configuration and management o Performance tuning and optimization • Strong experience designing and developing complex ETL pipelines in Databricks. • Experience building reusable and parameterized data pipelines. • Strong SQL development skills, including: o Complex queries o Stored procedures o Views and functions o Data reconciliation o Query optimization o Performance troubleshooting • Experience with Azure Data Lake Storage and cloud-based data ingestion. • Experience with structured, semi-structured, and unstructured data formats, including CSV, JSON, Parquet, and Delta. • Strong understanding of ETL and ELT patterns, data integration principles, and data processing frameworks. • Experience implementing data quality checks, logging, monitoring, alerting, and exception handling. • Experience with job orchestration, scheduling, dependency management, and restartability. • Experience with source control and CI/CD deployment practices using Azure DevOps. • Experience supporting Databricks pipelines in production environments. • Strong analytical, troubleshooting, and problem-solving skills. • Strong written and verbal communication skills. **Good-to-Have Skills:** • Experience with Databricks Unity Catalog, including access controls, governance, and data lineage. • Experience with Databricks Asset Bundles or other Databricks deployment frameworks. • Experience with Databricks SQL Warehouses. • Knowledge of medallion and lakehouse architecture patterns. • Experience with Auto Loader and incremental data ingestion. • Experience with streaming pipelines using Spark Structured Streaming. • Experience with Delta Live Tables or Lakeflow Declarative Pipelines. • Experience with Python software engineering practices, including modular development and unit testing. • Experience integrating Databricks with Azure Data Factory. • Experience with Azure API Management and REST API integrations. • Experience with PowerShell or other scripting languages. • Familiarity with data modelling techniques, including dimensional and relational models. • Experience migrating legacy ETL workloads from Informatica, SSIS, or similar platforms to Databricks. • Previous experience working within banking, capital markets, commercial banking, corporate banking, or another regulated f