U

Enterprise Architect - Databricks

Unisys · Bangalore, KA, India

12–20 yrs experiencePosted Today
Apply now →

Job description

What success looks like in this role: Databricks Platform Architecture • Architect enterprise Databricks Lakehouse platforms on Azure, AWS, or GCP — including Unity Catalog, Delta Live Tables (DLT), and Databricks SQL. • Design multi-workspace, multi-region Databricks deployments with workspace federation, network isolation (Private Link / V Net injection), and identity federation via Azure AD / Okta / SCIM. • Define medallion architecture (Bronze / Silver / Gold) standards, enforcing schema evolution, data contracts, and SLA-tiered pipeline SLOs. • Lead migration of legacy data warehouses (Teradata, Netezza, Snowflake, Synapse) to the Databricks Lakehouse — including SQL translation, workload profiling, and TCO modeling. • Own cluster architecture decisions: autoscaling policies, instance fleet configurations, job vs. interactive cluster strategies, Photon engine enablement, and cost-per-query optimization. • Design and implement Databricks Workflows and Delta Live Tables for mission-critical streaming and batch pipelines, including CDC (Change Data Capture) patterns using Autoloader and Structured Streaming. Data Governance & Security • Implement Unity Catalog as the enterprise meta store: fine-grained access control (row/column-level security), lineage tracking, and cross-workspace catalog federation. • Define data classification frameworks, PII masking strategies, and dynamic views for regulatory compliance (GDPR, HIPAA, SOX) within the Databricks platform. • Architect Delta Sharing for secure, governed cross-organizational data sharing without data movement. • Lead data mesh and data product design patterns on Databricks — aligning ownership, SLAs, and discoverability across domains. AI & ML Platform (Databricks ML) • Design end-to-end ML platforms using Databricks ML Runtime, ML flow (Tracking, Registry, Projects), and Feature Store for model lifecycle management. • Architect LLM / Generative AI workloads on Databricks: RAG pipelines, fine-tuning with Mosaic AI, vector search (Databricks Vector Search), and Model Serving endpoints. • Define MLOps frameworks covering CI/CD for models, drift detection, A/B testing, and shadow deployments using ML flow and Databricks Workflows. • Integrate Databricks AI/BI (Genie, Dashboards) for self-service analytics and natural language data querying. Stakeholder & Delivery Leadership • Lead architecture reviews, technical workshops, and proof-of-concept engagements with C-suite and VP-level client stakeholders. • Author solution architecture documents, reference architectures, and RFP/RFI responses for Databricks-led pursuits. • Define center-of-excellence (CoE) standards, reusable accelerators, and Databricks best-practice playbooks for the Unisys Data & AI practice. • Mentor and upskill a team of data architects and engineers; drive Databricks certification paths across the practice. • Partner with Databricks field engineering on joint go-to-market opportunities and co-delivery engagements. #LI-SS1 You will be successful in this role if you have: • 12&#43; years of experience in data architecture, data engineering, or enterprise analytics — with at least 5 years focused on Databricks platform delivery. • Databricks Certified Data Engineer Professional or Databricks Certified Associate Developer for Apache Spark — required. Additional Databricks certifications strongly preferred. • Expert-level proficiency in Apache Spark (PySpark, Scala Spark) — performance tuning, DAG optimization, shuffle management, and broadcast strategies. • Deep hands-on expertise with Delta Lake: ACID transactions, Z-ordering, OPTIMIZE, VACUUM, time travel, and schema enforcement/evolution. • Proven production experience with Delta Live Tables (DLT): expectations, quarantine patterns, SCD Type 1/2 in DLT, and pipeline monitoring. • Strong Unity Catalog implementation experience: catalog/schema/table hierarchy design, privilege inheritance, service principals, and attribute-based access control. • Experience architecting Databricks on at least one major cloud: Azure Databricks (ADLS Gen2, Azure AD), AWS (S3, IAM Instance Profiles, Glue Catalog), or GCP (GCS, Dataproc comparison). • Proficiency in infrastructure-as-code for Databricks: Terraform (databricks provider), Databricks Asset Bundles (DABs), and CI/CD integration (GitHub Actions, Azure DevOps, Jenkins). • Hands-on MLflow experience: experiment tracking, model registry, model serving, and custom pyfunc models. • Strong SQL expertise for Databricks SQL / Photon query optimization, materialized views, and lakehouse serving patterns. • <span