Senior Data Engineer(15-30 days joiners only)
Quantiphi · Bengaluru, Karnataka, India
Free to search · AI fit score against your CV · tailor your résumé in one click
Quantiphi · Bengaluru, Karnataka, India
Summary We are looking for an experienced Senior Data Engineer to join our team at Quantiphi, a leader in AI-first engineering and innovative GenAI solutions. In this pivotal role, you will design and architect robust data pipelines that power advanced AI applications, specifically focusing on building high-performance ingestion engines for unstructured data, vector databases, and scalable real-time streaming architectures. You will bridge the gap between complex data engineering and modern AI, ensuring data quality, lineage, and observability as we deploy cutting-edge Agentic AI solutions. Responsibilities • Design and build batch ETL pipelines that ingest large unstructured document corpora - handling PDF, RTF, and JSON file formats, consuming pre-extracted text and metadata, applying document chunking strategies, and loading processed outputs into a vector store and analytical platforms. • Design partitioning, parallelism, and throughput strategies to sustain high-volume ingestion. • Build real-time streaming ingestion pipelines consuming document creation and update events, processing them through chunking and embedding generation. • Build vector and indexing pipelines targeting vector store platforms (e.g., Vertex AI Vector Search, Pinecone, Weaviate, Milvus, pgvector) • Manage vector store index operations: bulk load, incremental upsert, index refresh, and consistency validation post-ingestion. • Implement metadata tagging on vector records to support filtered retrieval at query time. • Implement end-to-end lineage tracking for vector records - linking each chunk back to its source document. • Implement document-level change tracking to detect upstream updates, deletions, and re-ingestion events and trigger appropriate vector store operations (upsert, delete, re-index) without full corpus re-processing. • Tag all vector records with pipeline run metadata - ingestion timestamp, pipeline version, chunking strategy version, embedding model version, to support retrieval quality debugging and model refresh traceability. Must Have Skills • Data Engineering - Experience building production-grade batch and streaming data pipelines supporting enterprise analytics and AI workloads. • Python- Strong proficiency in Python for data processing, pipeline development, automation, and transformation frameworks. • Streaming Technologies - Experience building real-time ingestion pipelines using Apache Kafka, GCP Pub/Sub, Azure Event Hubs, or equivalent event streaming platforms. • Data Platforms - Experience working with BigQuery/Synapse, Databricks, Snowflake, or equivalent cloud-native analytical platforms. • Data Modeling & SQL - Strong SQL skills and experience designing scalable analytical and operational data models. • Vector databases - Familiarity with Pinecone, Vertex AI Vector Search, Weaviate, Milvus, pgvector, or similar retrieval platforms. • Document processing - Familiarity with PDF, RTF, JSON, OCR pipelines, and metadata extraction workflows. • Data Quality & Governance - Experience implementing validation frameworks, reconciliation, lineage, metadata management, monitoring, and pipeline observability. • Cloud Platforms - Hands-on experience building cloud-native data solutions using GCS/ADLS, BigQuery/Synapse, Dataflow/ADF, Cloud Composer/Airflow, or equivalent services. • Regulated Data - Experience working with regulated data environments where security, auditability, governance, and compliance are critical.