21 Aug
|
Harbinger Techventures
|
Pune
21 Aug
Harbinger Techventures
Pune
Data Architect Consultant
Position: Data Architect Consultant | Modern Data & AI Architecture
Experience: 10+ years
Mode- Freelancer/ Consultant
Role Overview
We are looking for an experienced Data Architect Consultant to design scalable, secure and high-performance enterprise data architectures across cloud, data platforms, analytics and AI workloads.
The candidate will work with business, engineering, analytics and AI teams to define end-to-end data architecture, including data ingestion, integration, storage, processing, modeling, governance, security and consumption.
The role also requires the ability to design AI-ready data platforms supporting Generative AI, RAG, AI agents and advanced analytics.
Required Skills
Must Have
10+ years of experience in Data Engineering / Data Architecture.
Strong experience designing enterprise data architectures.
Strong SQL and data modeling expertise.
Experience with Data Warehouse / Data Lake / Lakehouse architectures.
Strong experience with at least one major cloud platform AWS, Azure or GCP.
Experience with modern data platforms such as Databricks, Snowflake, Microsoft Fabric or BigQuery.
Strong understanding of ETL/ELT, APIs, batch and streaming architectures.
Strong understanding of data governance, security, quality and lineage.
Experience working with senior business and technical stakeholders.
AI / GenAI Required
Understanding of AI-ready data architecture.
Practical exposure to GenAI / LLM / RAG architectures.
Understanding of vector databases and semantic/hybrid search.
Understanding of how enterprise data is consumed by AI applications and agents.
Good to Have
Data Mesh / Data Fabric experience.
Apache Kafka / event-driven architecture.
Apache Iceberg / Delta Lake.
Microsoft Fabric.
Databricks.
Snowflake.
Data Vault.
Knowledge Graphs / Graph databases.
Experience with MCP / Agentic AI architectures.
Experience with AI data governance.
Experience with ML/AI platforms and MLOps.
Terraform / Infrastructure as Code.
FinOps / cloud cost optimization.
Key Responsibilities
1. Data Architecture & Strategy
Define enterprise and solution-level data architecture strategies and roadmaps.
Design scalable architectures across Data Warehouse, Data Lake, Lakehouse, Data Fabric and Data Mesh patterns.
Evaluate build-vs-buy, technology and platform choices based on business, scalability, cost and performance requirements.
Define architecture standards, principles and reusable patterns.
2. Modern Data Platforms
Design cloud-native data platforms across AWS, Azure and/or GCP.
Architect modern lakehouse solutions using technologies such as:
oDatabricks oSnowflake oMicrosoft Fabric oBigQuery oDelta Lake / Apache Iceberg
Design batch, near-real-time and real-time data processing architectures.
3. Data Engineering & Integration
Define architecture for
oETL/ELT oAPI-based integration oStreaming pipelines oEvent-driven architectures oData ingestion and transformation
Work with Data Engineers to establish scalable and reusable pipeline patterns.
Define integration strategies across enterprise applications and data sources.
4. Data Modeling
Design
oConceptual, logical and physical data models o3NF models oDimensional models oStar/Snowflake schemas oData Vault where applicable
Define enterprise data models, master data and business semantics.
Establish standards for structured and unstructured data.
5. Data Governance & Security
Define enterprise data governance frameworks covering:
oData ownership oData quality oMetadata oData lineage oData classification oRetention oAccess control
Work with platforms such as Microsoft Purview, Unity Catalog, Collibra, Alation or equivalent.
Ensure compliance, privacy and security requirements are incorporated into architecture.
AI / GenAI Data Architecture
This should be a key differentiator for the modern Data Architect profile.
6. AI-Ready Data Architecture
Design data architectures that support Generative AI, RAG, AI copilots and Agentic AI.
Define how structured, unstructured and semi-structured enterprise data can be made available to AI systems.
Design data pipelines for AI ingestion, preprocessing, metadata and retrieval.
Define architecture for enterprise knowledge bases and AI-ready data products.
Modern AI systems increasingly require more than a basic vector database; production RAG architectures need ingestion, metadata, retrieval, evaluation, governance and access controls.
7.
RAG & Retrieval Architecture
Design enterprise RAG architectures using:
oEmbeddings oVector databases oHybrid search oMetadata filtering oSemantic search oRe-ranking
Evaluate technologies such as Pinecone, Azure AI Search, OpenSearch, PostgreSQL/pgvector, Databricks Vector Search or equivalent.
Define strategies for document ingestion, chunking, indexing, retrieval and data freshness.
8. Agentic AI Data Foundation
Design data access patterns for AI agents and multi-agent systems.
Define how agents securely access enterprise data, APIs and knowledge repositories.
Design structured and unstructured data interfaces suitable for agent consumption.
Understand concepts such as MCP, tool calling, agent memory and knowledge retrieval.
The shift toward AI-native lakehouse architectures is specifically driving requirements for continuous data, multimodal data and reliable context for AI agents.
Data Quality & Observability
Establish enterprise data quality frameworks and standards.
Define quality rules, validation frameworks and data contracts.
Implement monitoring for
oData freshness oCompleteness oAccuracy oConsistency oPipeline failures
Define observability and lineage requirements across the data platform.
Performance & Cost Optimization
Design architectures for high-volume and high-throughput workloads.
Optimize storage, compute and data processing costs.
Define partitioning, caching, indexing and query optimization strategies.
Evaluate cloud/data platform cost-performance trade-offs.
Establish architecture-level scalability and performance benchmarks.
AI Governance & Responsible Data Use
Ensure AI solutions follow enterprise data governance and security requirements.
Define access controls for data used by LLMs, RAG systems and AI agents.
Address
oPII / sensitive data oData leakage oAuthorization oData provenance oAuditability oRetention
Ensure AI systems access only the data users are authorized to access.
This is increasingly significant because AI/RAG governance needs to enforce authorization and data controls at runtime, not simply at the catalog level.
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Consultant-Data Architect / Modern Data & AI Architecture (Pune)
🏢 Harbinger Techventures
📍 Pune