09 Aug
|
CriticalRiver
|
Hyderabad
09 Aug
CriticalRiver
Hyderabad
Role Overview
We are looking for a Principal Architect, Data & AI/ML who is equally comfortable whiteboarding an enterprise data strategy and rolling up their sleeves to write production-grade code, design reference architectures, and mentor delivery teams. This is a hands-on leadership role that sits at the intersection of data engineering, machine learning engineering, and solution architecture.
You will serve as the technical anchor across client engagements — owning architecture decisions end-to-end, driving pre-sales pursuits, and building reusable accelerators that amplify the practice's delivery capability. The ideal candidate has deep expertise across the full data & AI/ML value chain — from data foundations to decision intelligence.
Key Responsibilities
Solution Architecture & Technical Leadership
- Design and own end-to-end reference architectures for data platforms, lakehouses, AI/ML pipelines, and GenAI products across multi-cloud environments (AWS, Azure, GCP).
- Lead architecture reviews, proof-of-concepts, and technical due-diligence for client engagements.
- Define architectural principles, design patterns, and guardrails for the practice; contribute to CriticalRiver's internal IP and accelerator library.
- Translate business requirements into scalable, cost-optimised, and secure technical blueprints.
Generative AI & LLM-Ops
- Architect and build enterprise-grade Generative AI solutions using large language models (LLMs), retrieval-augmented generation (RAG), vector databases, and AI agents.
- Design LLM-Ops pipelines covering fine-tuning, prompt engineering, evaluation harnesses, guardrails, and model observability.
- Enable "no-data" and "agents-on-data" patterns — embedding AI agents into structured data workflows and decision processes.
Data Engineering & Integration
- Architect modern data integration pipelines: batch, micro-batch, and real-time; ELT/ETL on cloud-native platforms.
- Lead lakehouse and warehouse modernisation engagements — design migration patterns, medallion architectures, and data contract frameworks.
- Define data pipeline standards using tools such as dbt, Apache Airflow, AWS Glue, Azure Data Factory, and Fivetran.
Data Warehouse & Lakehouse
- Provide deep expertise on Snowflake, Databricks, and cloud-native warehouses (BigQuery, Synapse, Redshift).
- Design multi-hop storage architectures (Bronze/Silver/Gold), partition strategies, indexing, and query optimisation.
- Guide clients through platform selection, TCO analysis, and migration roadmaps.
Data Streaming & Edge Computing
- Architect streaming and event-driven platforms using Apache Spark Structured Streaming, Kafka, Flink,
and cloud-native event services.
- Design edge-to-cloud IoT data pipelines — from device ingestion to real-time analytics and actionable insights.
- Define SLA/SLO frameworks for latency-sensitive workloads.
AI/ML Enablement & MLOps
- Oversee the full ML lifecycle: feature engineering, model development, training infrastructure, hyperparameter tuning, deployment, and drift monitoring.
- Design MLOps platforms using MLflow, Kubeflow, SageMaker, Azure ML, or Vertex AI.
- Champion responsible AI practices — fairness, explainability, bias detection, and model governance.
Data Science & Advanced Analytics
- Guide data science teams on forecasting, prescriptive analytics, and decision-intelligence models that connect directly to business outcomes.
- Architect feature stores, experiment tracking systems, and model registries.
- Evaluate and adopt emerging frameworks for causal inference, simulation, and reinforcement learning as applicable.
BI, Semantic Layer & Self-Service Analytics
- Design semantic layers, metrics frameworks, and governed BI architectures that enable self-service analytics at scale.
- Advise on BI platform selection and implementation: Power BI, Tableau, Looker, ThoughtSpot, and headless BI tools.
- Define data products and data mesh principles — domain ownership, data contracts, and discoverability.
Data Strategy & Governance
- Co-develop data strategy, roadmaps, and operating models with client CDOs, CTOs, and data leadership.
- Design and implement data governance frameworks covering data quality, data lineage, cataloguing, master data management (MDM), and privacy/compliance.
- Champion data literacy and centre-of-excellence models within client organisations.
Pre-Sales & Practice Development
- Support business development — respond to RFPs, present technical architecture in client pitches, and estimate delivery effort.
- Publish thought leadership: blogs, white papers, reference architectures, and conference talks.
- Mentor senior engineers and architects; drive the internal CoE agenda across upskilling and certification.
Required Qualifications & Skills
Experience
- 15+ years in data engineering, analytics, or AI/ML roles, with at least 4 years in a principal or lead architect capacity.
- Proven track record of delivering large-scale data platforms and AI/ML solutions in complex enterprise environments.
- Hands-on coding is mandatory — architecture without execution is not sufficient for this role.
Core Technical Proficiency (hands-on expected)
- Languages: Python, SQL, Scala (Must Have)
- Data platforms: Snowflake, Databricks, BigQuery, Synapse, or Redshift — at least two in depth.
- Data integration: dbt, Airflow, Spark, Kafka, ADF, Glue, Fivetran, or equivalent.
- ML frameworks: scikit-learn, XGBoost, PyTorch, TensorFlow; MLOps tools: MLflow, Kubeflow, SageMaker, or Vertex AI.
- GenAI & LLM: LangChain / LlamaIndex, OpenAI / Azure OpenAI / Bedrock / Gemini APIs, vector DBs (Pinecone, Weaviate, pgvector).
- Cloud: AWS, Azure, or GCP — certified preferred; multi-cloud experience strongly valued.
- Infrastructure-as-Code: Terraform, Bicep, or CloudFormation.
- Containerisation & orchestration: Docker, Kubernetes.
- Snowflake and Databricks (Must Have)
Architecture & Design
- Solid command of data modelling (dimensional, Data Vault 2.0, entity-centric), API design, and microservices patterns.
- Experience with data mesh, data fabric, and event-driven architecture patterns.
- Ability to produce detailed architecture documents, C4 diagrams, and ADRs independently.
Soft Skills & Leadership
- Excellent executive-level communication — ability to present complex technical topics clearly to C-suite and non-technical stakeholders.
- Strong consulting mindset: structured problem-solving, rapid context switching, and client-facing delivery confidence.
- Self-starter with the ability to work in ambiguous, rapid-paced environments.
Good to Have
- Relevant cloud and data certifications: AWS Solutions Architect Professional, Azure Data Engineer / AI Engineer, GCP Professional Data Engineer, Databricks Certified Associate/Professional, Snowflake SnowPro Core.
- Contribution to open-source data or ML projects.
- Experience in regulated industries (BFSI, Healthcare, Retail) with data privacy compliance (GDPR, HIPAA).
- Exposure to graph databases, time-series platforms, or geospatial analytics.
- Prior consulting or Big-4 technology advisory background.
What We Offer
- Opportunity to shape the technical direction of a rapidly growing Data & AI/ML practice.
- Exposure to cutting-edge client challenges across industries and geographies.
- Access to the latest cloud, data, and GenAI platforms and tooling.
- Competitive compensation with performance-linked incentives.
- Sponsored certifications and continuous learning budget.
- Collaborative, high-trust culture with strong engineering values.
📌 Practice Architect – Data & AI/ML - Snowflake and Databricks (Hyderabad)
🏢 CriticalRiver
📍 Hyderabad