Key Responsibilities
- Design and implement Medallion Architecture (Bronze / Silver / Gold) pipelines with data quality gates, transformation logic, and lineage tracking at production scale.
- Lead end-to-end POCs for HVR, Oracle GoldenGate, AWS DMS, Dremio, and Snowflake — benchmark performance and deliver clear go/no-go recommendations.
- Write production-quality Python, SQL, and PySpark code; pair-program with data engineers on complex pipeline and integration challenges.
- Own architecture decisions end-to-end — recommend solutions backed by trade-off analysis and TCO models; push back on misaligned vendor proposals.
- Define and maintain CI/CD pipelines for data workflows using GitHub Actions, AWS CodePipeline — covering automated testing, deployment, and promotion across Bronze / Silver / Gold environments.
- Define IaC templates (Terraform / AWS CDK), coding standards, and reusable frameworks that accelerate team delivery.
- Integrate data quality and observability tooling (Great Expectations, Monte Carlo) into ingestion pipelines.
- Produce Architecture Decision Records (ADRs) and participate in design reviews, stand-ups, and architecture calls in PST hours.
Required Qualifications
- 8+ years in data engineering / architecture; 4+ years hands-on with AWS data services.
- Proven delivery of Medallion Architecture pipelines at production scale.
- Hands-on CDC experience — HVR, Oracle GoldenGate, AWS DMS, or Debezium.
- Practical experience with PowerBI, Dremio, Snowflake, Apache Iceberg, or Delta Lake on AWS S3.
- Deep expertise in AWS S3, Glue, Redshift, Athena, Kinesis, Step Functions, Lake Formation.
- Experience with Terraform or AWS CDK for infrastructure provisioning.
Preferred Qualifications
- AWS Certified Solutions Architect – Qualified or AWS Certified Data Analytics Specialty.
- Experience with Apache Kafka / Amazon MSK for streaming ingestion alongside batch CDC.
- Prior experience in US–India distributed team models.
- Familiarity with dbt, Apache Airflow, or DataOps CI/CD practices.
- AI Agent development experience — hands-on exposure to building, deploying, or orchestrating AI agents using frameworks such as LangChain, LangGraph, AWS Bedrock Agents, AutoGen, or Agent Core; including agentic workflows for data pipeline automation, anomaly detection, metadata management, or self-healing ingestion processes.
📌 AWS Data Architect (India)
🏢 DynPro
📍 India