20 Sep
|
The UrbanIndia
|
Bengaluru
20 Sep
The UrbanIndia
Bengaluru
Experience: 3–5 years in Data Engineering, specifically with distributed systems and cloud-native architectures
Coding: Expert-level Python/PySpark and SQL. Familiarity with Go/Java/Scala is a plus
Infrastructure: Hands-on experience with AWS (S3, EKS, MSK) and Infrastructure-as-Code
Orchestration: Experience with Airflow or Temporal for complex workflow management
AI-Native: Proficiency in using AI tools (Claude, Codex, Copilot) to write, test, and document code efficiently
Systems Thinking: Ability to explain the trade-offs between different storage formats and processing frameworks
Tech Execution : Drive key tech initiatives by preparing TRD and actively involve in design reviews
Domain Modelling - Should be hands on in designing Domain models for OLAP like Fact, Dimension, Cumulative , types of SCD’s and OBT pattern tables
Self Starter - Lead the team technically and bring in new ideas to contribute to the growth of the charter
Stakeholder Interaction - Interact with the Product & Key Stakeholders & help them by adding value to the business workflow with data & analytics
Good to have -
Real-time CDC: Ownership of high-throughput ingestion from RDBMS to Lakehouse using Debezium, PeerDB
Lakehouse Architecture: Designing and optimizing table formats (Iceberg, Delta, Hudi) for both performance and storage efficiency
Unified Compute: Developing robust ETL/ELT frameworks in PySpark and Flink (handling both batch and streaming workloads)
Infrastructure & Ops: Managing data workloads on AWS (EMR, EKS, MSK, S3) and automating everything via Gitlab/Github Actions
Query & BI: Tuning Trino or Clickhouse to power real-time dashboards in Metabase, Superset, and PowerBI
Our Tech Stack-
Ingestion & CDC: OLake and PeerDB for near real-time sync from production systems into Alchemy; Debezium/Kafka for CDC-heavy use cases, with support for Kafka and S3-based sources
Lakehouse / Storage: Alchemy on Apache Iceberg with S3 as the data lake storage layer and AWS Glue Catalog for metadata; exposure to Delta/Hudi is a plus
Processing & Compute: PySpark on EMR/EKS for batch and streaming workloads; Flink and Spark Structured Streaming fundamentals for low-latency pipelines
Streaming Platform: MSK / Kafka for event-driven ingestion, CDC propagation, replay, backfills, and operational monitoring through Kafka UI
Query & Serving Layer: Trino over Alchemy/Iceberg for lakehouse analytics, ClickHouse for high-throughput operational and real-time dashboards, and BigQuery exposure where applicable
Workflow Orchestration: Airflow for scheduled data pipelines, DQ/reconciliation DAGs, backfills, and SLA-driven jobs; Temporal for durable workflow execution in ingestion services
Data Quality & Governance: DQ checks, freshness/SLA monitoring, source-to-lake reconciliation, deduplication, schema evolution handling, and cataloging/lineage through OpenMetadata
Infrastructure & DevOps: AWS (S3, EKS, EMR, MSK), Kubernetes, Terraform/IaC, GitLab/GitHub Actions, observability via Grafana/CloudWatch, and production runbook discipline
BI & Analytics: Metabase, Superset, Tableau, and PowerBI for business dashboards; solid ability to model curated marts, fact/dimension tables, SCDs, and OBT patterns for stakeholder-facing analytics
Skills: analytics,kafka,cdc,s3,infrastructure,data,alchemy,storage,aws,workflow
📌 Data Engineer (Bengaluru)
🏢 The UrbanIndia
📍 Bengaluru