Consultant (DBT) (Chennai)

Consultant (DBT) (Chennai)

04 Aug
|
GainInsights
|
Chennai

04 Aug

GainInsights

Chennai

Experience: 3 to 5 years

Location: Chennai / Bangalore

Job purpose:

Client-facing consultant responsible for designing and delivering secure, scalable, and cost-efficient data platforms using Python, PySpark (Apache Spark), dbt (ELT), and Apache Kafka for streaming. Owns end-to-end engineering across ingestion (batch + streaming), transformation, modelling, governance, performance tuning, CI/CD, and operations. Partners with stakeholders to translate business requirements into technical solutions, produce design artifacts, and ensure reliable delivery.

Duties and responsibilities of position:

General areas of responsibility relate to:

Lead/co-lead requirements workshops; translate business needs into user stories and architecture patterns combining PySpark, dbt, and Kafka.

Design secure, scalable, cost-optimized data solutions across lakehouse + warehouse and streaming paradigms.

Prepare HLD/LLD, architecture diagrams, data flow/mapping specifications, estimations, and runbooks.

Ensure compliance and governance (e.g., GDPR, HIPAA, SOC 2) and enterprise standards across batch and streaming.

Support production reliability through monitoring, incident response, and continuous improvement.

Collaborate with cross-functional teams (BI, Cloud Ops, Data Science) and mentor junior engineers.

Specific areas of responsibility include but are not limited to:

Design and implement Kafka producers/consumers (Python or JVM-based) with exactlyonce/at-least-once semantics per use case.

Configure topics, partitions, replication, schemas (Avro/JSON/Protobuf) with Schema Registry and compatibility rules.

Implement stream processing via Kafka Streams, Spark Structured Streaming (PySpark), or Flink (if applicable).

Build CDC pipelines from databases using Debezium/CDC connectors; manage offsets, replays, and dead-letter queues.

Enforce resilience patterns: idempotency, retries/backoff, circuit breakers, exactly-once sinks (Delta/Snowflake/etc.).

Build ingestion frameworks using Python (files/APIs) and PySpark for distributed processing; parameterize environments.

Standardize medallion zones (bronze/silver/gold) with file formats (Parquet/Delta/CSV/JSON) and metadata conventions.

Implement incremental loads and hybrid micro-batch strategies for streaming/batch convergence.

Develop scalable ETL/ELT using PySpark (DataFrames, Spark SQL) and dbt for modular SQL models (sources staging marts).

Implement SCD Type 1 & Type 2:

PySpark/Delta Lake MERGE with surrogate keys, effective/expiry dates, current flags,



audit columns.

dbt snapshots and incremental models with materializations (table/incremental).

Optimize jobs: partitioning, Z-ordering/bucketing, broadcast joins, predicate pushdown, and shuffle tuning.

Enforce coding standards: modular Python packages, parameterized notebooks/jobs, dbt macros and packages.

Design dimensional models (star/snowflake) with conformed dimensions and KPI-aligned fact tables.

Build dbt semantic layers with tests (unique/not null/relationships), docs, exposures, and data contracts.

Manage schema evolution and late-arriving data; ensure reproducibility with versioned models.

Apply RBAC in data platforms; manage secrets via Key Vault/Secrets Manager and environment variables.

Implement encryption in transit/at rest, data masking, and row/column-level policies in target systems.

Maintain lineage and catalog (dbt docs/manifest; Purview/Collibra/Atrium where available); steward metadata and tagging.

Profile PySpark jobs (stage/task metrics); fix data skew, excessive shuffles, small-file problems; tune cluster configs.

Tune dbt models (query plans, indices/distributions, materializations, incremental strategies).

Optimize Kafka workloads: topic partitioning, consumer group balancing, backpressure, batch size/linger.ms tuning.

Version control with Git; implement CI/CD for Python packages, PySpark jobs, dbt projects, and Kafka artifacts (schemas/config).

Parameterize dev/test/prod; manage secrets/variables; approvals; maintain rollback/DR strategies.

Automate quality gates: unit tests, data tests (dbt), schema compatibility checks, linting, and release notes.

Implement observability: job/app telemetry, logs, metrics dashboards for SLAs/SLOs, freshness, and latency (batch + streaming).

Set alerts for failures/anomalies; conduct RCA and remediation plans; track data downtime.

Optimize costs: autoscaling/serverless (where applicable), warehouse/cluster rightsizing, storage lifecycle, usage tagging/reporting.

Other duties:

Lead solution walkthroughs and trade-off analysis (PySpark vs. dbt, streaming vs. batch, Delta vs. warehouse).

Build POCs (CDC with Debezium,



SCD2 with Delta/dbt snapshots, Kafka throughput/latency benchmarks) to de-risk delivery.

Manage risks/assumptions, change requests, and stakeholder communications; maintain decisions & actions log.

Contribute to proposals, estimations, and SOWs for Python/PySpark/dbt/Kafka engagements.

Prepare executive reports on pipeline health, governance posture, streaming reliability, and cost optimization.

Prepare executive reports on pipeline health, governance posture, streaming reliability, and cost optimization.

Conduct internal knowledge-sharing; build accelerators (dbt macros, PySpark templates, Kafka schema standards); mentor associates.

Qualifications:

Bachelor’s degree in Computer Science, Information Technology, or a related field; or equivalent professional experience.

Experience: 4+ years in data engineering/BI consulting with strong hands-on work in Python, PySpark, dbt, and Kafka.

Python: data processing, packaging, logging, testing; API/file ingestion; CLI utilities.

PySpark: DataFrame APIs, Spark SQL, performance tuning (partitions, caching, joins, shuffle), structured streaming.

dbt: sources/staging/marts, materializations (table/incremental), snapshots, tests, docs; Jinja/macros; packages.

Kafka: producers/consumers, topics/partitions/replication, Schema Registry (Avro/JSON/Protobuf), Streams/Structured Streaming, offset management, DLQs.

CDC/SCD: incremental loads, watermarks, SCD1/SCD2 via Delta MERGE or dbt snapshots/incremental patterns; Debezium/CDC connectors (nice-to-have).

Modelling: star/snowflake, conformed dimensions, semantic consistency.

Orchestration/Tooling: Airflow/ADF/Prefect (any), Git, CI/CD (GitHub Actions/Azure DevOps), secrets management.

Storage/Compute: Lakehouse/warehouse (Delta/Parquet; Snowflake/Synapse/BigQuery/Databricks— any preferred).

Security/Governance: RBAC, secrets, encryption, masking/row policies; catalogue/lineage basics.

Hands-on experience in data modeling, indexing, and query performance tuning.

Experience with backup, disaster recovery, and business continuity planning.

Strong understanding of cloud security risks, including data breaches, malware, and social engineering.

Excellent collaboration skills for working with cross-functional teams and providing leadership in cloud security projects.

Experience working in a team, able to lead, prioritize and execute tasks in a high-pressure setting and provide or make sound decisions in emergencies.

📌 Consultant (DBT) (Chennai)
🏢 GainInsights
📍 Chennai

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: consultant (dbt) (chennai) / chennai

Subscribe to this job alert:

Get the latest job offers by email for: consultant (dbt) (chennai) / chennai