About the Role
We are looking for a
Big Data Engineer
to design, build, and operate large-scale data pipelines and analytical infrastructure that transform high-volume raw data into reliable, query-ready datasets for analytics, reporting, and data-driven products.
Our data platform ingests and processes data from multiple sources and serves analytics, data science, product, and downstream applications. In this role, you will own data pipelines end-to-end—from ingestion and transformation to warehousing, orchestration, data quality, and observability.
A key part of the role is owning
ClickHouse as our primary analytical data store
. You will be responsible for designing scalable data models, optimizing query performance, and ensuring the platform remains reliable and cost-effective as data volumes and workloads grow.
You will work closely with
data scientists, analysts, product engineers, and other engineering teams
to build a modern, cloud-native data platform.
What You'll Do
- Design, build, and maintain robust
batch and streaming data pipelines
that ingest data from multiple sources into analytical data stores.
- Build and operate
Apache Airflow DAGs
, including scheduling, dependencies, retries, backfills, idempotency, concurrency, and failure handling.
- Develop analytics-ready datasets using
dbt
, following well-structured staging, intermediate, and mart layers with appropriate tests, documentation, and incremental models.
- Own
ClickHouse
as the primary analytical store, including:
- Schema and table design using the MergeTree family of engines
- Partitioning and sorting/primary key strategies
- Materialized views
- Distributed and replicated table architectures
- Query and memory optimization
- High-volume data ingestion and performance tuning
- Work with
BigQuery
where cloud data-warehouse patterns are appropriate, including data modeling and query/cost optimization.
- Design and operate
NoSQL and key-value data stores
, including Bigtable, DynamoDB, an
📌 Big Data Engineer (India)
🏢 Wingify
📍 India