We are seeking a skilled and proactive Data Engineer to design, build, and maintain robust, scalable, and optimized data pipelines and architectures. In this role, you will collaborate closely with Data Scientists, Business Intelligence Analysts, and Software Engineers to transform raw enterprise data into clean, accessible, and high-performance data models.
The ideal candidate has hands-on experience with modern cloud data platforms, distributed computing, streaming and batch ETL/ELT pipelines, data modeling, and robust data governance practices.
2. Key Responsibilities
Pipeline Development & Architecture
• Design & Build ETL/ELT Pipelines: Develop scalable batch and real-time streaming data ingestion and transformation pipelines from diverse sources (APIs, transactional DBs, Kafka, logs, flat files).
• Data Modeling & Storage: Architect and optimize dimensional data models (Star/Snowflake schemas, One Big Table), data lakes, data lakehouses, and relational/NoSQL data stores.
• Workflow Orchestration: Implement and maintain pipeline scheduling, dependency management,
and monitoring using contemporary orchestrators (e.g., Apache Airflow, Prefect, Dagster).
Performance, Reliability & Scalability
• Query & Storage Optimization: Fine-tune SQL queries, indexing, partitioning, caching, and distributed compute workloads (e.g., Spark, Trino, DuckDB).
• Data Quality & Testing: Implement automated data validation, anomaly detection, and schema evolution testing (e.g., Great Expectations, Soda, dbt tests).
• Monitoring & Alerting: Build observability frameworks for pipeline uptime, latency, SLA adherence, and data freshness.
Collaboration & Governance
• Data Governance & Security: Enforce enterprise data security, role-based access control (RBAC), data lineage, and compliance standards (GDPR, CCPA, SOC2, HIPAA).
• Stakeholder Enablement: Partner with downstream data consumers (BI developers, ML engineers, business stakeholders) to understand data requirements and deliver self-ser
📌 Data Engineer (Pune)
🏢 Citi
📍 Pune