21 Aug
|
Sourcebae
|
Bengaluru
21 Aug
Sourcebae
Bengaluru
Role summary
We are seeking a Senior Data Engineer to design and build reliable batch and streaming data platforms. The role requires advanced Python and SQL, mandatory hands-on Apache Spark experience, strong data modelling and practical expertise with at least one cloud platform. The engineer will own ingestion, transformation, quality, orchestration, performance and operational reliability.
Key responsibilities
- Design and build scalable ETL/ELT pipelines for structured, semi-structured and streaming data.
- Develop distributed data-processing workloads using Apache Spark and Python/PySpark.
- Implement ingestion, transformation, validation, reconciliation and publishing workflows.
- Design cloud data lake, lakehouse and data-warehouse components.
- Implement incremental loading, change data capture, schema evolution and idempotent processing.
- Optimise Spark jobs, SQL queries, storage formats, partitioning and cloud resource utilisation.
- Build data-quality checks, monitoring, alerting, retry and recovery mechanisms.
- Implement orchestration and dependency management for production data pipelines.
- Collaborate with analytics, AI/ML and application teams to deliver trusted data products.
- Maintain documentation covering source-to-target mapping, lineage, operational runbooks and architecture.
Required qualifications and experience
- 5+ years of relevant experience in data engineering, big-data engineering or cloud data platforms.
- Robust Python and advanced SQL skills.
- Mandatory hands-on experience with Apache Spark, including performance tuning and distributed-processing concepts.
- Experience with at least one cloud platform: AWS, Microsoft Azure or Google Cloud Platform.
- Experience building and supporting production-grade data pipelines.
- Strong understanding of data modelling, data quality, orchestration and operational reliability.
📌 Data Engineer (Bengaluru)
🏢 Sourcebae
📍 Bengaluru