11 Sep
|
CloudBoson
|
Noida
Role & responsibilities
- Design, build, and optimize ETL/ELT pipelines using PySpark on Azure Databricks.
- Develop real-time and near-real-time ingestion pipelines using Spark Structured Streaming and Azure Event Hubs
- Develop and tune complex SQL models in BigQuery for analytics and reporting
- Build ingestion pipelines from APIs and third-party sources into Azure Data Lake / BigQuery
- Implement incremental loads, CDC, and MERGE/upsert patterns for large datasets.
- Experience on data lake/warehouse solutions across GCP and Azure
- Build dashboards and reporting layers in Looker Studio
- Lead code reviews, mentor engineers, and drive Git-based CI/CD practices for data pipelines
- Partner with product/business teams to translate requirements into scalable data models
- Optimize cost & performance of already running queries
- Carefully plan and implement data load requirement.
- Regular monitoring of the data pipelines and systems and maintains the uptime of the overall data availability.
- Co-ordinate with stake holders and provide relevant data and business insights.
Required Skills:
- 5+ years in data engineering
- Strong SQL (joins, window functions, query optimization)
- Well versed in Python and PySpark (Azure Databricks)
- Spark Structured Streaming for real-time data processing
- Azure Event Hubs for event ingestion and streaming architecture
- BigQuery , AWS DynamoDB, Azure Storage Tables, Azure SQL, Azure Event Hub.
- Azure Data Lake (ADLS Gen2)
- Multi-cloud experience: GCP + Azure
- Looker Studio for BI/reporting
- Git version control
- REST API integration (auth, pagination, error handling)
Positive to have:
- Orchestration tools (Airflow, Databricks Workflows)
- Data quality/governance frameworks
- Cloud cost optimization experience
📌 Senior Data Engineer (Noida)
🏢 CloudBoson
📍 Noida