08 Aug
|
Smart IMS
|
Bengaluru
08 Aug
Smart IMS
Bengaluru
Role & Responsibilities
- Design, develop, and maintain scalable ETL/ELT pipelines to process large volumes of structured and unstructured data.
- Write, optimize, and review complex SQL queries for high-performance data processing and analytics.
- Build and manage robust data models, dimensional models (Star/Snowflake Schema), and enterprise data warehouse solutions.
- Develop scalable data pipelines using Python, Apache Spark/PySpark, and cloud-native services.
- Work extensively with cloud data platforms such as GCP (BigQuery, Dataflow, Dataproc), Snowflake, Databricks, or Amazon Redshift.
- Ensure data quality, integrity, reconciliation, and validation across multiple data sources.
- Optimize query performance, partitioning, clustering, indexing, and overall data processing efficiency.
- Collaborate with Data Scientists, Analysts, Product Managers, and Business stakeholders to understand data requirements and deliver reliable datasets.
- Design reusable data frameworks, automation solutions, and monitoring for production pipelines.
- Perform code reviews, mentor junior engineers, and establish engineering best practices for SQL, Python, and Spark development.
- Troubleshoot production issues, perform root cause analysis, and implement preventive measures.
- Follow CI/CD, version control (Git), and DevOps best practices for data engineering workflows.
- Document technical designs, data lineage, and architecture for maintainability and governance.
- Ensure adherence to data security, compliance, and governance standards.
Preferred Candidate Profile
- 57 years of hands-on experience in Data Engineering, Big Data,
or Data Platform development.
- Strong expertise in Advanced SQL with experience in writing complex joins, CTEs, window functions, query optimization, and performance tuning.
- Strong programming skills in Python for data processing and automation.
- Hands-on experience with Apache Spark/PySpark for large-scale distributed data processing.
- Experience with cloud platforms, preferably Google Cloud Platform (GCP), including BigQuery, Dataflow, Dataproc, Cloud Storage, and Composer.
- Practical experience with modern data warehouse technologies such as BigQuery, Snowflake, Databricks, or Amazon Redshift.
- Strong understanding of ETL/ELT, data integration, workflow orchestration, and batch/streaming data pipelines.
- Positive knowledge of Data Modeling, including Star Schema, Snowflake Schema, and dimensional modeling concepts.
- Experience with workflow orchestration tools such as Apache Airflow or Cloud Composer.
- Familiarity with modern data lake/lakehouse technologies such as Delta Lake, Apache Iceberg, or Apache Hudi is an added advantage.
- Experience working with version control systems (Git) and CI/CD pipelines.
- Strong analytical, debugging, and problem-solving skills with a focus on data quality and reliability.
- Experience conducting code reviews, mentoring junior engineers, and driving technical best practices.
- Excellent communication and stakeholder management skills with the ability to collaborate across cross-functional teams.
- Bachelor's or Master's degree in Computer Science, Information Technology, or a related engineering discipline.
📌 Senior Data Engineer (Bengaluru)
🏢 Smart IMS
📍 Bengaluru