Technical Skills
- Expertise in DBT (preferred) for data modeling.
- Strong SQL skills with hands-on experience in: o Snowflake o Azure SQL o Stored procedures
- Proficiency in Python for data engineering workflows.
- Hands-on experience with: o Databricks & PySpark o Azure Data Services (ADF, ADLS, Synapse)
- Robust knowledge of Snowflake architecture and schema design.
- Experience with data validation, quality frameworks, and analysis tools.
- Familiarity with GitHub and CI/CD practices. Role demands a highly skilled Data Engineer to design, build, and optimize scalable data pipelines and data platforms. The ideal candidate will have strong expertise in data modeling, cloud-based data architectures, and modern data engineering tools across Azure, Snowflake, and Databricks environments. Key Responsibilities Data Engineering & Pipeline Development
- Design, develop, and maintain robust ETL/ELT pipelines using Databricks, PySpark, and Azure Data Factory (ADF).
- Build scalable and efficient data ingestion frameworks for structured and unstructured data.
- Optimize pipeline performance through performance tuning and orchestration best practices. Data Modeling & Management
- Develop and maintain data models using modern tools (DBT preferred).
- Implement Master Data Management (MDM) solutions to ensure data consistency and integrity.
- Design scalable and efficient Snowflake schemas (star/snowflake schema, dimensional modeling). Database & Query Optimization
- Write and optimize advanced SQL queries across Snowflake, Azure SQL, and Synapse.
- Develop and manage stored procedures and database objects.
- Ensure efficient data retrieval through indexing, partitioning, and query optimization. Cloud & Platform Integration
- Work with Azure data services including: o Azure Data Factory (ADF) o Azure Data Lake Storage (ADLS) o Azure Synapse Analytics o Azure SQL Database
- Integrate and maintain Snowflake with Azure ecosystem. Python Development
- Develop data transformation and automation scripts using Python libraries: o pandas o pyodbc o SQLAlchemy
- Build reusable components for data processing and validation. Data Quality, Validation & Monitoring
- Implement data validation rules, quality checks, and anomaly detection frameworks.
- Perform root cause analysis for data inconsistencies.
- Develop dashboards or tools for data quality monitoring. Collaboration & DevOps
- Use GitHub for version control, branching strategies, and code reviews.
- Manage workload scheduling and dependency management for pipelines.
- Collaborate with cross-functional teams including data analysts, data scientists, and business stakeholders. Required Skills & Qualifications
- Bachelor's or Master's degree in Computer Science, Information Systems, or related field.
- Strong experience in data engineering and data platform development. Preferred Qualifications
- Experience implementing data governance and MDM solutions.
- Knowledge of performance tuning in distributed processing systems.
- Familiarity with workflow orchestration tools.
- Experience in agile environments (SCRUM/Kanban).
📌 Data Engineer (Pune)
🏢 Infosys
📍 Pune