Educational Requirements
- Bachelor of Engineering, Bachelor Of Technology, Master Of Technology, Master Of Engineering
Service Line
- Engineering Services
Responsibilities
Role demands a highly skilled Data Engineer to design, build, and optimize scalable data pipelines and data platforms. The ideal candidate will have strong expertise in data modeling, cloud-based data architectures, and contemporary data engineering tools across Azure, Snowflake, and Databricks environments.
- Data Engineering Pipeline Development
- Design, develop, and maintain robust ETL/ELT pipelines using Databricks, PySpark, and Azure Data Factory (ADF).
- Build scalable and efficient data ingestion frameworks for structured and unstructured data.
- Optimize pipeline performance through performance tuning and orchestration best practices.
- Data Modeling Management
- Develop and maintain data models using modern tools (DBT preferred).
- Implement Master Data Management (MDM) solutions to ensure data consistency and integrity.
- Design scalable and efficient Snowflake schemas (star/snowflake schema, dimensional modeling).
- Database Query Optimization
- Write and optimize advanced SQL queries across Snowflake, Azure SQL, and Synapse.
- Develop and manage stored procedures and database objects.
- Ensure efficient data retrieval through indexing, partitioning, and query optimization.
- Cloud Platform Integration
- Work with Azure data services including:
- Azure Data Factory (ADF)
- Azure Data Lake Storage (ADLS)
- Azure Synapse Analytics
- Azure SQL Database
- Integrate and maintain Snowflake with Azure ecosystem.
- Python Development
- Develop data transformation and automation scripts using Python libraries: pandas, pyodbc, SQLAlchemy.
- Build reusable components for data processing and validation.
- Data Quality, Validation Monitoring
- Implement data validation rules, quality checks, and anomaly detection frameworks.
- Perform root cause analysis for data inconsistencies.
- Develop dashboards or tools for data quality monitoring.
- Collaboration DevOps
- Use GitHub for version control, branching strategies, and code reviews.
- Manage workload scheduling and dependency management for pipelines.
- Collaborate with cross-functional teams including data analysts, data scientists, and business stakeholders.
Required Skills Qualifications
- Bachelors or Masters degree in Computer Science, Information Systems, or related field.
- Strong experience in data engineering and data platform development.
Additional Responsibilities:
Preferred Qualifications
- Experience implementing data governance and MDM solutions.
- Knowledge of performance tuning in distributed processing systems.
- Familiarity with workflow orchestration tools.
- Experience in agile environments (SCRUM/Kanban).
Technical and Skilled Requirements:
- Technical Skills
- Expertise in DBT (preferred) for data modeling.
- Strong SQL skills with hands-on experience in:
- Snowflake
- Azure SQL
- Stored procedures
- Proficiency in Python for data engineering workflows.
- Hands-on experience with:
- Databricks PySpark
- Azure Data Services (ADF, ADLS, Synapse)
- Strong knowledge of Snowflake architecture and schema design.
- Experience with data validation, quality frameworks, and analysis tools.
- Familiarity with GitHub and CI/CD practices.
Preferred Skills:
- Technology->Cloud Platform->Databases on Azure->Azure SQL Database
- Technology->Cloud Integration->Azure Data Factory (ADF)
- Technology->Oracle->PL/SQL
- Technology->Cloud Platform->Azure Analytics Services->Azure Data Lake
- Technology->Data on Cloud-DataStore->Snowflake
- Technology->Data Engineering->Databricks
- Technology->Big Data - Data Processing->PySpark
📌 Data Engineer (Pune)
🏢 Infosys
📍 Pune