18 Sep
|
systems plus
|
Pune
Location: Pune Employment Type: full time
Job Details
ROLE OVERVIEW
We are seeking a motivated and technically proficient Senior Data Engineer to join our Data Engineering practice. In this role, you will be responsible for the end-to-end design, development, and optimization of scalable data pipelines, integration workflows, and analytical data solutions across cloud and modern data platforms. You will translate complex business requirements and architectural blueprints into robust, production-grade engineering solutions, contributing meaningfully to the quality and reliability of our enterprise data ecosystem.
We welcome candidates with 4 or more years of strong core data engineering experience. Proficiency in modern cloud platforms such as Snowflake or Microsoft Fabric is an added advantage, but not a prerequisite. What matters most is a solid foundation in SQL, data warehousing, ETL/ELT engineering, workflow orchestration, and API-based integrations, paired with a collaborative mindset and a drive to grow.
KEY RESPONSIBILITIES
Data Pipeline & Integration Development
Design, develop, and maintain scalable data pipelines and ETL/ELT workflows that meet performance, reliability, and maintainability standards.
Build and manage data ingestion, transformation, and integration solutions across structured and semi-structured data sources.
Implement REST and SOAP API integrations to connect enterprise systems, third-party services, and data platforms within the data ecosystem.
Implement data integration workflows using platforms such as IDMC CDI, CData, Azure Data Factory, Microsoft Fabric Data Factory (DF2), or equivalent tools.
Develop and optimize data models for data warehouses, data lakes, and analytical datasets, ensuring alignment with enterprise modeling standards.
Leverage Apache Kafka to design and implement event-driven and streaming data pipeline architectures.
Platform Engineering & Cloud Operations
Develop and optimize solutions on cloud data platforms; experience with Snowflake or Microsoft Fabric is a strong plus.
Collaborate with cloud infrastructure and platform teams to deploy, monitor, and manage data solutions in Azure-based environments.`
Use Apache Airflow for workflow orchestration and automation to enhance deployment efficiency and pipeline reliability.
Support data migration and modernization initiatives from legacy on-premises or traditional platforms to cloud-native architectures.
Data Quality, Governance & Security
Implement data quality validation, monitoring, alerting,
and error-handling mechanisms to ensure end-to-end pipeline integrity.
Embed data governance, security, and privacy controls within engineering solutions in compliance with organizational and regulatory standards.
Author technical documentation covering data flows, data dictionaries, architectural decisions, and operational runbooks.
Performance Optimization & Engineering Excellence
Write and optimize complex SQL queries and data processing logic for performance at enterprise scale.
Develop reusable frameworks, modular components, and engineering accelerators to improve team delivery velocity.
Conduct thorough code reviews and uphold adherence to engineering standards, coding best practices, and version control guidelines.
Diagnose and resolve production incidents, providing structured root-cause analysis and preventive recommendations.
Collaboration & Knowledge Sharing
Partner with data architects, product owners, business analysts, and cross-functional engineering teams to deliver data solutions aligned with business objectives.
Provide technical guidance and mentorship to junior engineers, fostering a culture of continuous learning and engineering excellence.
Participate in agile ceremonies, sprint planning, and technical design reviews, contributing to team-level delivery planning.
REQUIRED TECHNICAL SKILLS & QUALIFICATIONS
✓ MUST-HAVE — Core Engineering Foundation (4–7 Years Experience)
Strong SQL skills — complex queries, query optimization, execution plan analysis, and large-scale dataset management.
Solid data warehousing experience — dimensional modeling, star/snowflake schemas, data vault principles, and warehouse lifecycle management.
Hands-on ETL/ELT development using one or more tools such as Apache Airflow, Azure Data Factory, IDMC CDI, CData, or equivalent.
Apache Airflow orchestration — building, scheduling, and maintaining DAGs for robust workflow automation.
API integration experience — consuming and building REST/SOAP APIs to connect enterprise systems and external data sources.
Familiarity with CI/CD pipelines, version control (Git), and DevOps practices for data engineering workflows.
Working knowledge of data quality frameworks, observability tooling, and metadata management concepts.
Proven track record of delivering production-grade data pipelines and integration solutions.
- ADDED ADVANTAGE — Modern Cloud Platform Exposure
Hands-on experience with Snowflake — performance tuning, clustering, data sharing, and Snowpipe.
Practical experience with Microsoft Azure services and Microsoft Fabric (Lakehouse, Data Factory, Pipelines, OneLake).
Experience with IDMC or similar enterprise integration platform for large-scale data integration.
Experience implementing event streaming and real-time data pipelines using Apache Kafka.
Exposure to data governance principles and regulatory data privacy requirements (e.g. GDPR).
PREFERRED QUALIFICATIONS
Relevant cloud or data platform certifications (e.g., Snowflake SnowPro Core, Microsoft Certified: Azure Data Engineer Associate, or equivalent).
Experience working in Agile/Scrum delivery environments with cross-functional teams.
Exposure to DataOps practices and familiarity with pipeline automation and deployment best practices.
Familiarity with data lakehouse architectures (Delta Lake, Apache Iceberg) or open table format standards.
Experience with data observability or data cataloging platforms (e.g., Monte Carlo, Alation, Microsoft Purview).
AI & Machine Learning (Plus-to-Have)
The following AI/ML capabilities are not mandatory but are considered strong differentiators for candidates looking to grow in data-intensive, AI-augmented environments:
Programming proficiency in Python or R for data analysis, modeling, and pipeline scripting.
Familiarity with classical ML models such as linear regression, logistic regression, time-series forecasting (ARIMA, Prophet), and vector-based / embedding models for similarity and search use cases.
Exposure to Conversational Analytics and AI applications including Large Language Models (LLMs) and Natural Language Processing (NLP) models for analytics and insight generation.
Awareness or hands-on experience with deep learning frameworks, particularly TensorFlow, for building or consuming model-driven data pipelines.
CORE COMPETENCIES
Problem-solving and analytical thinking
Ownership and accountability for deliverables
Clear written and verbal communication
Collaborative cross-team partnership
Attention to detail and engineering rigor
Adaptability in fast-paced, evolving environments
Mentorship and knowledge-sharing mindset
Continuous learning orientation
📌 Data Engineer (Pune)
🏢 systems plus
📍 Pune