09 Sep
|
Infinityquest It Services
|
Pune
09 Sep
Infinityquest It Services
Pune
We are looking for an experienced Data Engineer who will play a key role in modernizing our data platform by analyzing existing ETL pipelines and migrating them to the new platform built on Databricks. The ideal candidate should have strong hands-on experience with data engineering best practices, AWS ecosystem, and Databricks-based transformations.
Key Responsibilities
Analyze and understand existing ETL pipelines, data sources, and transformation logic across current platforms
Drive end-to-end migration of data pipelines to the new platform using Databricks
Design, build, and maintain scalable data pipelines for batch and near real-time processing in the new platform
Re-engineer legacy ETL/ELT workflows into optimized, scalable, and modular pipelines using Spark/Databricks
Integrate data from multiple heterogeneous sources and ensure seamless data flow across systems
Implement data transformation logic using PySpark/SQL on Databricks and ensure adherence to best practices
Validate migrated pipelines to ensure data accuracy, reconciliation, and consistency with source systems
Optimize pipelines for performance, scalability, and cost efficiency within cloud environments
Work closely with business stakeholders, analysts, and data teams to understand data requirements and use cases
Support reporting, analytics, and downstream applications by providing reliable and high-quality datasets
Monitor production pipelines, troubleshoot failures, and proactively resolve data-related issues
Ensure proper logging, alerting, and observability for all pipelines (using tools like CloudWatch / Databricks monitoring)
Follow data governance, security,
and compliance standards during data handling and migration
Implement CI/CD pipelines for data workflows and maintain proper documentation of pipelines and processes
Required Skills & Qualifications
Strong proficiency in SQL and relational databases (Oracle, PostgreSQL, etc.)
Hands-on experience with ETL/ELT pipeline development and migration projects
Strong programming skills in Python and/or Scala
Good experience with Apache Spark (preferably PySpark)
Hands-on experience with Databricks (Delta Lake, notebooks, jobs, workflows)
Strong knowledge of AWS data engineering services including:
AWS Glue, Lambda, S3, EventBridge, SQS
Amazon Redshift, Firehose, CloudWatch
Strong understanding of data modeling, data warehousing concepts, and data lake architectures
Experience working with large-scale datasets and distributed processing
Knowledge of pipeline orchestration and workflow management
Preferred Skills (Good to Have)
Experience with data migration to contemporary data platforms (e.g., Databricks, Lakehouse architecture)
Understanding of CDC (Change Data Capture), Delta Lake, and incremental processing patterns
Experience with monitoring tools (Grafana, CloudWatch)
Familiarity with CI/CD tools (Azure DevOps, Git-based workflows)
Key Expectations from Candidate
Ability to quickly understand existing systems and reverse-engineer data pipelines
Strong problem-solving skills to modernize and optimize legacy data processes
Ownership mindset to drive migration tasks end-to-end with minimal supervision
Strong collaboration and communication skills to work across cross-functional teams
📌 Data Engineer (Pune)
🏢 Infinityquest It Services
📍 Pune