09 Sep
|
Scaletrix.AI
|
Gurugram
09 Sep
Scaletrix.AI
Gurugram
## About the Role
We are looking for a highly skilled Data Engineer with 4–6 years of experience in designing, building, and maintaining scalable data platforms and data pipelines. The ideal candidate should possess strong expertise in cloud-based data engineering, ETL/ELT development, data warehousing, and big data technologies. The candidate will work closely with Data Scientists, Analysts, BI teams, and business stakeholders to enable reliable and efficient data solutions.
## Key Responsibilities
Design, develop, and maintain scalable ETL/ELT pipelines using Python, PySpark, SQL, Pandas, and DBT.
Build and optimize batch and real-time data processing workflows using Apache Airflow and Databricks.
Develop and manage cloud-based data lakes and data warehouses using AWS services such as S3, EMR, Redshift, Lambda, and RDS.
Implement and maintain Snowflake data warehouse solutions, including Snowpipe, external stages, and data-sharing capabilities.
Develop robust data ingestion frameworks for structured and semi-structured data from multiple sources.
Design and implement event-driven architectures using AWS Lambda, Kafka, S3 Event Notifications, and other cloud-native services.
Collaborate with cross-functional teams including Data Science, Analytics, BI, and Product teams to support data requirements.
Build and expose APIs for data integration and data-sharing use cases.
Ensure data quality, governance, privacy, and compliance through validation, masking, and monitoring frameworks.
Implement CI/CD pipelines for data engineering projects using GitHub Actions, GitLab, JFrog, and automated testing frameworks.
Optimize data pipelines for performance, scalability, and cost efficiency.
Create operational dashboards and reporting solutions using Tableau or similar BI tools.
## Required Skills & Qualifications
### Technical Skills
Strong programming skills in Python and SQL.
Hands-on experience with PySpark, Pandas, and large-scale data processing.
Expertise in Databricks, DBT, and Apache Airflow.
Experience with Snowflake Data Warehouse and Snowpipe.
Strong understanding of AWS ecosystem including:
S3
EMR
Redshift
Lambda
EC2
RDS
Experience working with Apache Kafka and event-driven data architectures.
Hands-on experience with PostgreSQL, MongoDB, and relational databases.
Knowledge of Docker and containerized deployments.
Experience with REST APIs and data integration frameworks.
Familiarity with GitHub, GitLab, CI/CD pipelines, and automated testing practices.
Understanding of Data Lake and Data Warehouse architectures.
### Preferred Qualifications
Experience in Healthcare, Life Sciences, Retail, or eCommerce domains.
Knowledge of data governance, security, and compliance best practices.
Exposure to cloud certifications (AWS, OCI, Azure, or GCP).
Experience working in Agile/Scrum environments.
## Education
Bachelor's Degree in Computer Science, Information Technology, Engineering, or a related field.
## Nice to Have
Experience with DuckDB.
Exposure to Generative AI, Machine Learning, or MLOps platforms.
Experience designing enterprise-scale data platforms and modern lakehouse architectures.
## What Success Looks Like
Deliver highly reliable and scalable data pipelines.
Ensure high-quality and trusted data for analytics and business decision-making.
Improve data processing efficiency and reduce operational overhead.
- Contribute to the evolution of the organization's up-to-date data platform strategy.
📌 Data Engineer (Gurugram)
🏢 Scaletrix.AI
📍 Gurugram