Role Summary
We are seeking a highly skilled Data Engineer specializing in AWS and Databricks. The ideal candidate will design, build, and maintain scalable data pipelines, ensuring efficient data ingestion, processing, and integration from multiple sources—including Google Analytics event data. This role requires deep expertise in AWS Glue, Lambda, Athena, Redshift, Databricks, PySpark, and SQL, alongside strong performance tuning, data security, and cost optimization skills.
Candidates with prior experience working in the retail domain will be strongly preferred. Cost optimization skills are essential.
Data Engineering & Pipeline Development:
- Develop and manage ETL pipelines for structured, semi-structured, and unstructured data using AWS Glue, PySpark, and SQL.
- Handle real-time event stream data ingestion and processing from multiple source systems.
- Ensure efficient data integration into Databricks for advanced processing and analytics.
Cloud & Infrastructure Management
- Build and optimize backend systems leveraging AWS services (Glue, Athena, Lambda, SNS, S3).
- Implement, configure, and manage Databricks environments, including clusters, notebooks, and libraries for performance optimization.
- Ensure optimal resource utilization for AWS and Databricks clusters to improve efficiency and reduce costs.
- Integrate Databricks with various cloud services while following governance and security best practices.
Testing & CI/CD Best Practices
- Write unit test cases and integration tests to ensure data pipeline reliability.
- Establish best practices for Databricks CI/CD and implement automation for deployment.
Optimization & Security
- Apply performance tuning techniques to optimize queries, storage, and processing times.
- Ensure compliance with security, governance, and industry best practices across AWS and Databricks environments.
- Monitor system performance and proactively address issues to maintain high availability and reliability.
People Management
The incumbent holds no direct supervisory responsibilities but is expected to engage effectively within their function and collaboratively across cross-functional teams. This role contributes meaningfully, whether through operational contributions and/or by offering specialized expertise, guidance, and support, to ensure alignment with either functional and/or strategic organizational goals and objectives.
Technical Competencies
Knowledge of Glue, PySpark, SQL, Athena, Lambda, SNS, S3
Knowledge of Databricks: Cluster setup, Notebooks, Libraries, CI/CD, Optimization
Data Processing: Event stream ingestion and batch processing
Testing: Writing unit test cases and integration tests
Security & Governance: AWS/Databricks governance standards and best practices
Performance Optimization: Query tuning, cluster performance improvements, cost reduction
Strong problem-solving and analytical skills
Ability to work in a fast-paced, cloud-based data workplace
Excellent collaboration and communication skills
Strong attention to detail and commitment to best practices
Experience
Minimum 6 experience in Data engineering (Core development/design), in which 3+ years on AWS with strong hands on (AWS glue, pyspark, SQL, Athena, lambda, SNS, S3) and 2+ year on Databricks.
📌 Data Engineer (Gurugram)
🏢 GMG
📍 Gurugram