31 Jul
|
Nineleaps
|
Karnataka
31 Jul
Nineleaps
Karnataka
Location: Bangalore
Department: Data & Analytics / Data Engineering
Employment Type: Full-Time
Position Summary
We are looking for an experienced Data Engineer with 4 to 6 years of hands-on experience to design, build, and maintain scalable, robust batch and real-time data pipelines. In this role, you will leverage Databricks, PySpark, Python, SQL, Apache Airflow, and Apache Hive to process complex datasets and deliver high-quality data products across the enterprise.
Key Responsibilities
- Pipeline Development & Maintenance: Design, build, and deploy production-grade ETL/ELT data pipelines using PySpark and Python on Databricks.
- Workflow Orchestration: Author, schedule, and maintain complex Directed Acyclic Graphs (DAGs) using Apache Airflow to automate and monitor data integration processes.
- Data Lakehouse & Warehouse Optimization: Manage Delta Lake/Hive Metastore architectures, utilizing Medallion Architecture (Bronze/Silver/Gold layers) for optimal data organization and governance.
- Performance Tuning: Diagnose and resolve performance bottlenecks in Apache Spark jobs, tuning partition strategies, caching, memory allocation, and custom Spark SQL queries.
- Data Modeling & SQL: Write complex, highly optimized SQL queries for data transformations, analytical aggregations, and data validation.
- Data Quality & Monitoring: Implement automated data quality checks, alerting mechanisms,
and exception handling using tools like Excellent Expectations or DLT Expectations.
- Collaboration & CI/CD: Partner with Data Analysts, Data Scientists, and Business Stakeholders to gather requirements and enforce version control (Git) and CI/CD deployment practices.
Education & Experience
- 4-6 years of overall experience as a Data Engineer, Big Data Engineer, or Analytics Engineer.
- Bachelors degree in Computer Science, Information Technology, Software Engineering, or a related field.
Primary / Core Technical Skills
- Python & PySpark: Advanced proficiency in Python programming and extensive hands-on experience building distributed data processing applications using PySpark (DataFrames, RDDs, Spark SQL).
- Databricks Platform: Deep experience working with Databricks Workspaces, Delta Lake, Auto Loader, Unity Catalog, and Delta Live Tables (DLT).
- Advanced SQL: Mastery in writing, debugging, and tuning complex SQL queries, CTEs, window functions, and analytical functions.
- Orchestration (Apache Airflow): Solid experience constructing custom DAGs, operators, sensors, XComs, and managing dependencies in Apache Airflow.
- Hive Ecosystem: Strong understanding of Apache Hive architectures, HQL, metastore management, schema management, and partitioning strategies.
📌 Data Engineer (Karnataka)
🏢 Nineleaps
📍 Karnataka