Company: Celebal Technologies
Location: Hyderabad/Bangalore
Experience: 3-10 Years
Role Overview
We are looking for a skilled Data Engineer with strong experience in building scalable data pipelines using modern big data technologies. The ideal candidate should have hands-on experience with streaming + batch processing, along with deep knowledge of Databricks and Delta Lake.
Key Responsibilities
- Design and build end-to-end ETL/ELT pipelines
- Ingest data from multiple sources like Kafka, databases, APIs, etc.
- Develop and optimize data pipelines using PySpark / Spark
- Work on both batch and real-time (streaming) data processing
- Implement Delta Lake architecture for reliable and scalable data storage
- Handle large-scale data (100GB1TB+) efficiently
- Perform data transformation, cleansing, and validation
- Optimize performance using partitioning, caching, and file formats
- Work with orchestration tools like Airflow for scheduling workflows
- Collaborate with Data Analysts and Business teams for requirements
Required Skills (Must Have)
- Strong experience in PySpark / Apache Spark
- Hands-on with Databricks
- Good understanding of Delta Lake
- Experience with Apache Kafka
- Strong SQL skills (joins, CTEs, window functions, recursion basics)
- Experience in designing scalable data pipelines
- Understanding of data warehousing concepts
Valuable to Have
- Experience with Microsoft Azure
- Knowledge of Apache Airflow
- Familiarity with Databricks Autoloader
- Experience with CI/CD tools like Jenkins
- Exposure to data modeling concepts
What Were Looking For
- Someone who has worked on real production pipelines
- Strong problem-solving mindset
- Ability to handle large-scale data efficiently
- Clear understanding of streaming + batch architecture
Interested candidate apply at
[email protected]
📌 Data Engineer (Hyderabad)
🏢 Celebal Technologies
📍 Hyderabad