10 Aug
|
Ekfrazo Technologies Private
|
India
10 Aug
Ekfrazo Technologies Private
India
About the Role :
We are looking for an experienced Principal Data Engineer to design, build, and optimize scalable big data platforms and real-time data processing pipelines. The ideal candidate should have strong expertise in Data Modelling, Advanced SQL, Spark Streaming, Python, and Big Data ecosystems, with hands-on experience handling terabyte-scale datasets and building high-performance batch and streaming solutions.
Key Responsibilities :
- Design, develop, and optimize scalable data pipelines for batch and real-time data processing.
- Build and maintain robust Big Data platforms capable of handling terabyte-scale datasets.
- Develop Spark applications and optimize Spark SQL jobs for high-performance data processing.
- Design and implement semantic data models to support analytics and business intelligence.
- Process streaming data using Apache Spark Streaming and Kafka.
- Write and optimize complex SQL and Hive queries involving joins, UDFs, views, partitions, and large datasets.
- Build efficient data ingestion frameworks for structured, semi-structured, and unstructured data.
- Configure, schedule, and monitor workflows using Airflow and/or Oozie.
- Work with multiple file formats including ORC, AVRO, and Parquet.
- Develop cloud-based data solutions using AWS, Azure, or GCP.
- Collaborate with cross-functional teams to design scalable data architectures and analytics solutions.
- Troubleshoot performance bottlenecks and optimize distributed data processing workloads.
- Participate in code reviews and ensure adherence to engineering best practices.
Required Skills &
Experience :
- 7 years of experience in Data Engineering or Big Data Engineering.
- Strong expertise in :
- Data Modelling
- Advanced SQL
3.
Semantic
Modelling
4.
Handling
Terabyte-scale datasets
- Hands-on experience with :
1.
Apache
Spark
- Spark SQL
3.
Spark
Streaming (Mandatory)
- Python (Preferred) or Scala
5.
Apache
Kafka or other messaging platforms
- Hive
- Airflow and/or Oozie
- Strong knowledge of :
- Batch and Real-time Streaming Data Processing
2.
Big Data
Ecosystems
3.
Data Ingestion
Frameworks
4.
Distributed
Computing
- Experience working with :
- ORC
- AVRO
- Parquet
4.
Unstructured
Data
- Experience with cloud platforms such as AWS, Azure, or GCP.
- Experience with at least one contemporary data warehouse :
- Snowflake
- AWS Redshift
3.
Google
BigQuery
- Experience with NoSQL storage solutions such as Amazon S3 or similar object storage.
- Excellent analytical, debugging, and performance tuning skills.
Good to Have :
- AWS services such as EMR, S3, Redshift, ECS/EKS.
- GCP services such as Dataproc and Google Cloud Storage.
- Apache Iceberg.
- Hadoop MapReduce.
- Apache Flink.
- Kubernetes.
- ELK Stack (especially Elasticsearch).
- Experience working with large-scale Big Data clusters containing millions of records.
Interview Process :
- Round 1 : Technical Interview GlobalLogic Engineering Team
- Round 2 : Client Technical Interview
- Round 3 : Client Technical/Managerial Interview (Mandatory)
Why Join Us :
- Work on enterprise-scale Big Data and real-time streaming platforms.
- Build high-performance data solutions for global clients.
- Gain exposure to cloud-native data engineering and modern analytics technologies.
- Collaborate with experienced engineering teams on cutting-edge data transformation initiatives.
📌 Ekfrazo Technologies - Principal Data Engineer - Big Data & Spark Streaming (India)
🏢 Ekfrazo Technologies Private
📍 India