Job Summary
We are looking for a skilled
Big Data Engineer
to design, develop, and maintain scalable data processing solutions. The ideal candidate should have hands-on experience with big data technologies, data pipelines, distributed computing, and cloud/data platforms.
Key Responsibilities
- Design and develop scalable big data pipelines and ETL/ELT workflows.
- Process and analyze large volumes of structured and unstructured data.
- Develop data solutions using Hadoop, Spark, Kafka, or similar technologies.
- Build and optimize batch and real-time data processing pipelines.
- Work with data engineers, analysts, developers, and business teams.
- Optimize data processing jobs for performance, reliability, and scalability.
- Implement data quality, validation, and monitoring processes.
- Troubleshoot pipeline failures and production data issues.
- Work with cloud-based data platforms and storage solutions.
- Maintain technical documentation and follow data engineering best practices.
Required Skills
- Robust programming experience in Python, Java, or Scala.
- Hands-on experience with Apache Spark / PySpark.
- Knowledge of Hadoop, HDFS, Hive, or similar big data technologies.
- Experience building ETL/ELT pipelines.
- Knowledge of SQL and relational databases.
- Experience with Apache Kafka or other streaming technologies.
- Understanding of distributed computing and data processing concepts.
- Familiarity with AWS, Azure, or GCP is preferred.
- Good understanding of data warehousing and data modeling.