- Bachelor s or Master s degree in Computer Science, Engineering, or related field.
- Minimum 8+ years of work experience
- Proven experience in building and managing large-scale distributed data processing systems.
- Solid experience with big data frameworks like Hadoop, Spark, Hive, Kafka, etc.
- Proficiency in data processing languages such as Python, Java, or Scala(Any one)
- Hands-on experience with cloud-based big data solutions (AWS, GCP, or Azure).
- Experience with data orchestration tools like Apache Airflow.
- Strong problem-solving skills and the ability to work in a fast-paced environment.
Key Responsibilities:
- Design and develop scalable and high-performance data pipelines for collecting, processing, and storing large datasets.
- Work with Hadoop, Spark, Kafka, and other big data technologies to manage and process structured and unstructured data.
- Optimize and troubleshoot data pipelines for speed and efficiency.
- Collaborate with data scientists and analysts to integrate data-driven insights into business decision-making processes.
- Implement data security measures and maintain data integrity across the systems.
- Analyze and improve performance and scalability of data pipelines and storage solutions.
- Monitor and maintain data workflows, ensuring consistency and reliability.
- Work with cross-functional teams, including product, engineering, and operations, to understand and implement data requirements.
- Continuously evolve the data architecture to meet the future needs of the company.
- Knowledge of containerization technologies like Docker and Kubernetes is a plus.