- Must have 5+ years in Pyspark.
- Strong programming experience; Python, Pyspark, Hadoop, Hive is preferred.
- Experience in designing and implementing CI/CD, Build Management, and Development strategy.
- Experience with SQL and SQL Analytical functions, and participating in key business, architectural, and technical decisions.
- Scope to get trained on AWS cloud technology.
- Proficient in leveraging Spark for distributed data processing and transformation.
- Skilled in optimizing data pipelines for efficiency and scalability.
- Experience with real-time data processing and integration.
- Familiarity with Apache Hadoop ecosystem components.
- Robust problem-solving abilities in handling large-scale datasets.
- Ability to collaborate with cross-functional teams and communicate effectively with stakeholders.
Primary Skills:
- Pyspark
- SQL
- Hadoop
- Hive
Secondary Skill:
- Experience on AWS/Azure/GCP would be an added advantage.