05 Oct
|
GK HR Consulting India Private
|
India
05 Oct
GK HR Consulting India Private
India
We are seeking an experienced Big Data Engineer with strong expertise in PySpark, Scala, Apache Spark, and distributed data processing. The ideal candidate should have hands-on experience building scalable data pipelines, processing large datasets, and working with cloud-based data platforms. Exposure to Generative AI, LLMs, or AI/ML technologies is desirable but not mandatory. Internal AI engineering role definitions also highlight cloud, MLOps, and GenAI familiarity as valuable complementary skills.
Mandatory Skills
- PySpark
- Scala
- Apache Spark
- Hadoop Ecosystem
- Hive
- SQL
- Data Warehousing Concepts
- ETL/ELT Development
- Python
- Git
- Linux/Unix
- Cloud Platforms (AWS/Azure/GCP)
Valuable to Have
- Databricks
- Kafka
- Airflow
- Delta Lake
- Snowflake
- Generative AI / LLMs
- Lang Chain
- Vector Databases
- AI/ML Fundamentals
- MLOps
- Azure OpenAI / OpenAI APIs
Key Responsibilities
- Design and develop scalable big data solutions using PySpark and Scala.
- Build and optimize ETL pipelines for processing large volumes of structured and unstructured data.
- Develop data ingestion, transformation, and orchestration frameworks.
- Work with Spark, Hive, Hadoop, and cloud-based data platforms.
- Optimize Spark jobs for performance, scalability, and reliability.
- Collaborate with data scientists, analysts, and business teams to deliver data solutions.
- Implement best practices for data quality, governance, and security.
- Support deployment automation and CI/CD for data engineering workloads.
- Contribute to AI-powered data initiatives where applicable.
Required Qualifications
- Bachelor's or Master's degree in Computer Science, Engineering, or related field.
- 5+ years of hands-on experience in Big Data Engineering.
- Strong experience with PySpark and Scala development.
- Extensive experience with Spark SQL, Hive, and distributed data processing.
- Experience working with cloud data platforms and modern data architectures.
- Strong problem-solving and debugging skills.
Preferred Experience
- Databricks implementation and optimization.
- Real-time data processing using Kafka/Spark Streaming.
- Exposure to Generative AI, LLMs, or AI-assisted analytics solutions.
- Experience integrating AI/ML workloads with big data platforms.
📌 Pyspark + AI (India)
🏢 GK HR Consulting India Private
📍 India