10 Oct
|
Tata Consultancy Services
|
Mumbai
10 Oct
Tata Consultancy Services
Mumbai
Role & responsibilities
Mandatory Technical Skills
- Strong experience with Cloudera Data Platform, Cloudera Runtime, HDFS, YARN, Hive, Impala
- Hands-on development in Spark Structured Streaming using Scala or PySpark, including DataFrames, Spark SQL, Kafka integration, checkpointing, joins, aggregations, and performance tuning.
- Hands-on development with Apache Flink using Java or Scala and Flink SQL, including DataStream or Table APIs, event-time processing, state management, checkpoints, savepoints, connectors, and job deployment.
- Strong experience designing Kudu schemas, primary keys, hash and range partitions, tablet distribution, replication, upsert patterns, and query access through Impala.
- Solid knowledge of Apache Kafka architecture, producers, consumers, consumer groups, partitions, offsets, retention, serialization,
schema evolution, and secure connectivity.
- Proficiency in Java, Scala, or Python, with strong SQL and Linux scripting skills.
- Understanding of distributed systems, fault tolerance, high availability, data consistency, and streaming delivery guarantees.
- Experience troubleshooting production-grade streaming applications using logs, metrics, dashboards, and platform administration tools.
Preferred Skills
- Experience with Cloudera Streaming Analytics, SQL Stream Builder, Cloudera Data Engineering.
Exposure to Git, Maven, Jenkins or similar CI/CD tools, and automated deployment
📌 Cloud Data Platform (Mumbai)
🏢 Tata Consultancy Services
📍 Mumbai