We are seeking an experienced and motivated Java Spark Developer to design, develop, and maintain high-performance, large-scale data processing pipelines and distributed applications. In this role, you will leverage Core Java and Apache Spark to build resilient batch and real-time streaming data architectures, optimize distributed data workloads, and collaborate with cross-functional teams including Data Scientists, Cloud Engineers, and Solution Architects.
1. Data Pipeline & Application Development
• Design, implement, and maintain robust, scalable data ingestion and ETL/ELT pipelines using Apache Spark (Core, SQL, Streaming) written in Java (or Scala interoperability).
• Develop performant, low-latency microservices and distributed processing modules integrated with messaging platforms (e.g., Apache Kafka).
• Build and maintain interfaces to relational databases, distributed data lakes, and NoSQL stores (e.g., Hive, Cassandra, HBase, MongoDB, Delta Lake, Snowflake).
2. Performance Tuning & Optimization
• Profile, debug, and optimize Spark jobs by managing partitioning strategies,
caching, broadcast variables, memory allocation (driver/executor memory), and data serialization (Kryo).
• Analyze query execution plans, DAGs, and Spark UI metrics to eliminate data skew, reduce shuffle overhead, and minimize bottleneck latencies.
• Monitor resource utilization on cluster managers such as Kubernetes, Apache YARN, or cloud-native orchestration engines.
3. Architecture & Data Modeling
• Design structured, semi-structured, and unstructured data storage schemas using columnar file formats (e.g., Parquet, ORC, Avro).
• Implement robust data validation, cleansing, data governance, and error-handling mechanisms across the ingestion lifecycle.
• Ensure data privacy and enterprise compliance by applying encryption at rest/transit and access-control policies.
4. Collaboration, CI/CD & Best Practices
• Participate in Agile/Scrum ceremonies, sprint planning, and code reviews to
📌 Java Spark developer (Chennai)
🏢 Citi
📍 Chennai