Spark-Java, Databricks
Key Responsibilities:
• Design, develop, and maintain Spark-based data processing jobs using Java aligned to business and technical requirements.
• Build and enhance ETL workflows, ensuring accuracy, completeness, and consistency of processed datasets.
• Implement integrations and workflows on DBX, supporting scalable execution and operational reliability.
• Optimize Spark jobs for performance (partitioning, caching, shuffle tuning) and cost efficiency.
• Write clean, maintainable code with appropriate logging, error handling, and unit/integration tests.
• Troubleshoot production issues, perform root-cause analysis, and implement preventive fixes.
• Collaborate with cross-functional teams to refine requirements, plan deliveries, and ensure smooth releases.
• Document technical designs, data flows, and operational runbooks to support long-term maintainability. Minimum Qualifications:
• 3–5 years of hands-on experience in software/data engineering roles.
• Bachelor’s or Master’s degree: BTECH, MTECH, MCA, MSC.
• Solid programming experience in Java with solid understanding of OOP and coding best practices.
• Practical experience with Apache Spark for batch data processing.
• Experience building and supporting ETL pipelines and data transformations.
• Working knowledge of DBX for developing and running data workflows. Preferred Qualifications:
• Experience designing end-to-end data pipelines including ingestion, transformation, validation, and publishing layers.
• Strong understanding of distributed processing concepts and Spark internals for performance tuning and stability.
• Exposure to CI/CD practices for data/engineering workflows and disciplined release management.
• Experience with production monitoring, alerting, and operational support for data pipelines.
• Proven ability to collaborate with stakeholders, communicate trade-offs, and deliver within timelines.
📌 Spark-Java, DBX (India)
🏢 Infosys
📍 India