- Join a team where data engineering meets real world impact
- In this role you ll help design and deliver scalable data solutions using Databricks and PySpark enabling teams to turn raw data into trusted analytics ready assets
- You ll collaborate closely with consultants data engineers and stakeholders to understand business needs build reliable pipelines and support high quality releases
- This is a great opportunity for someone with 2 3 years of experience who enjoys solving data challenges improving performance and learning contemporary lakehouse practices
- If you re motivated by clean engineering continuous improvement and working in a collaborative environment where your contributions are visible and valued this role offers the right mix of ownership guidance and growth
Key Responsibilities:
- Key Responsibilities
- Develop and maintain data pipelines and transformations using Databricks and PySpark
- Implement scalable ETL ELT workflows to ingest cleanse and curate data for downstream analytics and reporting
- Optimize Spark jobs for performance and cost by tuning partitions caching joins and cluster configurations
- Build reusable notebooks and modular code to support consistent development and easier maintenance
- Perform data validation reconciliation and quality checks to ensure accuracy and reliability of datasets
- Collaborate with cross functional teams to gather requirements clarify data definitions and deliver aligned solutions
- Support deployments and production operations by troubleshooting failures analyzing logs and resolving incidents
- Contribute to documentation coding standards and best practices for Databricks based development
- Minimum Qualifications
- Bachelor s or Master s degree in BTECH MTECH MCA or MSC or equivalent
- 2 3 years of hands on experience working with Databricks in data engineering or analytics engineering projects
- Strong experience in PySpark for building transformations and distributed data processing
- Solid understanding of data pipeline concepts data modeling basics and structured semi structured data handling
- Ability to debug and troubleshoot Spark jobs and collaborate effectively within delivery teams
Technical Requirements:
- ETL PYSPARK DATABRICKS Delta Lake Spark SQL Data Modeling Workflow Orchestration Performance Tuning
Additional Responsibilities:
- Experience with Spark optimization techniques and practical performance tuning in Databricks environments
- Familiarity with Delta Lake concepts such as ACID tables schema evolution and incremental processing patterns
- Exposure to orchestrating workflows and managing dependencies for end to end pipeline execution
- Experience working in agile delivery models with strong ownership of tasks timelines and quality outcomes
- Strong communication skills to translate requirements into implementable data solutions and clearly document outcomes
Preferred Skills:
Technology->Big Data - Data Processing->PySpark,Technology->Data Engineering->Databricks
📌 Databricks (Bengaluru)
🏢 Infosys
📍 Bengaluru
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.