21 Aug
|
Jobgether
|
India
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Databricks Engineer - Senior/Lead based in India.
This is a senior-level data engineering opportunity focused on designing and delivering scalable data platforms and analytics solutions on Microsoft Azure. You will work extensively with Azure Databricks, Python, SQL, and Apache Spark to build high-performance data pipelines and transformation workflows. The role involves working with large and complex datasets across structured and semi-structured formats while applying modern data engineering practices. You will optimize distributed processing workloads for performance, scalability, reliability, and cost efficiency. You will also collaborate with data scientists, analysts, and business stakeholders to enable analytics and machine learning use cases. This is an opportunity to take ownership of critical data engineering initiatives within a technology-driven, AI-focused environment.
Accountabilities:
- Design, develop, and maintain scalable data pipelines and data engineering solutions using Azure Databricks.
- Build robust ETL/ELT workflows using Python, PySpark, Spark SQL, and Apache Spark.
- Develop complex SQL queries and transformations to process and prepare large datasets for analytics and machine learning.
- Optimize Spark jobs and distributed processing workloads for performance, scalability, reliability, and cloud cost efficiency.
- Work with structured and semi-structured data formats including Parquet, Delta, JSON, and CSV.
- Design, build, and maintain Delta Lake tables, leveraging capabilities such as ACID transactions, time travel, and schema evolution.
- Integrate Databricks workloads with Azure Data Lake Storage Gen2 and other Azure data services.
- Apply distributed computing principles to efficiently process large-scale datasets.
- Implement data quality, validation, monitoring, and error-handling processes across data pipelines.
- Collaborate with data scientists, analysts, and business stakeholders to understand requirements and deliver reliable data solutions for analytics and ML use cases.
- Apply Azure security, access control, governance, and RBAC best practices across data engineering environments.
- Use Git-based version control and follow collaborative software development practices.
- Contribute to CI/CD pipelines and automated deployment processes using tools such as Azure DevOps or GitHub Actions.
- Support data modeling initiatives and help ensure data structures are optimized for downstream analytics and reporting.
- Contribute to streaming data solutions using technologies such as Spark Structured Streaming, Azure Event Hubs, or Kafka.
- Identify opportunities to improve pipeline architecture, engineering standards, performance, and operational efficiency.
- Provide technical guidance and contribute to engineering best practices appropriate for a Senior or Lead-level position.
Requirements:
- 6+ years of skilled experience in Data Engineering or a closely related field.
- Strong hands-on experience designing and developing solutions with Azure Databricks.
- Advanced proficiency in Python for data processing, transformation, and pipeline development.
- Strong SQL skills, including complex joins, window functions, data transformations, and query performance optimization.
- Hands-on expertise with Apache Spark and PySpark, including distributed data processing and Spark job optimization.
- Solid experience working with Delta Lake, including ACID transactions, time travel, and schema evolution.
- Strong knowledge of Azure Data Lake Storage Gen2 and cloud-based data architecture.
- Understanding of distributed computing concepts and the ability to design solutions for large-scale data workloads.
- Experience using Git for version control and collaborative development.
- Strong understanding of modern ETL/ELT patterns, data pipeline architecture, and data engineering best practices.
- Experience with Azure Data Factory is preferred.
- Exposure to CI/CD practices and tools such as Azure DevOps or GitHub Actions is an advantage.
- Basic understanding of data modeling and its application to analytical workloads.
- Familiarity with Azure security principles, RBAC, access control, and data governance.
- Exposure to real-time and streaming data technologies such as Spark Structured Streaming, Event Hubs, or Kafka is a plus.
- Strong analytical and problem-solving abilities, with a structured approach to debugging and performance optimization.
- Strong communication and collaboration skills, with the ability to work effectively with technical teams, data scientists, analysts, and business stakeholders.
- Ability to take ownership of complex engineering initiatives and provide technical direction at a senior or lead level.
Benefits:
- Opportunity to work on large-scale data engineering and analytics initiatives within an AI and technology-focused environment.
- Hands-on experience with Microsoft Azure, Azure Databricks, Apache Spark, Delta Lake, and modern cloud data platforms.
- Opportunity to design and optimize production-grade data pipelines handling complex and large-scale datasets.
- Exposure to advanced distributed data processing, cloud architecture, and modern ETL/ELT practices.
- Opportunity to contribute to analytics and machine learning use cases in collaboration with data scientists and analysts.
- Exposure to CI/CD, cloud security, data governance, and streaming data technologies.
- Senior/Lead-level scope with opportunities to influence engineering standards and technical architecture.
- Collaborative environment involving data engineering, analytics, data science, and business stakeholders.
- Opportunities for continued technical growth across cloud data engineering, AI, and emerging data technologies.
How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. #LI-CL1We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
📌 Databricks Engineer - Senior/ Lead (India)
🏢 Jobgether
📍 India