We are seeking a highly skilled Senior Data Engineer with solid expertise in PySpark, Databricks, and SQL. The ideal candidate should have hands-on experience building enterprise-scale batch and real-time data pipelines, optimizing high-volume workloads, and designing modern data architectures.
Key Responsibilities
- Design and develop scalable batch and streaming data pipelines using PySpark and Databricks
- Build and optimize enterprise data models and transformation pipelines using Snowflake and DBT
- Develop advanced SQL scripts, stored procedures, and performance-optimized queries
- Design and implement real-time streaming architectures
- Build integrations using APIs, event-driven architectures, and cloud-native services
- Optimize workloads for performance, scalability, and cost efficiency
- Design enterprise data platforms using Lakehouse, Data Mesh, Data Vault, and Fabric principles
- Work with cross-functional teams to build scalable analytics and data engineering solutions
- Implement CI/CD and DevOps practices for data engineering pipelines
- Provide technical leadership and mentor engineering teams
Required Skills
- Strong hands-on experience with Python / PySpark
- Expertise in Databricks
- Advanced SQL expertise
- Strong experience with Stored Procedures and complex SQL optimization
- Hands-on experience with streaming frameworks and real-time processing
- Experience with REST APIs and enterprise integrations
- Strong understanding of distributed data processing
- Working knowledge of AWS and/or Azure
- Excellent analytical and debugging skills
Strong Knowledge Of
- Data Warehousing
- Data Vault Modeling
- Data Mesh
- Data Lake & Lakehouse Architecture
- Enterprise Data Architecture Patterns