6 – 10 yearsOnsite — Bengaluru, HyderabadFull-time
Design, develop, and maintain scalable ETL/ELT pipelines using Python and PySpark, handling large volumes of structured and unstructured data across cloud-native data services.
PythonPySparkSQLETL/ELTAWSAzureGCP
Roles & responsibilities
- Design, develop, and maintain scalable ETL/ELT pipelines using Python and PySpark.
- Build and optimize data processing workflows handling large volumes of structured and unstructured data.
- Develop complex SQL queries, stored procedures, views, and data transformation logic.
- Implement data ingestion, cleansing, validation, and enrichment processes.
- Work with cloud-native data services on AWS, Azure, or GCP.
- Optimize Spark jobs for performance, scalability, and cost efficiency.
- Collaborate with business stakeholders, analysts, and architects to understand data requirements.
- Implement data quality checks, monitoring, and alerting mechanisms.
- Support CI/CD deployments and version control using Git-based repositories.
- Participate in Agile/Scrum ceremonies and contribute to solution design discussions.
- Troubleshoot production issues and provide performance tuning recommendations.
Prefer to send a file? Email your resume to
[email protected] with the role in the subject line. Both routes reach the same team.
📌 Senior Data Engineer Python, PySpark & GCP (Bengaluru)
🏢 Zetamicron
📍 Bengaluru