- Design, develop, and maintain scalable data pipelines on Google Cloud Platform
- Build and optimize batch and streaming data processing frameworks using Scala
- Develop ETL/ELT workflows using GCP services such as:
- BigQuery
- Dataflow (Apache Beam)
- Pub/Sub
- Cloud Storage
- Cloud Composer (Airflow)
- Collaborate with data scientists, analysts, and business stakeholders to support analytics and ML use cases
- Ensure data quality, reliability, and performance across pipelines
- Implement best practices for cloud security, governance, and cost optimization
- Troubleshoot production issues and perform root cause analysis
- Participate in architecture discussions, code reviews, and DevOps processes
Required Skills & Qualifications
- 4+ years of experience in Data Engineering
- Strong hands-on experience with Google Cloud Platform (GCP)
- Proficiency in Scala (mandatory)
- Solid experience with BigQuery and Dataflow / Apache Beam
- Strong understanding of distributed systems, data modeling, and performance tuning
- Experience building ETL/ELT pipelines for large-scale datasets
- Familiarity with SQL, NoSQL, and cloud-based data warehouses
- Knowledge of CI/CD pipelines, version control (Git), and Agile methodologies
- Strong problem-solving and communication skills
Valuable to Have (Preferred Skills)
- Experience with Apache Spark / Dataproc
- Exposure to real-time/streaming pipelines using Pub/Sub or Kafka
- Experience with Python alongside Scala
- Understanding of ML data pipelines and feature engineering
- GCP certifications (Professional Data Engineer preferred)