- Design, develop, test, deploy and maintain large-scale data pipelines using GCP tools such as BigQuery, Dataflow, Pub/Sub.
- Collaborate with cross-functional teams to gather requirements and design solutions for complex data processing needs.
- Develop scalable machine learning models using Python libraries like scikit-learn or TensorFlow.
- Troubleshoot issues related to data quality, performance, and security in the pipeline.
Job Requirements :
- 7-15 years of experience in software development with a focus on big data technologies like Spark (Scala) or Kafka.
- Solid proficiency in programming languages such as Java or Node.js.
- Experience with cloud-based platforms like Google Cloud Platform (GCP).
- Knowledge of machine learning concepts and ability to apply them in real-world scenarios.