We are looking for a highly skilled Data Engineer to design, build, and optimize scalable data platforms that power analytics, machine learning, and business-critical applications. The ideal candidate will have solid expertise in building large-scale ETL/ELT pipelines, distributed data processing, cloud-native technologies, and modern data engineering practices.
You will work closely with Data Scientists, Analytics teams, and Platform Engineers to develop robust data solutions, enable ML workflows, and drive data-driven decision-making across the organization.
Key Responsibilities
- Design and develop scalable ETL/ELT pipelines to ingest and process high-volume telemetry, log, and business data.
- Build and maintain data lakes, data warehouses, feature stores, and ML training pipelines.
- Develop and optimize data models using Star Schema, Snowflake Schema, and Slowly Changing Dimensions (SCD Type 1, 2, and 3).
- Implement distributed data processing solutions using Apache Spark (PySpark/Scala).
- Develop, schedule, and monitor workflows using Apache Airflow.
- Build modular and scalable data transformation frameworks using dbt.
- Develop and maintain real-time streaming solutions using Apache Kafka or similar technologies.
- Deploy and manage containerized applications using Docker and Kubernetes.
- Utilize Infrastructure as Code (Terraform) for provisioning and managing cloud resources.
- Optimize query performance, data processing workloads, and cloud infrastructure costs.
- Ensure data quality, governance, security, and compliance across production environments.
- Collaborate with Data Scientists to support feature engineering and machine learning initiatives.
📌 Google Cloud Engineer (Bengaluru)
🏢 AMS
📍 Bengaluru