Key Responsibilities:
- Build & optimize ETL pipelines using Apache Spark for batch and real-time data processing.
- Develop scalable data ingestion solutions integrating MSSQL, PostgreSQL, and NoSQL sources into Apache Iceberg tables.
- Optimize Spark SQL queries and Dremio performance tuning for high-speed analytics.
- Implement streaming data ingestion using Apache Kafka or Spark Structured Streaming.
- Manage data catalog and metadata storage with Hive Metastore and Iceberg.
- Automate data workflows using Apache Airflow.
- Deploy and monitor Spark jobs in containerized environments (Docker, Kubernetes).
- Implement data security & governance policies using Apache Ranger to enforce RBAC, ABAC, and encryption rules.
- Collaborate with Data Analysts, Scientists, and Business Teams to support data-driven decision-making.
Required Skills & Qualifications:
- 3+ years of hands-on experience in Data Engineering / Big Data Technologies.
- Strong expertise in Apache Spark , ETL Design, and SQL Optimization.
- Data Storage & Processing:
Proficiency in Apache Iceberg, Hive Metastore, MinIO, S3.
- Streaming Pipelines: Experience with Apache Kafka, Spark Streaming, or Flink for real-time processing.
- SQL Query Optimization: Hands-on experience with Spark SQL, Dremio.
- Workflow Orchestration: Experience with Apache Airflow for job automation.
- Containerization & CI/CD: Working knowledge of Docker, Kubernetes and GitHub.
- Security & Data Governance: Experience with Apache Ranger for access control, encryption, and compliance.
- Robust understanding of data partitioning, caching, and indexing strategies for big data workloads
Preferred Skills: (Good-to-have)
- Experience implementing Lakehouse architectures with Apache Iceberg, Delta Lake, or Hudi.
- Familiarity with REST APIs, and BI integrations (Metabase, Tableau, Power BI).
- Experience using Python-based data frameworks (pandas, SQLAlchemy, DuckDB,MSSQL, MongoDb).
📌 Data Engineer (Bengaluru)
🏢 CSC
📍 Bengaluru