01 Aug
|
Nam Info
|
Karnataka
01 Aug
Nam Info
Karnataka
Key Responsibilities:
- Design & implement distributed data pipelines using Apache Spark for batch and streaming workloads.
- Develop real-time streaming systems with Apache Kafka (Kafka Streams, Flink, or similar).
- Build ETL/ELT workflows to ingest data from multiple sources into data lakes/warehouses.
- Optimize Spark jobs (partitioning, caching, broadcast joins, cluster tuning).
- Manage data lakes (S3, GCS, Azure Data Lake) with formats like Parquet, Delta Lake, or Iceberg.
- Ensure data quality, governance, and security compliance across pipelines.
- Collaborate with cross-functional teams to translate business requirements into technical solutions.
- Mentor junior engineers and lead projects when in senior roles.
Required Skills & Qualifications
- Core technologies:
Apache Kafka, Apache Spark, Hadoop ecosystem, Flink (optional).
- Programming languages: Proficiency in Java/Scala, Python, or SQL.
- Cloud platforms: AWS, GCP, or Azure with certifications (Databricks, GCP Data Engineer, AWS Data Engineer).
- Experience:7-17 Years
- Education: Bachelors/Masters in Computer Science, Data Science, or related quantitative fields.
- Preferred extras: Knowledge of data lakehouse formats (Iceberg, Delta), orchestration tools (Airflow, DolphinScheduler), and real-time analytics stores (Pinot, Clickhouse).
📌 Kafka,Spark,Bigdata (Karnataka)
🏢 Nam Info
📍 Karnataka