11 Sep
|
Halcer
|
Bengaluru
About Halcer Halcer is a leading IT services, consulting, and staffing organization committed to delivering innovative technology solutions and top-tier technical talent to global enterprises.
Job Overview We are seeking an experienced Data Engineer (6+ Years of Experience) to build and maintain scalable, high-performance data pipelines and infrastructure for our next-generation data platform. The platform ingests and processes real-time and historical data from diverse industrial sources such as airport systems, sensors, cameras, and APIs. You will work closely with AI/ML engineers, data scientists, and DevOps teams to enable reliable analytics, forecasting, and anomaly detection use cases.
Key Responsibilities
- Design and implement real-time (Kafka, Spark/Flink) and batch (Airflow, Spark) pipelines for high-throughput data ingestion, processing, and transformation.
- Develop data models and manage data lakes and warehouses (Delta Lake, Iceberg, etc.) to support both analytical and ML workloads.
- Integrate data from diverse sources: IoT sensors, databases (SQL/NoSQL), REST APIs, and flat files.
- Ensure pipeline scalability, observability, and data quality through monitoring, alerting, validation, and lineage tracking.
- Collaborate with AI/ML teams to provision clean and ML-ready datasets for training and inference.
- Deploy, optimize, and manage pipelines and data infrastructure across on-premise and hybrid environments.
- Participate in architectural decisions to ensure resilient, cost-effective, and secure data flows.
- Contribute to infrastructure-as-code and automation for data deployment using Terraform, Ansible, or similar tools.
Qualifications & Experience Requirements
- Total Experience: 6+ years of total experience in core data engineering roles.
- Streaming & Real-Time Experience: 2+ years of direct experience building and maintaining real-time or streaming pipelines.
- Education: Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.
Required Skills
- Strong programming proficiency in Python or Java, alongside expert-level SQL.
- Hands-on experience with Apache Kafka, Apache Spark, or Apache Flink for real-time and batch processing.
- Proficiency with workflow orchestration tools like Airflow, dbt, or similar technologies.
- Deep familiarity with data modeling (OLAP/OLTP), schema evolution, and file formats (Parquet, Avro, ORC).
- Hands-on experience with hybrid/on-premise and cloud platform deployments (AWS, GCP, or Azure).
- Proven track record working with data lakes and modern data warehouses (Snowflake, BigQuery, Redshift, or Delta Lake).
- Solid working knowledge of DevOps practices, Docker, Kubernetes, and IaC tools like Terraform or Ansible.
- Knowledge of data observability, data cataloging, and quality frameworks (e.g., Great Expectations, OpenMetadata).
Good-to-Have Skills
- Experience with time-series databases (e.g., InfluxDB, TimescaleDB) and processing high-frequency sensor data.
- Prior domain experience in aviation, manufacturing, or logistics.
📌 Data Engineer (Bengaluru)
🏢 Halcer
📍 Bengaluru