01 Oct
|
Dhruv Technology Solutions
|
Bengaluru
01 Oct
Dhruv Technology Solutions
Bengaluru
Owns the pipelines, data models, and analytical output built on top of the platform
Core Responsibility Summary
Designs and builds real-time CDC pipelines, Flink stream processing jobs, and ClickHouse data models that deliver factory analytics, dashboard data, and agentic integration endpoints. Consumes the platform provided by the Platform Engineer.
Core Technical Skills
Streaming & CDC
- Apache Kafka (producers, consumers, consumer groups, offsets — as a pipeline developer)
- Kafka Connect source/sink connectors, SMTs, connector configuration
- Debezium CDC — relational database source connectors, event schemas, slot management, schema evolution handling
Stream Processing
- Apache Flink (DataStream API, Table API/SQL, stateful processing, windowing, joins)
- PyFlink or Flink SQL for job development
- Real-time aggregation, enrichment, deduplication, and late-data handling patterns
- Complex Event Processing (CEP) for factory event pattern detection
Databases & Data Modeling
- ClickHouse — table engine selection, materialized views, TTL policies, query optimization
- Relational database proficiency (SQL, joins, indexes) for understanding CDC source schemas
- Time-series and analytical data modeling (wide tables, pre-aggregation, partitioning strategies)
- Schema Registry and schema evolution (Avro, JSON Schema)
Python Development
- Python pipeline development, data transformation logic, and tooling
- pandas, PyArrow, SQLAlchemy, and Kafka/Flink Python clients
- FastAPI or Flask for data service APIs consumed by dashboards and agentic systems
- Unit testing, mocking, and integration testing for pipeline code
AI-Assisted Development
- Proficiency with AI coding assistants (GitHub Copilot, Cursor, Claude, or equivalent)
- Using AI to scaffold Flink jobs, generate ClickHouse DDL, write Kafka Connect configurations
- Prompt engineering to produce correct, production-ready pipeline code
- Critical evaluation of AI-generated code for correctness, performance, and security
- Using AI for test generation, documentation, and debugging pipeline failures
Domain & Integration Skills
- Factory data source integration: MES, SCADA, factory tool messaging
- Familiarity with SECS/GEM, OPC-UA, or MQTT (preferred, not required)
- Dashboard integration: Apache Superset, Grafana, Tableau, or Power BI
- Agentic/LLM integration patterns — exposing real-time factory data as tool APIs or vector store input
📌 Data engineer (Bengaluru)
🏢 Dhruv Technology Solutions
📍 Bengaluru