11 Sep
|
Tata Electronics (TEPL)
|
Rajkot
11 Sep
Tata Electronics (TEPL)
Rajkot
Job Responsibilities -
- Architect and implement scalable offline data pipelines for manufacturing systems including AMHS, MES, SCADA, PLCs, vision systems, and sensor data.
- Design and optimize ETL/ELT workflows using Python, Spark, SQL, and orchestration tools (e.g., Airflow) to transform raw data into actionable insights.
- Lead database design and performance tuning across SQL and NoSQL systems, optimizing schema design, queries, and indexing strategies for manufacturing data.
- Enforce robust data governance by implementing data quality checks, lineage tracking, access controls, security measures, and retention policies.
- Optimize storage and processing efficiency through strategic use of formats (Parquet, ORC), compression, partitioning, and indexing for high-performance analytics.
- Implement streaming data solutions (using Kafka/RabbitMQ) to handle real-time data flows and ensure synchronization across control systems.
- Building dashboards using analytics tools like Grafana.
- Good Understanding of Hadoop ecosystem.
- Develop standardized data models and APIs to ensure consistency across manufacturing systems and enable data consumption by downstream applications.
- Collaborate cross-functionally with Platform Engineers, Data Scientists, Automation teams, IT Operations, Manufacturing, and Quality departments.
- Mentor junior engineers while establishing best practices, documentation standards, and fostering a data-driven culture throughout the organization.
Essential Attributes -
- Expertise in Python programming for building robust ETL/ELT pipelines and automating data workflows.
- Proficiency with Hadoops ecosystem.
- Hands-on experience with Apache Spark (PySpark) for distributed data processing and large-scale transformations.
- Strong proficiency in SQL for data extraction, transformation, and performance tuning across structured datasets.
- Proficient in using Apache Airflow to orchestrate and monitor complex data workflows reliably.
- Skilled in real-time data streaming using Kafka or RabbitMQ to handle data from manufacturing control systems.
- Experience with both SQL and NoSQL databases, including PostgreSQL, Timescale DB, and MongoDB, for managing diverse data types.
- In-depth knowledge of data lake architectures and effective file formats like Parquet and ORC for high-performance analytics.
- Proficient in containerization and CI/CD practices using Docker and Jenkins or GitHub Actions for production-grade deployments.
- Strong understanding of data governance principles, including data quality, lineage tracking, and access control.
- Ability to design and expose RESTful APIs using FastAPI or Flask to enable standardized and scalable data consumption.
Qualifications -
- BE/ME Degree in Computer science, Electronics, Electrical
Desired Experience Level -
- Masters+ 2 Years of relevant experience.
- Bachelors+4 Years of relevant experience.
- Experience with semiconductor industry is a plus.
Skills: Kafka, Apache Spark, Rabbitmq, Pyspark, Sql, Python
Experience: 8.00-13.00 Years
📌 Data Engineer (Rajkot)
🏢 Tata Electronics (TEPL)
📍 Rajkot