Job Description
We're are Hiring @Algo8AI
n
Role:
n
Data Engineer — Data Platform & Pipelines
n
Lake, canonical model and data quality
n
Location: Noida | Hybrid (4 days/week in office)
n
Experience: 5–7 years
n
Employment: Permanent | Algo8 Payroll
n
Reports to: Tech Lead
n
Joining: Early joiners / short notice preferred
n
You will build and own the spine of the programme — the on-premises data platform, the canonical model that everything lands into, and the quality framework that determines whether the foundation can be trusted.
n
About Algo8
n
Algo8 is an industrial and manufacturing AI firm. We work inside plants and R&D; centres — with engineers, scientists and plant teams — building the data foundations and AI capabilities that enable industrial organisations to make better decisions. Our work is hands-on and close to the client.
n
The Engagement
n
You will join a multi-year programme with the R&D; centre of excellence of a large Indian manufacturer.
n
The centre holds four decades of engineering knowledge — design records, formulations, laboratory results, test data, reports and studies — spread across enterprise systems, departmental databases, local servers and document archives.
n
The programme recovers that knowledge into a single governed foundation, connects it to manufacturing context, and then validates where engineering intelligence can be built on top of it.
n
This is a data engineering programme first. The AI comes later, and only where the evidence supports it.
n
The solution is on-premises.
n
Because of the proprietary nature of new product development, this foundation is being built inside the client's own infrastructure.
n
If your experience is entirely on managed cloud services, this will not be the right fit.
n
What You Will Own
n
The platform
n
Stand up and operate the on-premises data lake and processing layer — storage, compute, orchestration and access control.
n
The canonical model
n
Work with engineers and business analysts to design a shared data model across engineering and manufacturing domains, then keep it coherent as new sources land.
n
Ingestion pipelines
n
Build the pipelines that bring recovered and synchronised data into the foundation, with lineage from source to consumption.
n
Data quality
n
Define and implement quality thresholds per source. Each estate has to pass an acceptance gate before it is declared complete — you own the evidence behind that.
n
Enabling everyone else
n
The integration and document engineers land data into your platform. Analytics and validation work builds on top of it. Making both easy is part of the job.
n
The Expected Stack
n
The architecture is being finalised with the client. We expect to work with an on-premises stack along these lines, or close equivalents:
n
n
- Distributed processing and query — Spark, Trino or equivalent
n
- On-premises object storage — MinIO, HDFS or equivalent
n
- Relational stores — PostgreSQL, Oracle, SQL Server
n
- Orchestration — Airflow or equivalent
n
- Streaming and change capture — Kafka, Debezium or equivalent
n
n
We are more interested in whether you have run this class of system on your own infrastructure than in which specific tools you have used.
n
What We Need From You
n
n
- Strong on-premises data platform experience — you have built and run distributed processing and storage on infrastructure you or your organisation managed, not only on managed cloud services.
n
- Data modelling depth — you can design a canonical model across messy sources and defend the decisions behind it.
n
- Production pipeline engineering — Spark or equivalent, Python, solid SQL, and orchestration with Airflow or similar.
n
- Data quality and governance in practice — validation frameworks, lineage and reconciliation. Not as theory, but as something you have actually implemented.
n
- Clear communication — you will explain to non-technical stakeholders why data is or is not ready,
and that conversation matters as much as the pipeline.
n
n
Useful, Not Essential
n
n
- Manufacturing or industrial data experience
n
- Kafka, Debezium or other change-data-capture tooling
n
- Metadata catalogue and governance tooling
n
- Experience defining acceptance criteria on a client programme
n
n
How You Will Work
n
n
- Based in Noida, hybrid — four days a week in office.
n
- Reporting to the Tech Lead for the programme, alongside a large on-site business analyst team who own requirements gathering and client expectations.
n
- Occasional travel to client plant and R&D; locations as the work requires.
n
- You will be working with an existing enterprise estate that is not fully documented.
n
n
Discovery is part of the job — expect to map what is actually there before you design against it.
n
Who This Role Is Not For
n
We would rather be direct than waste your time. This role is probably not a fit if:
n
n
- Your data engineering experience is entirely on managed cloud services and you have never sized, deployed or operated infrastructure yourself.
n
- You are a BI or reporting developer — dashboards, SQL views and reports — rather than someone who has built and owned pipelines.
n
- Your work has been drag-and-drop ETL tooling without engineering underneath it.
n
- You are an application or backend developer whose data work has been incidental.
n
- You have never worked with enterprise systems such as ERP, PLM or MES.
n
n
Working at Algo8
n
n
- You will build capability that outlasts this engagement. What you learn here is directly reusable across our industrial client base.
n
- Small senior team, high ownership, and direct access to client engineering leadership.
n
n
Applying
n
We need this team in place early.
n
Please tell us your notice period when you apply — candidates who can join immediately or at short notice will be prioritised.
n
Our process is a CV review, a short written assignment, and interviews.
n
If you have the experience to build and own an on-premises data platform from the ground up, we'd like to hear from you.
n
You can share your CVs at:
[email protected]
📌 Data Engineer (Noida)
🏢 Algo8 AI
📍 Noida