? We're are Hiring @Algo8AI
Role:
Data Engineer — Data Platform & Pipelines
Lake, canonical model and data quality
? Location: Noida | Hybrid (4 days/week in office)
? Experience: 5–7 years
? Employment: Permanent | Algo8 Payroll
? Reports to: Tech Lead
⏱️ Joining: Early joiners / short notice preferred
You will build and own the spine of the programme — the on-premises data platform, the canonical model that everything lands into, and the quality framework that determines whether the foundation can be trusted.
About Algo8
Algo8 is an industrial and manufacturing AI firm. We work inside plants and R&D; centres — with engineers, scientists and plant teams — building the data foundations and AI capabilities that enable industrial organisations to make better decisions. Our work is hands-on and close to the client.
The Engagement
You will join a multi-year programme with the R&D; centre of excellence of a large Indian manufacturer.
The centre holds four decades of engineering knowledge — design records, formulations, laboratory results, test data, reports and studies — spread across enterprise systems, departmental databases, local servers and document archives.
The programme recovers that knowledge into a single governed foundation, connects it to manufacturing context, and then validates where engineering intelligence can be built on top of it.
This is a data engineering programme first. The AI comes later, and only where the evidence supports it.
The solution is on-premises.
Because of the proprietary nature of new product development, this foundation is being built inside the client's own infrastructure.
If your experience is entirely on managed cloud services, this will not be the right fit.
What You Will Own The platform
Stand up and operate the on-premises data lake and processing layer — storage, compute, orchestration and access control.
The canonical model
Work with engineers and business analysts to design a shared data model across engineering and manufacturing domains, then keep it coherent as new sources land.
Ingestion pipelines
Build the pipelines that bring recovered and synchronised data into the foundation, with lineage from source to consumption.
Data quality
Define and implement quality thresholds per source. Each estate has to pass an acceptance gate before it is declared complete — you own the evidence behind that.
Enabling everyone else The integration and document engineers land data into your platform. Analytics and validation work builds on top of it. Making both easy is part of the job.
The Expected Stack The architecture is being finalised with the client. We expect to work with an on-premises stack along these lines, or close equivalents:
- Distributed processing and query — Spark, Trino or equivalent
- On-premises object storage — MinIO, HDFS or equivalent
- Relational stores — PostgreSQL, Oracle, SQL Server
- Orchestration — Airflow or equivalent
- Streaming and change capture — Kafka, Debezium or equivalent
We are more interested in whether you have run this class of system on your own infrastructure than in which specific tools you have used.
What We Need From You
- Strong on-premises data platform experience — you have built and run distributed processing and storage on infrastructure you or your organisation managed, not only on managed cloud services.
- Data modelling depth — you can design a canonical model across messy sources and defend the decisions behind it.
- Production pipeline engineering — Spark or equivalent, Python, solid SQL, and orchestration with Airflow or similar.
- Data quality and governance in practice — validation frameworks, lineage and reconciliation. Not as theory, but as something you have actually implemented.
- Clear communication — you will explain to non-technical stakeholders why data is or is not ready,
and that conversation matters as much as the pipeline.
Useful, Not Essential
- Manufacturing or industrial data experience
- Kafka, Debezium or other change-data-capture tooling
- Metadata catalogue and governance tooling
- Experience defining acceptance criteria on a client programme
How You Will Work
- Based in Noida, hybrid — four days a week in office.
- Reporting to the Tech Lead for the programme, alongside a large on-site business analyst team who own requirements gathering and client expectations.
- Occasional travel to client plant and R&D; locations as the work requires.
- You will be working with an existing enterprise estate that is not fully documented.
Discovery is part of the job — expect to map what is actually there before you design against it.
Who This Role Is Not For
We would rather be direct than waste your time. This role is probably not a fit if:
- Your data engineering experience is entirely on managed cloud services and you have never sized, deployed or operated infrastructure yourself.
- You are a BI or reporting developer — dashboards, SQL views and reports — rather than someone who has built and owned pipelines.
- Your work has been drag-and-drop ETL tooling without engineering underneath it.
- You are an application or backend developer whose data work has been incidental.
- You have never worked with enterprise systems such as ERP, PLM or MES.
Working at Algo8
- You will build capability that outlasts this engagement. What you learn here is directly reusable across our industrial client base.
- Small senior team, high ownership, and direct access to client engineering leadership.
Applying
We need this team in place early.
Please tell us your notice period when you apply — candidates who can join immediately or at short notice will be prioritised.
Our process is a CV review, a short written assignment, and interviews.
If you have the experience to build and own an on-premises data platform from the ground up, we'd like to hear from you.
You can share your CVs at:
[email protected]
📌 Data Engineer (Noida)
🏢 Algo8 AI
📍 Noida