Experience : 5.00 + years
Salary : Confidential (based on experience)
Expected Notice Period : 7 Days
Shift : (GMT+05:30) Asia/Kolkata (IST)
Opportunity Type : Remote
Placement Type : Full Time Contract for 5 Months(40 hrs a week/160 hrs a month)
**(*Note: This is a requirement for one of Uplers' client - PredictiveOps AI)
**What do you need for this opportunity?
Must have skills required: Predictive Maintenance, Reliability Engineering, Backend Development, LLM-based Agent Systems, Machine Learning (production), MLOps, Python (advanced), Cloud Server (Google / AWS), React Js, RESTAPI
PredictiveOps AI is Looking for:**** We are is building an AI-powered predictive intelligence platform for data centre infrastructure.
Data centres are among the most critical infrastructure being built in the world, and they are run reactively. Existing tools report what is happening right now, not what is coming. At the same time, replacement lead times for critical equipment have gone from months to years. So when equipment fails or a delivery slips, operators find out too late to act cheaply, and a facility can run exposed for months or a build can miss its energisation date.
Critical information sits fragmented across spreadsheets, emails, procurement records, monitoring systems and project schedules that do not talk to each other.
We connect those signals and makes them predictive. Our long term vision is to become the intelligence layer for data centre infrastructure, moving operators from reactive workflows to predictive, AI-assisted decision making.
The platform combines predictive models and specialised AI agents to:
- Predict equipment failures and supply chain delays 30 to 90 days ahead
- Explain the evidence and reasoning behind every prediction
- Compute downstream impact and the date by which action must be taken
- Recommend and draft the actions required
- Preserve human oversight over consequential decisions
We are conducting customer discovery and preparing to build the first product alongside a small number of paid design partners. Initial product direction Predictive supply chain and infrastructure risk for data centre operators.
Critical equipment carries long lead times and complex dependencies. A delayed component or a degrading asset can aQect installation, commissioning, energisation and the facility go-live date, or leave a live site without redundancy for months.
Likely initial capabilities:
- Identifying supplier, procurement and schedule risk
- Connecting risks to downstream dependencies and computing cost and schedule impact
- Producing calibrated confidence scores with supporting evidence
- Computing the last responsible date to act
- AI agents that monitor changes, investigate risks, draft procurement documents and coordinate follow-up
The exact first workflow and user persona will be refined through customer discovery. The objective is one narrow, high value, repeatable use case, not bespoke software per customer. The role We are seeking a Founding AI/ML Engineer to become the company''''''''s first engineering hire and build the technical foundation of the product.
You will work directly with the founder, domain experts and early design partners to turn complex operational problems into a reliable, commercially valuable AI product.
This is a zero to one role with ownership across machine learning, agent systems, architecture, data infrastructure, backend engineering, evaluation and monitoring, and enterprise deployment.
This is not a research role. The goal is to translate sophisticated approaches into a product customers can deploy, understand and trust.
Key Responsibilities Predictive modelling This is the core of the product and where most of the technical diQiculty sits.
- Build and evaluate predictive models on real operational data: equipment telemetry, maintenance records, failure histories, purchase order histories and project schedules
- Select appropriate modelling approaches for time-to-event prediction — horizon-windowed classification, survival analysis, time series forecasting,
anomaly detection, probabilistic modelling — and justify the choice against the data
- Handle sparse events and small samples. Failure events are rare, and per-asset history is limited. Approaches such as hierarchical modelling, pooling across asset classes, transfer learning and informed priors will matter
- Produce calibrated probabilities. If the model says 70%, it must be right about 70% of the time. Calibration is a product requirement, not an internal metric, because operators make capital decisions on these numbers
- Build leakage-safe training and evaluation pipelines, including correct temporal validation where data drifts over time
- Measure precision, recall, calibration, lead time and operational usefulness — not accuracy alone
- Build backtesting infrastructure that proves accuracy against a customer''''''''s own history, so value can be demonstrated inside a short pilot despite long prediction horizons
- Quantify uncertainty and make it visible in the product
Trustworthy and explainable AI Operators will not act on a black box when the decision costs seven figures.
- Every prediction must surface a calibrated confidence score, its data sources, and a full reasoning chain that a domain expert can interrogate
- Design human-in-the-loop controls for consequential recommendations
- Detect and mitigate model failure modes: drift, degradation, bias in training data, and confidently wrong outputs caused by bad inputs such as sensor faults
- Build monitoring that catches silent failures, including stale data feeds producing confident but obsolete predictions
- Maintain a predictions ledger that scores every prediction against what actually happened, visible to the customer, misses included
AI agents and autonomous workflows
- Build specialised agents for monitoring, investigation, sourcing, document generation and workflow coordination
- Design tool use so agents query approved data sources and internal functions, and cannot assert anything not returned by a tool
- Implement grounding, retrieval, structured outputs and abstention so agents escalate rather than invent when confidence is low
- Prevent hallucination structurally rather than by instruction. Consequential numbers must be computeddeterministically, never generated
- Implement permissions, audit logs and human approval checkpoints
- Build evaluation harnesses for agent behaviour: grounding quality, action quality, restraint, and regression testing when prompts or models change
- Judge when agentic workflows add value and when conventional software is the better answer
Platform, data and deployment engineering
- Define the initial technical architecture
- Build backend services, APIs and core product infrastructure
- Design schemas and pipelines for assets, suppliers, components, purchase orders, milestones and dependencies
- Ingest and normalise messy real-world data from spreadsheets, exports, documents and limited APIs, with a mapping layer that does not require rewriting per customer
- Build the operator-facing interface showing risks, confidence, evidence and recommended actions
- Deploy into isolated per-customer environments. Each customer runs in their own dedicated cloud tenant (VPC) or on-premise container. Raw operational data never leaves the customer''''''''s environment. Only de-identified patterns travel centrally, and that boundary must be enforced in code
- Implement authentication, permissions, organisation management and customer data isolation
- Build deployment automation so standing up a new customer environment is repeatable, not manual
- Establish testing, logging, monitoring and model observability,
including self-health reporting from environments we cannot freely access
- Optimise inference cost and latency for production
Running it in production The role is build and run. Once deployments are live, ongoing ownership includes system health monitoring, accuracy and drift tracking, retraining pipelines, diagnostics, and fixing customer-facing bugs quickly.
Customer collaboration
- Join discovery and design partner meetings
- Assess technical feasibility, data availability and deployment requirements
- Identify the minimum viable dataset for a pilot
- Translate customer feedback into focused product improvements
- Communicate technical trade-oQs clearly to technical and non-technical audiences
What Success Looks Like In The First Five Months
- A validated initial use case
- A documented technical architecture and roadmap
- A functional core platform with working data ingestion and validation
- Predictive models with defensible, calibrated performance on real customer data
- An evaluation framework covering performance, calibration and operational value
- Predictions carrying confidence, evidence and traceable sources
- At least one specialised agent workflow with appropriate controls
- A repeatable process for deploying customer pilots into isolated environments
- At least one design partner pilot deployed or progressing to deployment
Essential Experience
- Strong software engineering fundamentals and advanced Python
- Demonstrable experience building and deploying machine learning systems in production, not only research or notebooks
- Experience with messy, sparse, real-world operational data
- Solid grounding in model evaluation, calibration and validation methodology
- Experience with LLM-based agent systems: tool use, grounding, structured outputs, evaluation
- Backend services, APIs, databases and cloud infrastructure
- Data pipelines, MLOps and production monitoring
- Strong technical judgement when requirements are incomplete
- Clear written and verbal communication, including the ability to explain and defend technical decisions
- High ownership and comfort in an early-stage workplace
Particularly valuable
- Predictive maintenance, reliability engineering, or industrial and infrastructure systems
- Time series forecasting or horizon-based event prediction
- Survival analysis or probabilistic modelling
- Anomaly detection on sensor or telemetry data
- Retrieval augmented generation and knowledge graphs
- Multi-tenant enterprise deployment and data isolation
- Frontend product engineering
- AWS or equivalent cloud infrastructure
- Reinforcement learning, particularly for sequential decision-making under uncertainty
Direct data centre experience is helpful but not required. The ability to learn a specialised domain quickly and translate operational problems into technical systems matters more.
Why join
- Become the first engineer at an ambitious AI infrastructure company
- Own major architecture, machine learning and product decisions
- Build trustworthy predictive models and agent systems for physical infrastructure
- Work directly with customers in a rapidly growing global industry
- Help define our engineering culture and future technical team
How to apply for this opportunity?
- Step 1: Click On Apply! And Register or Login on our portal.
- Step 2: Complete the Screening Form & Upload updated Resume
- Step 3: Increase your chances to get shortlisted & meet the client for the Interview!
About Uplers: Our goal is to make hiring reliable, simple, and fast. Our role will be to help all our talents find and apply for relevant contractual onsite opportunities and progress in their career. We will support any grievances or challenges you may face during the engagement. (Note: There are many more opportunities apart from this on the portal. Depending on the assessments you clear, you can apply for them as well).
So, if you are ready for a new challenge, a great work environment, and an opportunity to take your career to the next level, don't hesitate to apply today. We are waiting for you!
📌 Founding AI/ML Engineer (India)
🏢 Uplers
📍 India