04 Oct
|
Cloudxtreme
|
Hyderabad
04 Oct
Cloudxtreme
Hyderabad
Role & responsibilities
Machine Learning Engineer Intelligent Payroll (TPY)
Team: HC Forward — Intelligent Payroll Role type: Individual contributor, builder track (no client engagement) Experience: 8–10 years hands-on building production data/ML systems Openings: 2
---
Human Capital — HC Forward
The Human Capital Offering Portfolio helps organizations manage and sustain their performance through their most important asset: their people. We see Human Capital as a topic critical to the C-Suite, and we bring together technology, insights, and services to help our clients innovate faster and proactively drive their business strategies.
HC Forward is Deloitte’s innovation engine for Human Capital — integrating technology, data, and industry expertise to create scalable solutions and assets that extend client capabilities and drive ongoing value across all Human Capital offerings.
About Intelligent Payroll
Intelligent Payroll is an HC Forward platform that uses machine learning to detect anomalies, forecast variances, and reduce manual intervention in enterprise payroll processing. It runs on Azure, ingests data from multiple HRIS sources (SAP, Oracle, ADP, Workday), normalises it into a common data model, and serves both batch ML pipelines and rule-based validation engines.
This is a builder role. You will spend your time writing production Python code, designing pipelines, and shipping features. You will not be in client workshops, executive presentations, or pre-sales.
---
Work you’ll do
- Design and ship distributed data processing pipelines in PySpark, moving raw payroll data through a medallion architecture (raw standardised aggregated modelled)
- Build ML pipelines for time-series forecasting and anomaly detection over payroll data
- Write clean, class-based, testable Python — interfaces, factories, dependency injection, not 800-line procedural scripts
- Submit and orchestrate Spark jobs programmatically on a managed cloud ML platform (we use Azure ML Studio, but the patterns generalise)
- Own the contracts between subsystems: data pipelines, ML libraries, and orchestration layers are intentionally decoupled, and you will help keep them that way
- Write tests, write docs, review PRs,
and treat the pipeline as a product, not a notebook
---
Required Skills
These are non-negotiable. If you don’t have these, this isn’t the right role.
Production PySpark
- You have shipped PySpark code that ran in production — not just one notebook, but a pipeline that processed real data on a recurring schedule.
- You know the DataFrame API, schema validation, parquet I/O, partitioning, and how to debug a job that’s slow or wrong.
- You can write a PySpark job that runs locally for development and at scale in the cloud, without changing the business logic.
Strong Python & OOP
- You write production Python by default: classes, modules, type hints, tests.
- You know when to reach for an abstract base class, a factory, or dependency injection — and when not to.
- You can read someone else’s class-based codebase and explain what it does, end to end.
Engineering experience
- You design for boundaries: schemas between steps, interfaces between modules, contracts between services.
- You ask about data quality before you ask about model choice.
- You can work from a config file, a CLI entry point, and a YAML pipeline definition — not just python script.py.
Cloud ML platforms
- You have submitted Spark or ML jobs programmatically using a cloud ML SDK (Azure ML, AWS SageMaker / boto3, Vertex AI) and understand cloud environments.
Data science fundamentals
- Hands-on with core ML algorithms, with a solid grounding in regression.
- Time-series modelling is a plus, not required.
---
Desired Skills
You don’t need all of these. You don’t even need most. But each one is a real signal.
Area What it looks like
Cloud ML platforms Comfortable working with cloud ML SDKs (Azure ML, AWS SageMaker / boto3, Vertex AI) and
Area What it looks like submitting Spark or ML jobs programmatically; understands cloud environments
Azure identity model You can explain DefaultAzureCredential, managed identity vs service principal, why Key Vault exists
Time-series ML Statsforecast, ARIMA family, cross-validation, season-length tuning, stationarity (ADF), autocorrelation (ACF) — you’ve actually used these, not just read about them
Data engineering at scale Medallion architecture, slowly changing dimensions, schema evolution, dedup logic, CDC patterns
FastAPI / API design Built APIs with Pydantic schemas, dependency injection, async endpoints
SQL Advanced SQL
Contemporary Python tooling uv, ruff, pyproject.toml, pre-commit hooks
Domain context HR, payroll, or financial-services data — useful but learnable
---
What we are explicitly not looking for
To save everyone’s time:
- Notebook-only data scientists. If your entire portfolio is Jupyter notebooks with no production deployment, this isn’t the role. We have huge respect for DS work — it’s just not what these two seats are for.
- Pure model-builders. We need people who think about pipelines, contracts, and systems, not just model accuracy.
- Generic GenAI experience without engineering fundamentals. LangChain on a side project does not substitute for production Python.
- Client-facing consultants. No pre-sales, no workshops, no “translate ML to the C-suite.”
---
Qualifications
- Bachelor’s degree in Computer Science, Mathematics, Statistics, Engineering, or a related quantitative field. Master’s degree helpful, not required.
- 8–10 years hands-on building production data or ML systems.
- A code portfolio we can look at — GitHub, sample code, a take-home, or a detailed walkthrough of past work.
Location - All USI locations i.e, Hyderabad, Bengaluru, Chennai, Gurgaon, Mumbai, Pune & Kolkata. Candidate should be willing to work in Hybrid mode from any of these office locations from Day 1.
Notice Period – Immediate to 15 days
No. Of Interview Rounds – 2
Mode Of Interview – Virtual
PFA the updated JD & screening sheet. Ensure that the candidates are screened by your technical panels using the screening sheet. Only then, add the profiles.
📌 Data Scientist (Hyderabad)
🏢 Cloudxtreme
📍 Hyderabad