Data Scientist (Hyderabad)

Data Scientist (Hyderabad)

04 Oct
|
Cloudxtreme
|
Hyderabad

04 Oct

Cloudxtreme

Hyderabad

Role & responsibilities

Machine Learning Engineer Intelligent Payroll (TPY)

Team: HC Forward — Intelligent Payroll Role type: Individual contributor, builder track (no client engagement) Experience: 8–10 years hands-on building production data/ML systems Openings: 2

---

Human Capital — HC Forward

The Human Capital Offering Portfolio helps organizations manage and sustain their performance through their most important asset: their people. We see Human Capital as a topic critical to the C-Suite, and we bring together technology, insights, and services to help our clients innovate faster and proactively drive their business strategies.

HC Forward is Deloitte’s innovation engine for Human Capital — integrating technology, data, and industry expertise to create scalable solutions and assets that extend client capabilities and drive ongoing value across all Human Capital offerings.

About Intelligent Payroll

Intelligent Payroll is an HC Forward platform that uses machine learning to detect anomalies, forecast variances, and reduce manual intervention in enterprise payroll processing. It runs on Azure, ingests data from multiple HRIS sources (SAP, Oracle, ADP, Workday), normalises it into a common data model, and serves both batch ML pipelines and rule-based validation engines.

This is a builder role. You will spend your time writing production Python code, designing pipelines, and shipping features. You will not be in client workshops, executive presentations, or pre-sales.

---

Work you’ll do

- Design and ship distributed data processing pipelines in PySpark, moving raw payroll data through a medallion architecture (raw standardised aggregated modelled)
- Build ML pipelines for time-series forecasting and anomaly detection over payroll data
- Write clean, class-based, testable Python — interfaces, factories, dependency injection, not 800-line procedural scripts
- Submit and orchestrate Spark jobs programmatically on a managed cloud ML platform (we use Azure ML Studio, but the patterns generalise)
- Own the contracts between subsystems: data pipelines, ML libraries, and orchestration layers are intentionally decoupled, and you will help keep them that way
- Write tests, write docs, review PRs,



and treat the pipeline as a product, not a notebook

---

Required Skills

These are non-negotiable. If you don’t have these, this isn’t the right role.

Production PySpark

- You have shipped PySpark code that ran in production — not just one notebook, but a pipeline that processed real data on a recurring schedule.
- You know the DataFrame API, schema validation, parquet I/O, partitioning, and how to debug a job that’s slow or wrong.
- You can write a PySpark job that runs locally for development and at scale in the cloud, without changing the business logic.

Strong Python & OOP

- You write production Python by default: classes, modules, type hints, tests.
- You know when to reach for an abstract base class, a factory, or dependency injection — and when not to.
- You can read someone else’s class-based codebase and explain what it does, end to end.

Engineering experience

- You design for boundaries: schemas between steps, interfaces between modules, contracts between services.
- You ask about data quality before you ask about model choice.
- You can work from a config file, a CLI entry point, and a YAML pipeline definition — not just python script.py.

Cloud ML platforms

- You have submitted Spark or ML jobs programmatically using a cloud ML SDK (Azure ML, AWS SageMaker / boto3, Vertex AI) and understand cloud environments.

Data science fundamentals

- Hands-on with core ML algorithms, with a solid grounding in regression.
- Time-series modelling is a plus, not required.

---

Desired Skills

You don’t need all of these. You don’t even need most. But each one is a real signal.

Area What it looks like

Cloud ML platforms Comfortable working with cloud ML SDKs (Azure ML, AWS SageMaker / boto3, Vertex AI) and





Area What it looks like submitting Spark or ML jobs programmatically; understands cloud environments

Azure identity model You can explain DefaultAzureCredential, managed identity vs service principal, why Key Vault exists

Time-series ML Statsforecast, ARIMA family, cross-validation, season-length tuning, stationarity (ADF), autocorrelation (ACF) — you’ve actually used these, not just read about them

Data engineering at scale Medallion architecture, slowly changing dimensions, schema evolution, dedup logic, CDC patterns

FastAPI / API design Built APIs with Pydantic schemas, dependency injection, async endpoints

SQL Advanced SQL

Contemporary Python tooling uv, ruff, pyproject.toml, pre-commit hooks

Domain context HR, payroll, or financial-services data — useful but learnable

---

What we are explicitly not looking for

To save everyone’s time:

- Notebook-only data scientists. If your entire portfolio is Jupyter notebooks with no production deployment, this isn’t the role. We have huge respect for DS work — it’s just not what these two seats are for.
- Pure model-builders. We need people who think about pipelines, contracts, and systems, not just model accuracy.
- Generic GenAI experience without engineering fundamentals. LangChain on a side project does not substitute for production Python.
- Client-facing consultants. No pre-sales, no workshops, no “translate ML to the C-suite.”

---

Qualifications

- Bachelor’s degree in Computer Science, Mathematics, Statistics, Engineering, or a related quantitative field. Master’s degree helpful, not required.
- 8–10 years hands-on building production data or ML systems.
- A code portfolio we can look at — GitHub, sample code, a take-home, or a detailed walkthrough of past work.

Location - All USI locations i.e, Hyderabad, Bengaluru, Chennai, Gurgaon, Mumbai, Pune & Kolkata. Candidate should be willing to work in Hybrid mode from any of these office locations from Day 1.

Notice Period – Immediate to 15 days

No. Of Interview Rounds – 2

Mode Of Interview – Virtual

PFA the updated JD & screening sheet. Ensure that the candidates are screened by your technical panels using the screening sheet. Only then, add the profiles.

📌 Data Scientist (Hyderabad)
🏢 Cloudxtreme
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: data scientist (hyderabad) / hyderabad

Subscribe to this job alert:

Get the latest job offers by email for: data scientist (hyderabad) / hyderabad