About the Role:
Grade Level (for internal use):
11
Team:
As a lead data scientist in the EDO, Collection Platforms & AI - Cognitive Engineering team, you will own the technical direction of a large-scale entity resolution and linking
platform used across S&P; Global. This platform combines classical ML and GenAI techniques to power intelligent matching and linking services across multiple data domains, exposed as
production APIs and orchestrated through downstream workflow systems. You will set technical strategy, champion the adoption of LLMs and agentic AI across the platform, architect its
evolution, mentor senior and junior data scientists, and be the final technical authority on design decisions across the codebase. You will work in a truly global team and be expected to
drive thoughtful risk-taking, cross-team alignment, and technical/AI innovation.
What's in it for you:
- Own the technical vision for an enterprise-scale AI/ML platform used across S&P; Global
- Be the org's go-to authority on LLMs and agentic AI, with room to prototype and productionize novel approaches
- Lead, grow, and set technical direction for a highly skilled team of data scientists and engineers
- Architect solutions to the highest-complexity, highest-impact problems in the org, end-to-end
- Have final say on architecture, tooling, and technical trade-offs for a multi-domain production system
Responsibilities:
- Own the end-to-end technical architecture of the platform, including its API service layer, asynchronous distributed task orchestration, and multi-domain model registry
- Set the AI/ML roadmap across both classical ML (gradient boosting, embedding-based similarity, fuzzy/probabilistic matching) and GenAI approaches (LLM-based extraction, prompt engineering,
fine-tuning/customization, agentic and multi-agent workflows)
- Champion innovation with agentic workflows: design, prototype, and productionize multi-step, tool-using LLM agents that can plan, reason, call tools/APIs, and collaborate with other agents
to solve complex linking problems
- Evaluate and drive adoption of state-of-the-art agentic AI infrastructure, including MCP (Model Context Protocol) servers for tool/context integration, agent orchestration frameworks, agent
memory/state management, and retrieval-augmented generation (RAG) with vector stores
- Architect and oversee productionization of models and pipelines - packaging, versioning, deployment to containerized/orchestrated environments , and integration with
cloud infrastructure (object storage, secrets/parameter management)
- Define and enforce MLOps and LLMOps standards: CI/CD for models and prompts, experiment tracking, model/prompt/agent versioning, automated evaluation, and monitoring for both synchronous
services and background task workers
- Lead technical design reviews and code reviews across the codebase; set and enforce coding standards, testing discipline, and architectural consistency across all modules
- Drive build-vs-buy and fine-tune-vs-prompt-vs-agent decisions across LLM providers, embedding models, and agentic frameworks used across the platform
- Continuously scout the LLM/agentic AI landscape (new models, protocols, and tooling) and run rapid proof-of-concepts to assess production fit
- Own production reliability of deployed services - diagnosing and resolving issues across API, async tasks, databases, caching, and LLM/agent layers
- Mentor senior and junior data scientists on both classical ML and modern LLM/agentic AI techniques; run technical onboarding for current team members joining the platform
- Manage stakeholders across engineering, workflow orchestration teams, and business domain owners to align on roadmap and delivery timelines
- Represent the team's technical decisions and AI innovation initiatives to senior leadership and influence org-level AI tooling and platform standards
Technical Requirements:
- Deep,
hands-on experience architecting production systems combining classical ML (gradient boosting: LightGBM/XGBoost, scikit-learn) with GenAI (LLMs from providers such as Gemini, OpenAI,
Anthropic; prompt engineering; fine-tuning/customization; embedding-based retrieval via sentence-transformers/Hugging Face)
- Proven experience designing and shipping LLM-powered agents and agentic workflows - planning/reasoning loops, tool use, task decomposition, and multi-agent collaboration - beyond
single-turn prompting
- Working knowledge of MCP (Model Context Protocol) and other emerging standards for connecting LLMs/agents to tools, data sources, and context; experience integrating or building MCP servers
- Understanding of RAG architectures and vector retrieval systems, and when to apply them versus fine-tuning, agentic tool-use, or classical search
- Familiarity with agent orchestration frameworks (e.g., Google ADK, or similar)
- Expert proficiency in Python and its data/ML ecosystem (Pandas, NumPy, PyTorch/TF, Transformers, scikit-learn) at a level sufficient to review and set standards for a large, multi-module
production codebase
- Strong understanding of asynchronous, distributed task architectures (task queues, message brokers) and how they interact with ML/LLM inference at scale
- Experience architecting and operating web API services (e.g., FastAPI or comparable) in production, including containerization and orchestration
- Working knowledge of cloud infrastructure (object storage, secrets/parameter management) and relational databases as used in production ML services
- Deep understanding of entity resolution / record linkage techniques: fuzzy matching, name matching, phonetic algorithms, embedding-based similarity, and how to evaluate/benchmark them
- Strong grasp of statistics, probability, and the mathematics underpinning both classical ML and modern GenAI/agentic systems
- Demonstrated ability to track and synthesize current AI/ML research (including agentic AI, RAG, tool-use paradigms, and evaluation methods for LLM-based systems) and drive its adoption into
a production platform
- Proven track record leading at least one large-scale ML, GenAI, or agentic AI platform from design through production, including having made and defended architectural trade-offs
- Familiarity with MLOps/LLMOps and observability tooling sufficient to define standards for the team
Good to have:
- 8+ years of relevant experience in Data Science/AI/ML engineering, with at least 3-4 years leading technical direction for a team or platform
- Hands-on track record of shipping at least one production agentic AI system (multi-step agents, tool-calling, or multi-agent orchestration), not just prototypes
- Prior experience with entity resolution, record linkage, or master data management systems, especially in financial/market-intelligence or risk analytics domains
- Experience owning architecture decisions for systems that span multiple ML paradigms (classical ML + GenAI + agentic AI) in the same production platform
- Public contributions or demos on GitHub, Kaggle, StackOverflow, technical blogs, or publications - especially around LLMs, agents, or agentic tooling
What's In It For You?
Our Mission:
Advancing Essential Intelligence.
Our People:
We're more than 35,000 strong worldwide-so we're able to understand nuances while having a broad perspective.
Our team is driven by curiosity and a shared belief that Essential Intelligence can help build a more prosperous future for us all.From finding new ways to measure sustainability to analyzing energy transition across the supply chain to building workflow solutions that make it easy to tap into insight and apply it. We are changing the way people see things and empowering them to make an impact on the world we live in. We're committed to a more equitable future and to helping our customers find new, sustainable ways of doing business. Join us and help create the critical insights that truly make a difference.
Our Values:
Integrity, Discovery, Partnership
Throughout our history, the world's leading organizations have relied on us for the Essential Intelligence they need to make confident decisions about the road ahead. We start with a foundation of integrity in all we do, bring a spirit of discovery to our work, and collaborate in close partnership with each other and our customers to achieve shared goals.
Benefits:
We take care of you, so you can take care of business. We care about our people. That's why we provide everything you-and your career-need to thrive at S&P; Global.
Our benefits include:
- Health & Wellness: Health care coverage designed for the mind and body.
- Flexible Downtime: Generous time off helps keep you energized for your time on.
- Continuous Learning: Access a wealth of resources to grow your career and learn valuable new skills.
- Invest in Your Future: Secure your financial future through competitive pay, retirement planning, a continuing education program with a company-matched student loan contribution, and financial wellness programs.
- Family Friendly Perks: It's not just about you. S&P; Global has perks for your partners and little ones, too, with some best-in class benefits for families.
- Beyond the Basics: From retail discounts to referral incentive awards-small perks can make a big difference.
For more information on benefits by country visit: https://spgbenefits.com/benefit-summaries
Global Hiring and Opportunity at S&P; Global:
At S&P; Global, we are committed to fostering a connected and engaged workplace where all individuals have access to opportunities based on their skills, experience, and contributions. Our hiring practices emphasize fairness, transparency, and merit, ensuring that we attract and retain top talent. By valuing different perspectives and promoting a culture of respect and collaboration, we drive innovation and power global markets.
Recruitment Fraud Alert:
If you receive an email from a spglobalind.com domain or any other regionally based domains, it is a scam and should be reported to
[email protected] . S&P; Global never requires any candidate to pay money for job applications, interviews, offer letters, "pre-employment training" or for equipment/delivery of equipment. Stay informed and protect yourself from recruitment fraud by reviewing our guidelines, fraudulent domains, and how to report suspicious activity here .
-----------------------------------------------------------
Equal Opportunity Employer
S&P; Global is an equal opportunity employer and all qualified candidates will receive consideration for employment without regard to race/ethnicity, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, marital status, military veteran status, unemployment status, or any other status protected by law. Only electronic job submissions will be considered for employment.
If you need an accommodation during the application process due to a disability, please send an email to:
[email protected] and your request will be forwarded to the appropriate person.
US Candidates Only: Know Your Rights: Workplace discrimination is illegal
-----------------------------------------------------------
20 - Professional (EEO-2 Job Categories-United States of America), IFTECH202.2 - Middle Professional Tier II (EEO Job Group), SWP Priority - Ratings - (Strategic Workforce Planning)
📌 Lead Data Scientist (Hyderabad)
🏢 S&P Global
📍 Hyderabad