17 Sep
|
Crisil
|
Hyderabad
We are seeking a versatile engineer to build the data foundation of our agentic AI analytics platform, built on
LangGraph, Azure OpenAI, AWS Bedrock, Databricks and Postgres and to establish how we measure the quality of what that platform produces. Our workflows combine market data, proprietary S&P; Global datasets, and analystbuilt models into repeatable analytical products used across S&P; Global Energy and in client-facing settings. Two things determine whether those products can be trusted at scale: reliable, structured, well-governed access to the underlying data, and objective evidence that the output is correct.
This role owns both.
You will design and maintain the pipelines, databases, and integrations that allow AI agents to work directly with trusted energy market data. You will also build the reference datasets, scoring methods, and regression tests that establish how a change to a prompt, a model, or a retrieval step affects output quality. You will work alongside the platform and release engineer on the data layer, alongside the agentic AI engineer on evaluation, and alongside subject matter experts in refining, crude, and related markets to understand what the data means, not only how it is shaped.
Responsibilities
Data and integration
- Design, build, and maintain data pipelines that deliver structured, AI-ready data to analytical workflows and AI agents.
- Integrate internal and external data sources through APIs, cloud data platforms, and modern agent-tool protocols, including credential, entitlement, and environment configuration.
- Design, migrate, and maintain relational database schemas supporting workflow persistence, metadata, and analytical outputs.
- Extend and modernize analytical and refining databases as coverage grows, with automated update processes and explicit technical documentation.
- Normalize and structure raw data so that it can be consumed directly by AI agents within automated workflows.
- Build monitoring and validation that maintains data quality across sources and refresh cycles, including completeness, consistency, and timeliness checks.
- Work with subject matter experts and analysts to understand domain data, its provenance, and its limitations.
- Document data flows, schemas, and integration points so that the work is transferable and auditable.
- Support deployment and security readiness of data components, including access control, entitlements,
and audit requirements.
Evaluation and quality
- Design and build evaluation harnesses that score AI-generated output systematically for accuracy,
completeness, consistency, and source traceability.
- Establish ground-truth and gold-standard reference sets in partnership with subject matter experts,
including training and holdout methodology.
- Build regression test suites so that changes to prompts, models, retrieval, or workflow logic are validated ahead of release.
- Run structured comparisons across AI models and configurations, measuring quality, latency, and compute cost.
- Define and detect the characteristic failure modes of generative systems, including unsupported statements,
incomplete field population, conflicting values, and retrieval drift.
- Build discrepancy detection that surfaces conflicts between sources as explicit, reviewable flags.
- Work with domain experts to translate expert quality judgments into testable, repeatable criteria.
- Report findings clearly to both engineers and senior leadership, and drive prioritization of the resulting improvements.
- Supply evidence of systematic validation to governance, risk, and security review processes.
S&P; Global External
August 2026
Core Requirements
All candidates should be able to demonstrate the following, whatever path they took to acquire it:
- Proficiency in Python and SQL, sufficient to build both data pipelines and test harnesses and to analyze the results independently.
- Practical experience building data pipelines or automated data processes that other people relied on in production, including what happened when they failed.
- Working knowledge of relational databases and data modeling, including schema change against systems with live consumers.
- Demonstrated experience evaluating, validating, or testing analytical output, models, or systems against a defined standard.
- Sound grasp of statistics and experimental design, including sampling, controls, baselines, and sources of bias.
- Practical understanding of generative AI applications and prompt engineering, shown through something you have built rather than tools you have tried.
- Intellectual honesty and precision: a willingness to report an unwelcome result and defend the method that produced it.
- Clear written communication, including the ability to explain a measurement or a data model to someone who did not design it.
Backgrounds We Will Consider
- Data or analytics engineering with production ownership of the pipelines you built.
- AI or machine learning evaluation, large language model evaluation, or applied data science.
- Model validation, model risk management, or quantitative audit in banking, insurance, energy, or a similar regulated setting.
- Scientific or academic research background with strong experimental methodology, in any discipline.
- Refining, process, or engineering role in which you validated models or simulations against plant or market reality and write code.
- Software quality or test automation engineering with genuine analytical depth.
- Market or research analyst with strong quantitative method, coding ability, and a documented habit of checking things.
Degrees in engineering, the sciences, statistics, mathematics, economics, computer science, or a related field are all relevant, as is equivalent practical experience without a matching degree. Method and evidence matter more to us in this role than any particular credential or industry.
Preferred Skills
- Postgres, Databricks, or Azure data services run in production.
- Familiarity with agentic orchestration frameworks such as LangGraph or LangChain, retrieval-augmented generation, or prompt versioning.
- Direct experience evaluating large language model or generative AI output, including with tooling such as
LangSmith and Ragas.
- Experience with model documentation, governance frameworks, or regulatory validation standards.
- Understanding of token and compute cost economics in AI systems.
- Experience working in a refinery or chemical industry, or in commodities markets, or with energy domain data such as crude and refined products, trade flows, or price assessment.
- Experience in developing and deploying machine learning models in a business context.
- Experience with data visualization and reporting tools, such as Power BI, Tableau, or Matplotlib
📌 Gen AI Python Developer (Data integration & evaluation engineer) (Hyderabad)
🏢 Crisil
📍 Hyderabad