GroundTruth GroundTruth India BFSI Eval

Dataset release · GroundTruth

India BFSI Eval: systematically evaluating factuality and safety for India-specific finance assistants

General factuality suites measure whether models get facts right. Consumer BFSI chatbots in India fail in more specific ways: they apply the wrong country’s rules, cite superseded circulars, botch EMI and claim math, or helpfully explain illegal workarounds. India BFSI Eval is a gold set built to measure those failure modes—with reference answers that must survive official-source support checks before publication.

What we are releasing

Today we publish India BFSI Eval v0.1: a curated set of consumer-style Indian BFSI prompts with gold answers grounded in regulator and government sources (primarily RBI, SEBI, IRDAI, Income Tax / India Code, and related official pages). The public snapshot contains 167 items that cleared human review, question-quality checks, and source-support verification.

The benchmark is designed around four properties:

  • India-first jurisdiction. Prompts often omit an explicit “in India” cue. Correct behaviour still means Indian statutes and circulars—not FDIC, IRS, or SEC analogues.
  • Capability axes, not product taxonomies alone. Items are labelled by the failure mode they stress: jurisdiction, numerical reasoning, temporal currency, or red-team refusal—alongside product themes such as KYC, insurance, or credit cards.
  • Source-supported golds. A fluent gold answer is not enough. Citations must resolve, preferably to official domains, and the gold must be judged supported against those sources before it enters this set.
  • Human ∩ LLM question gate. Published items require human Approved status and LLM passes on question-quality rubrics before answer verification.

Browse every published item in the .

At a glance

English consumer prompts for Indian banking, insurance, markets, tax, and payments. Each item has one primary capability label and a product theme.

Capability slices in the public set
Slice Code N Share What a failure looks like

Motivation

Large language models are increasingly the first place consumers ask about deposits, loans, KYC freezes, mediclaim renewals, remittance limits, and tax withholdings. In India that stack is jurisdiction-specific: RBI master directions, SEBI circulars, IRDAI health and life rules, Income Tax Act updates, UPI and Aadhaar-linked identity, DICGC deposit insurance—not the US or EU analogues models absorb from pretraining.

General factuality benchmarks help measure whether a model can recall or retrieve true statements. They usually do not ask whether a chatbot quietly substituted FDIC rules for DICGC, quoted a superseded IRDAI turnaround time, invented a tax section number, or walked a user through hawala. For a consumer-facing BFSI assistant, those are the failures that create harm and regulatory exposure.

India BFSI Eval organises items by capability failure mode—the same axes product and safety teams use when writing eval plans—and publishes only golds that survive a multi-stage source-support pipeline. The goal is a gold set you can use to score chatbots on those risks before launch.

In particular, generic evals rarely stress-test:

  • Jurisdiction without an explicit “in India” hint in the user message
  • Which circular or product limit still applies when the year is omitted
  • Whether EMI, underinsurance, or TDS arithmetic is actually correct
  • Whether the model refuses illegal financial workarounds instead of partially complying

Scoring also needs care. Numerical items have method-and-number ground truth that can be checked objectively. Jurisdiction, temporal, and red-team items are more open-ended: we verify the gold against sources, and we recommend careful judge design when scoring model outputs—LLM judges alone can miss subtle wrong-country or outdated-rule answers that still sound fluent.

Capability slices

Every published item has one primary slice. Below: what we measure, and two illustrative prompts with a good vs weak answer pattern. Weak patterns are representative failure modes observed in finance chatbots—not quotes from a single model.

Composition

Beyond capability labels, items carry a product theme so teams can sample by banking, insurance, markets, and so on. Most golds cite official regulator or government pages; numerical items may use an inline method note instead.

Product themes
Theme N Share
Top citation domains
Domain family Items Share

Design choices

  1. Wrong-country answers are a first-class failure mode. Jurisdiction items deliberately omit “India” when a real user would. Scoring that rewards a fluent US answer as “helpful” will miss the behaviour BFSI products cannot ship.
  2. Supported ≠ exhaustive legal research. Source checks confirm the gold’s main claims are backed by citations. That is a publication bar, not a substitute for counsel review of a production assistant.
  3. Refusal is scored as behaviour, not as essay quality. Red-team golds reward clear refusal plus legitimate channels. Partial how-tos framed as “general knowledge” are treated as failures.
  4. Capability slices and product themes answer different questions. Sample by slice when you care about a failure mode; sample by theme when you care about coverage of insurance, KYC, credit cards, and so on.

Looking ahead

v0.1 is a starting public snapshot. Near-term directions include growing under-represented themes (tax, payments), hardening numerical scoring toward fully automatic checkers where possible, expanding temporal items tied to dated circulars, and expanding the shared model-run table beyond the first two open models we scored on jurisdiction and numerical slices.

If you are building or evaluating an India-facing BFSI assistant, start with the , sample by slice and theme, and treat unsupported or ungated drafts as out of scope for leaderboard-style claims.

Early model results

We scored two open models with search-enabled tool use on the jurisdiction and numerical slices (no India location hint in the prompt). Headline: Sarvam-30B leads Param2 on both axes. Full protocol and tables are on the page.

Cite

GroundTruth. India BFSI Eval v0.1.
Answer-verified prompts for Indian finance AI.
https://huggingface.co/spaces/ground-truth/india-bfsi-eval