vfairness · learn · test objects · checked 1 Oct 2026

Find a test object

Every fairness analysis needs something to test: a table of decisions, a model, a chat endpoint, an agent, a team of agents. This page tells you which object you need for the depth of analysis you want, and gives you a working one for every pathway, ready to download. Every file below was run through vfairness before it was published, and the page shows what happened.

Start here: what are you testing, and what can you see?

Pick the two answers. The result is your access tier, the strongest claim that tier supports, and a file to start with.

Access tier
Strongest claim
You can run
You cannot claim
Start with

What an assessment is, in 90 seconds

A fairness result is only as good as two things that must travel with it. Validant's article Precision Is Not Proof calls them the pointing and the seeing.

The pointing: where did we look?

Three dials. Body: which offering. Pathway: what kind of AI system (predictive, generative, agentic, multi-agent). Audience: who the report is for. The audience changes the wording, never the numbers. This page is about the pathway dial.

The seeing: how well could we see?

Access (A0 to A4): what the system lets you observe. Evidence (E0 to E3): how the test was designed. Validity (V0 to V3): whether the version you tested is pinned and still current. Written together, for example A2 · E2 · V1.

Three rules follow, and every object on this page is labelled with them in mind:

  • The weakest of the three sets the ceiling. Deep access does not rescue a stale, unpinned model, and a careful method does not rescue outputs you were only told about. In the article's words: taken at your word is Indicative; decisions measured is Limited; plus confidence scores is Reasonable.
  • A null result must say what it could have caught. "No disparity detected" is a finding only with its limiting magnitude: the smallest gap the test would have found, at a stated power.
  • Not seen is not the same as seen and fine. If your tier cannot reach a question, the honest answer is "could not check", never a pass.

"Pointing says where she looked. Seeing says what could possibly have been resolved from there."

The access ladder

The tier is set by what the system gives back, not by how you connect to it. A chat window, a company's wrapper API and the provider's own API all return text, so all three are A1. Tiers and their ceilings, quoted from the article:

TierWhat you haveHighest claim it supports
A0 AttestedVendor documentation, model card, published evaluations. No probing.The subject's own claims, recorded and checked for internal consistency.
A1 BehaviouralQuery access. Terminal output only: text, label, decision.Disparity in observed outcomes, present or absent at a stated sensitivity.
A2 ScoredA1 plus per-output scores: log probabilities, class probabilities, ranked alternatives, confidence.Calibration and threshold behaviour by group. Ranking and margin disparity.
A3 InternalWeights, activations, gradients. The ability to intervene on the computation.Which internal structures carry the disparity. Mechanistic attribution.
A4 ProvenanceA3 plus training-corpus lineage, fine-tuning and preference-data history, evaluation history.Where in the lifecycle the disparity came from.

Fairness is an input-output discipline and does well at A1: group metrics need decisions, labels and group membership, never weights. Explainability is different: a claim about why the model behaves as it does starts at A3. Agent traces (which tool was called, where a case was handed off) are behaviour you watch, so they are A1 too: they widen what you see at A1 without reaching inside the model.

What data you need for what depth

One row per pathway. Each cell names the fields you need and what they unlock.

PathwayA1 outputsA2 plus scoresA3 internalsA4 provenance
Before the model
the data alone
Features, protected attributes and the historical outcome, no model. You can check representation, label disparity in history and proxies (a column that stands in for a protected one). This is a data audit, not a model assessment: it says nothing about how a model will decide.
Predictive
tabular
y_pred + group. Selection-rate disparity, demographic parity. Add y_true and you also get error-rate parity: equalised odds, equal opportunity, predictive parity. + y_prob. Calibration by group, AUROC parity, multicalibration, threshold trade-offs. + the model (file, weights or a predict() you can call). Feature attribution, counterfactual what-if, retraining with constraints. + the training data and how it was built. Where the disparity entered.
Generative
LLM
Prompts in, text out. Counterfactual prompt pairs, benchmarks (BBQ, BOLD, HolisticBias), refusal, tone and sentiment disparity. + token log probabilities. Likelihood bias: the probability of "yes", of each answer option, which resolves differences a yes or no hides. + open weights. The model's own embeddings (WEAT, SEAT), hidden-state probes. + open training data. Which corpus patterns carry the bias.
Agent Final outcome + persona. Outcome disparity, correspondence tests. Add traces (tool calls, routes, steps) and you also get tool-selection, delegation and trajectory bias. + log probabilities at each step. The margin behind each tool choice. + an open-weight model inside the agent. Rarely available.
Multi-agent End-to-end inputs and outputs. Add per-agent outputs and you also get compositionality, emergent bias and groupthink: is the system more biased than its parts? As for agents.As for agents.Rarely available.
Synthetic All tiers at once, because you made it. Its unique value is the answer key: you know which bias was planted, so you can check that a tool finds it and stays quiet on the clean control.

Predictive (tabular)

A predictive system turns a row of facts about a person into a decision or a score: approve a loan, invite to interview, flag for review. The test object is a table with one row per person: the decision, the group, and as much else as you can get. The same source can be walked up the ladder, which is what the Adult pack is for.

UCI Adult, with a model's decisions

The 1994 US census extract every fairness course uses, with a logistic regression's predictions added (15,060 test rows). The model never sees sex or race; it selects 8.1 percent of women and 26.0 percent of men anyway, because other features carry it. From the published weights (an A3 reading), marital_status alone accounts for 59 percent of the gap in the model's scores between men and women. Removing the attribute did not remove the disparity.

A1: y_pred + group. A1 with labels: add y_true. A2: add y_prob. A3: every weight, enough to rebuild predict(); we checked that it reproduces every published probability.

COMPAS, the real deployed tool

The criminal-risk scores ProPublica analysed in "Machine Bias" (2016): real decile scores from a real product, with real two-year outcomes. Not re-hosted: the origin repository has no licence and names real defendants. get_compas.py downloads it from ProPublica, checks the file, drops every name, and adds y_pred. We ran it: 6,172 rows, false-positive rate 42.3 percent for African-American and 22.0 percent for Caucasian defendants, matching ProPublica's analysis.

A1 with labels. The decile is a rank, not a probability, so calibration claims need care.

Our own synthetic files (lending, hiring, healthcare, recruitment) are in the catalogue. We measured what each contains before writing a word about it; the descriptions say what is in the data, not what was intended.

import pandas as pd
from vfairness import FairnessAnalyzer

URL = "https://vfairness.validant.ai/test-objects/data/"
df = pd.read_csv(URL + "predictive/adult_test_with_predictions.csv")

# A1 with labels: decisions, true outcomes, groups
a1 = FairnessAnalyzer(df.y_true, df.y_pred, df.sex).get_report(include_ci=False)
# A2: the same, plus the model's probabilities
a2 = FairnessAnalyzer(df.y_true, df.y_pred, df.sex, y_prob=df.y_prob).get_report(include_ci=False)
print(sorted(set(a2["metrics"]) - set(a1["metrics"])))
# ['auroc_parity', 'calibration_difference', 'integrated_calibration_index', 'multicalibration']

No labels at all? Then you have decisions and groups only, and run_pulse works in its label-free mode (FairnessAnalyzer needs y_true):

from vfairness.operations.pulse import run_pulse

decisions_only = df[["sex", "race", "age", "y_pred"]]
result = run_pulse(decisions_only, {"source_kind": "predictive", "protected_attributes": ["sex", "race"]})
print(result["data"]["verdict"]["headline"])

Synthetic: data with a known answer

Real data never tells you whether your tool is right, because nobody knows the true bias in it. Synthetic data does: you plant one mechanism and check that the tool finds that one and nothing else. The pack has a clean control and seven single-bias variants of the same 500-applicant recruitment funnel, each with its answer key in _bias_flags (drop it, and true_qualification_score, before you analyse). They are regenerated byte for byte by scripts/synth_single_bias_datasets.py.

One variant is a deliberate trap. surname_proxy.csv has no race column: surnames carry the signal and drive the decision. Nothing can make a direct claim about race from this file, because race is not in it. That is a declared ceiling, and a tool that reports race as "fine" here is claiming something it could not see.

This is what happened when we ran every variant through run_pulse as a US employment screen, with all ten protected attributes declared:

FilePlantedPlanted mechanismConfirmed but not plantedTop-level verdict
synthetic/clean.csvnothingnothing to findage, national_origincritical
synthetic/gender_penalty.csva gap on gendermissednonecritical
synthetic/race_penalty.csva gap on race_ethnicityfoundnational_origincritical
synthetic/age_cliff.csva gap on agefoundnonecritical
synthetic/disability_penalty.csva gap on disability_statusfoundagecritical
synthetic/zip_redlining.csva proxy, zip_minority_majority (a race_ethnicity gap follows from it)foundnonecritical
synthetic/surname_proxy.csvnot checkable (race is not in the file)nothing to findnonecritical
synthetic/photo_laundering.csva proxy, photo_attractiveness_score (a race_ethnicity gap follows from it)foundnational_origin, primary_languagecritical
What the pack showed about vfairness itself. Graded on what Pulse statistically confirmed (significant after its own correction), it found five of the six mechanisms that can be found, and missed one: the gender penalty, which removes half of women's invitations in technical roles only and is about as large, at 500 applicants, as the gaps chance alone produces there. It also confirmed gaps where nothing was planted, including on the clean control: age (the youngest band really was invited less often in this random draw, raw p 0.025, but that should not survive a correction across ten attributes) and national origin (which fails an overall test across origins, p 0.31). And its top-level verdict was "critical" on all eight files, the clean one included, because in US employment the four-fifths screen sets the severity from the ratio, and a 27-person group falls below 0.80 easily. So on this pack the headline verdict could not tell planted from clean; the per-attribute confirmed results mostly could. That is the reason this pack exists: to measure the tool before trusting it, and to publish what it found. The false confirmations on the clean control are being investigated as a defect.

Tone alone is not detection. Graded on tone (any warn or critical), every file, the clean control included, would show six or more attributes "flagged", which says nothing.

Generative (LLM)

An LLM is tested by what it says. Change only the person in the prompt (a name, an age, a stated gender) and compare what comes back. The prompts can come from a benchmark; the answers can come from you running the model, or from someone who already did.

discrim-eval, answered by an open model

Anthropic's discrim-eval prompts (loans, hiring, medical priority, and more), each asked once for every combination of age, gender and race. We asked qwen2.5:7b through Ollama and kept both the answer (A1) and the probability of "yes" from the log probabilities (A2). The model digest is in every row, so the version is pinned (V1).

A row where neither "yes" nor "no" was among the 20 most likely tokens has an empty p_yes: could not measure, never 0.5.

Support bot replies

Prompts and replies from a synthetic support assistant, labelled with the customer's gender and ethnicity. A1 at its simplest: text only, no model needed.

What this model showed, and what A2 saw that A1 could not. On 24 decisions asked 135 ways each (3,240 answers), qwen2.5:7b said yes more often to some groups than others: Black versus white personas 3.7 points more often (95 percent interval 0.9 to 7.4, resampling whole decisions), women versus men 3.8 points, non-binary versus men 4.8 points. That is broadly the direction Anthropic reported for its own model in the discrim-eval paper. Then the case that shows why A2 matters. Hispanic personas got "yes" 1.4 points more often than white personas, an interval of −0.15 to 3.6 that includes no difference, so the answers alone cannot say. The probability of "yes" is 1.0 points higher, interval 0.1 to 2.4: small, only just resolved, but resolved. Same prompts, same file; the deeper tier saw what the shallower one could not. (Compare probabilities, not logits, here: 589 answers were so certain that no "no" appeared among the 20 most likely tokens, and a logit of those depends on where you clip.)

Where to get each tier: any chat endpoint is A1. Log probabilities (A2) come from OpenAI, vLLM, and Ollama from version 0.12.11; hosted routers such as Hugging Face Inference Providers and OpenRouter list the parameter, but whether it comes back depends on the model behind them, so check one response first. A3 means open weights you run yourself: GPT-2 is the cheapest full white box; Apertus 8B publishes its training data too, which makes it one of the few A4 subjects. vfairness's LLMApiProxy reads text only, so A2 analysis works the way the pack does: collect the probability yourself and analyse it as a score.

llm = pd.read_csv(URL + "llm/discrim_eval_qwen25_7b.csv")
measured = llm.dropna(subset=["p_yes"])            # A2: rows where p_yes could be read
print(len(llm) - len(measured), "rows could not be measured")
print(measured.groupby("race")["p_yes"].mean().round(3))
print(llm.groupby("race")["answer"].apply(lambda a: a.str.lower().str.startswith("yes").mean()).round(3))  # A1

Agent

An agent decides by acting: which tool it calls, whom it hands a case to, how many steps it takes. The test object is one row per episode with the group and the action, or the raw trace your framework already writes (OpenTelemetry, Langfuse, LangSmith). vfairness flattens raw traces into that table for you.

A real loan agent, counterfactual personas

A tool-calling agent on qwen2.5:7b with four tools (approve, ask for documents, refer to a human, decline). 30 fixed applications, each sent once per persona: twelve names that signal gender and origin (Swiss, Kosovar, Nigerian) plus an unnamed control. Every persona of the same application shares one sampling seed, so a difference is the name, not the dice; two repeats measure the noise floor. That makes it a controlled test (E2), not a log we happened to have.

Download the harness and point it at your own agent: it needs only an OpenAI-compatible endpoint.

What this agent showed: no name effect we could detect, and how small one could hide. Approvals were 84.2 percent for Swiss, 84.6 for Nigerian and 85.0 for Kosovar names; 83.9 for men and 85.3 for women (780 episodes). Asked twice with the same name, the agent changed its tool 14.2 percent of the time; with the same seed, a name changed it 5.3 percent of the time. The difference between names is smaller than the agent's own noise. This null has a limit: at 80 percent power and alpha 0.05, the test would have caught an approval gap of about 10 points between two origins, or 8 between genders. A smaller gap is not ruled out. That sentence is the result, not a footnote to it.

Benefits-triage traces, three formats

The same synthetic agent as a per-episode CSV, as raw OpenTelemetry GenAI spans, and as a Langfuse export. Use them to check that your own traces load before you run anything.

runs = pd.read_csv(URL + "agent/loan_agent_persona_runs.csv")
named = runs[runs.group != "control"]
result = run_pulse(named, {"source_kind": "agent", "protected_attributes": ["group"]})
agent = result["data"]["agent"]
print(agent["sampleAdequacy"]["perGroup"])   # episodes per group
print(agent["summary"])                       # what the probe measured

Print the probe's own summary, not only the top-level verdict. On this file the probe measures all six groups (120 episodes each) and finds no tool-selection bias, while the top-level verdict reads "Insufficient assessable data". That wording is wrong: it reports a measured null as a could-not-check. It is a known defect in how Pulse words a clean agent result, and the table below shows both readings until it is fixed.

Multi-agent

Several agents can each look fair and still be unfair together: one agent's note can steer the next. The test object needs each agent's output, not only the final one, so the question "is the system more biased than its parts?" can be answered.

A two-agent loan committee

A screener scores the risk; a reviewer decides twice, once alone and once after reading the screener's note. Same applications, personas and seeds as the agent above. The CSV has every agent's output per sample; the harness trace is the same run captured with MultiAgentRunHarness, ready for CompositionalityAnalyzer and EmergentBiasDetector. A reply the model did not return in a usable form stays empty and is counted, never filled in.

What this committee showed: a real gap at the first agent, carried to the end; amplification not established. The screener, working alone, scores names differently: average risk runs from 43.8 for Swiss women to 47.1 for Nigerian men on a 0 to 100 scale, a 3.4 point spread (permutation test keeping each application's personas together, p 0.0005). The reviewer working alone does not separate the groups clearly (2.9 points, p 0.06). After reading the screener's note, the reviewer's spread is 4.9 points (p 0.0005), so the screener's gap reaches the final decision. Whether the committee amplifies it is a separate question, and the answer is no evidence yet: the emergent-bias test gives amplification of 1.30 (women and men), 1.35 (Swiss and Nigerian names) and 0.65 (Swiss and Kosovar names), none significant. A small, measured gap and an unproven amplification: both statements are only possible because each agent's output was kept.
committee = pd.read_csv(URL + "multi_agent/loan_committee_runs.csv")
named = committee[committee.group != "control"]
shift = (named.reviewer_after_risk - named.reviewer_alone_risk).groupby(named.persona_origin).mean()
print(shift.round(2))   # how far the screener's note moves the reviewer, by origin

Notebooks

Three notebooks walk through the objects above with every result executed, not described:

  • notebooks/vfairness_10_test_objects_tour.ipynb: one source (Adult) up the ladder from A1 to A3, then the synthetic pack, planted against clean.
  • notebooks/vfairness_11_llm_testing_demo.ipynb: discrim-eval at A1 and A2, and a live counterfactual test against your own endpoint.
  • notebooks/vfairness_12_agent_testing_demo.ipynb: the loan agent with its noise floor, and the committee through the multi-agent analyzers.

What happened when we ran each file

Each object below was run through run_pulse by scripts/verify_test_objects.py. This is what the library returned, unedited. A verdict here describes the file, not a recommendation to use it.

Read the verdict as a triage screen, not as the size of the gap. Its line depends on the domain and jurisdiction you pass: in lending, severity needs a statistically confirmed gap with a meaningful effect size, and the four-fifths ratio (0.80) is shown as a guide. The lending file has real gaps (women approved 9.2 points less often, Black applicants 6.0 points less), but their ratios are 0.82 and 0.88, so it can read "within fairness budget". Read the measured description in the catalogue next to the tone, and decide your own threshold before you look.

Run on 2026-10-01.

FileRows usedVerdict returnedPer attribute
predictive/adult_test_with_predictions.csv15,060Not fit to deploy as-is: material disparity for "Female".sex (critical), race (critical)
predictive/adult_test_with_predictions.csv
A1 without labels
15,060Not fit to deploy as-is: material disparity for "Female".sex (critical), race (critical)
predictive/lending_fairness.csv1,000Within fairness budget on the assessed attributes.gender (pass), race (pass)
predictive/hiring_fairness.csv1,000Within fairness budget on the assessed attributes.gender (pass), ethnicity (pass)
predictive/healthcare_fairness.csv1,000Within fairness budget on the assessed attributes.sex (pass), race (pass)
predictive/recruitment_fairness_dataset.csv2,700Not fit to deploy as-is: material disparity for "Cognitive/Neurodivergent".gender (warn), race_ethnicity (critical), age (critical), disability_status (critical)
llm/generative_support_fairness.csv210Qualified opinion: deployable only with the documented remediations and active monitoring below; material but bounded fairness issues were found.none assessable
llm/discrim_eval_qwen25_7b.csv3,240Within fairness budget on the assessed attributes.gender (pass), race (pass)
agent/agent_traces_support.csv240Adverse opinion: material fairness defects make this system unfit to deploy as-is. Trajectory gap on 'steps': man averages 6.97 vs 5.10 for woman (steps per episode; permutation test, BH-adjusted p 0.002997, Cohen's d 1.32); Episodes for 'woman' lean on memory-carried context in 100% of cases vs 60% for 'nonbinary' (chi-square p 8.176e-09).
Agent probe: 13 agentic-bias finding(s) after omnibus gating and family-wise Benjamini-Hochberg correction.
none assessable
agent/agent_traces_otel.jsonl2,130Adverse opinion: material fairness defects make this system unfit to deploy as-is. Episodes for 'woman' lean on memory-carried context in 100% of cases vs 60% for 'nonbinary' (chi-square p 8.176e-09).
Agent probe: 12 agentic-bias finding(s) after omnibus gating and family-wise Benjamini-Hochberg correction.
none assessable
agent/agent_traces_langfuse.jsonl240Adverse opinion: material fairness defects make this system unfit to deploy as-is. Trajectory gap on 'steps': woman averages 6.37 vs 5.30 for nonbinary (steps per episode; permutation test, BH-adjusted p 0.002997, Cohen's d 0.85); Episodes for 'woman' lean on memory-carried context in 100% of cases vs 60% for 'nonbinary' (chi-square p 8.176e-09).
Agent probe: 10 agentic-bias finding(s) after omnibus gating and family-wise Benjamini-Hochberg correction.
none assessable
agent/loan_agent_persona_runs.csv720Insufficient assessable data to issue a fairness opinion. Provide grouped protected attributes and, ideally, outcome labels.
Agent probe: No significant tool/action selection bias detected across groups (omnibus-gated, family-wise corrected). 6 per-tool test(s) could not fire at any data (too few uses of that tool to reach p<0.05) and are reported as not assessed, not as absence of bias.
none assessable
agent/loan_agent_otel.jsonl1,550Insufficient assessable data to issue a fairness opinion. Provide grouped protected attributes and, ideally, outcome labels.
Agent probe: No significant tool/action selection bias detected across groups (omnibus-gated, family-wise corrected). 7 per-tool test(s) could not fire at any data (too few uses of that tool to reach p<0.05) and are reported as not assessed, not as absence of bias.
none assessable

Catalogue of downloads

Every file is listed in manifest.json with its sha256, size, source and licence. Check the hash after you download.

FilePathwayTierWhat is in itSource and licenceSize and sha256
predictive/lending_fairness.csv
Synthetic credit decisions with model scores
PredictiveA2Measured: Black applicants approved 6.0 points less often than White (42.0 vs 48.0 percent), Hispanic 4.5 points; women 9.2 points less than men (41.0 vs 50.2).Validant (synthetic)
CC BY 4.0 (Validant)
1,000 rows
57 KB
6e74356aa7b6
predictive/hiring_fairness.csv
Synthetic hiring decisions with model scores, includes non-binary
PredictiveA2Measured: small gaps only. At equal interview and technical scores women are selected 2.0 points less often (not significant); non-binary n = 50. Useful to check that a tool does not over-flag.Validant (synthetic)
CC BY 4.0 (Validant)
1,000 rows
52 KB
6f4a42b61c0e
predictive/healthcare_fairness.csv
Synthetic clinical risk scores
PredictiveA2Measured: at similar clinical profiles Black patients get risk scores 4.2 points lower. Probabilities understate risk for every group by 9 to 17 points, so the miscalibration is global, not group-specific.Validant (synthetic)
CC BY 4.0 (Validant)
1,000 rows
58 KB
32fd1f4a032d
predictive/recruitment_fairness_dataset.csv
Synthetic recruitment funnel, 9 planted biases, ground truth included
PredictiveA1Nine mechanisms, flagged per row in _bias_flags (B1 to B9). All personal details are invented.Validant (synthetic)
CC BY 4.0 (Validant)
2,700 rows
1.3 MB
446e1dff99ca
predictive/adult_test_with_predictions.csv
UCI Adult test split with a pinned logistic regression's predictions
PredictiveA2Nothing planted: a real 1994 census extract. The model never sees sex or race.UCI Adult, Becker and Kohavi (1996), doi:10.24432/C5XW20; predictions by Validant
CC BY 4.0
15,060 rows
1.8 MB
304a3c8ee6fc
predictive/adult_logreg_model.json
The same model, every weight: rebuild predict() in ten lines
PredictiveA3n/aValidant, trained on UCI Adult
CC BY 4.0

4 KB
6ebd59bf8e45
scripts/get_compas.py
Download COMPAS from ProPublica and add prediction columns
PredictiveA2Nothing planted: the real Northpointe decile scores.ProPublica, compas-analysis
MIT (script); data is not re-hosted

2 KB
2b98e40af859
synthetic/clean.csv
Control: no bias planted
SyntheticA1None. A correct detector reports nothing here.Validant (generated)
CC BY 4.0 (Validant)
500 rows
118 KB
a1bd6f577d67
synthetic/gender_penalty.csv
Gender penalty in tech roles
SyntheticA1Half of female invites in tech roles removed (B1).Validant (generated)
CC BY 4.0 (Validant)
500 rows
119 KB
bc0a91743c0a
synthetic/race_penalty.csv
Race penalty
SyntheticA1Half of Black and Hispanic invites removed (B2).Validant (generated)
CC BY 4.0 (Validant)
500 rows
119 KB
6aebab36bfe1
synthetic/age_cliff.csv
Age cliff
SyntheticA1Applicants under 25 never invited (B3).Validant (generated)
CC BY 4.0 (Validant)
500 rows
118 KB
d34b7859d7f6
synthetic/disability_penalty.csv
Disability penalty
SyntheticA180 percent of invites removed for any reported disability (B7).Validant (generated)
CC BY 4.0 (Validant)
500 rows
119 KB
94b95cd77859
synthetic/zip_redlining.csv
ZIP code redlining (a proxy)
SyntheticA1Minority-majority ZIP invites halved; race itself untouched (B4).Validant (generated)
CC BY 4.0 (Validant)
500 rows
119 KB
b3ef176c8707
synthetic/surname_proxy.csv
Surname proxy, race column removed
SyntheticA1Surnames typical of Black and Hispanic applicants lose half their invites; race_ethnicity is not in the file (B5).Validant (generated)
CC BY 4.0 (Validant)
500 rows
115 KB
7f9580a9e311
synthetic/photo_laundering.csv
Photo score laundering
SyntheticA1A photo score correlated with race drives half the decision (B9).Validant (generated)
CC BY 4.0 (Validant)
500 rows
119 KB
a5205d150dbc
llm/generative_support_fairness.csv
Support-bot prompts and replies by customer gender and ethnicity
Generative (LLM)A1Tone and refusal differ by group in the replies.Validant (synthetic)
CC BY 4.0 (Validant)
210 rows
40 KB
e3e86121d222
llm/discrim_eval_qwen25_7b.csv
Real yes/no decisions with log probabilities from a pinned open model
Generative (LLM)A2Nothing planted: measured behaviour of the model.Prompts: Anthropic discrim-eval (Tamkin et al. 2023); answers and logprobs: Validant run of qwen2.5:7b
CC BY 4.0
3,240 rows
445 KB
cae9945e172f
agent/agent_traces_support.csv
Benefits-triage agent, one row per episode
AgentA1Tool choice and routing differ by gender.Validant (synthetic)
CC BY 4.0 (Validant)
240 rows
38 KB
6ecd2d323ffc
agent/agent_traces_otel.jsonl
The same agent as raw OpenTelemetry GenAI spans
AgentA1As above.Validant (synthetic)
CC BY 4.0 (Validant)
2,130 rows
498 KB
5c7dd211f12c
agent/agent_traces_langfuse.jsonl
The same agent as a Langfuse export
AgentA1As above.Validant (synthetic)
CC BY 4.0 (Validant)
240 rows
255 KB
d3bf2ef9f64a
agent/loan_agent_persona_runs.csv
A real tool-calling loan agent, counterfactual personas, one row per episode
AgentA1Nothing planted: measured behaviour of the agent.Validant run of qwen2.5:7b
CC BY 4.0 (Validant)
780 rows
138 KB
bf3ae4f22d98
agent/loan_agent_otel.jsonl
The same runs as OpenTelemetry GenAI spans
AgentA1Nothing planted.Validant run of qwen2.5:7b
CC BY 4.0 (Validant)
1,550 rows
416 KB
3ab291bd8c5e
scripts/loan_agent_persona_harness.py
The harness that produced the runs: point it at your own agent
AgentA1n/aValidant
MIT

12 KB
1f8593708553
multi_agent/loan_committee_runs.csv
Two-agent loan committee: screener, then reviewer, per-agent outputs
Multi-agentA1Nothing planted: measured behaviour of the system.Validant run of qwen2.5:7b
CC BY 4.0 (Validant)
780 rows
118 KB
57e1d55e27e4
multi_agent/loan_committee_harness_trace.json
The same run as a MultiAgentRunHarness trace
Multi-agentA1Nothing planted.Validant run of qwen2.5:7b
CC BY 4.0 (Validant)

49 KB
81c55b698ef5

Classic public datasets, at their origin

Where a well-known dataset can be redistributed, we host a ready-to-run version above and link the origin here. Where it cannot, we link the origin only. Always cite the origin.

DatasetPathwayReal model outputs?Tier you can reachLicence and notes
COMPAS (ProPublica)PredictiveYes, a deployed tool's scoresA1 with labelsNo licence file; names real people. Use our script.
Adult (UCI)PredictiveNo, labels onlyA1 to A4 once you train a modelCC BY 4.0. Hosted above with predictions. Known problems: 1994 data, an arbitrary $50K threshold; see Retiring Adult.
folktables (US Census ACS)PredictiveNoA1 to A4 once you trainMIT (code). The modern replacement for Adult: any US state, any year, several tasks.
German Credit (UCI)PredictiveNoA1 to A4 once you trainCC BY 4.0. Only 1,000 rows; sex must be derived from a combined field with known coding problems.
Default of Credit Card Clients (UCI)PredictiveNoA1 to A4 once you trainCC BY 4.0. 30,000 rows; sex, age, marriage, education.
Bank Marketing (UCI)PredictiveNoA1 to A4 once you trainCC BY 4.0. Age and marital status only.
Diabetes 130-US Hospitals (UCI)PredictiveNoA1 to A4 once you trainCC BY 4.0. About 100,000 encounters; race, gender, age.
Fairness dataset survey (Law School, Dutch census, and cleaned copies of the above)PredictiveNoA1 to A4 once you trainThe repository has no licence; each dataset keeps its original terms.
HMDA (US mortgage disclosure)PredictiveA lender's decision, not one model'sA1US public data. Millions of rows per year. Treat the decision as an organisation's outcome, not a model output.
Boston HousingDo not use. It encodes a race-derived variable and was removed from scikit-learn 1.2; the link explains why.
synthcitySyntheticn/aAllApache-2.0. Includes DECAF, a fairness-aware causal generator.
ml-fairness-gymSyntheticn/aAllApache-2.0. Feedback loops over time (lending, attention allocation).
SDVSyntheticn/aAllBusiness Source License, not open source: check the terms before commercial use.
discrim-eval (Anthropic)GenerativeNo, prompts onlyA1; A2 with log probabilitiesCC BY 4.0. Answered by an open model in our pack above.
BBQGenerativeNoA1; A2 with option probabilitiesCC BY 4.0. Nine bias categories; a subset ships with vfairness (BenchmarkRunner.run_bbq).
BOLDGenerativeNoA1CC BY 4.0. Open-ended continuations scored for sentiment and toxicity; a subset ships with vfairness.
HolisticBiasGenerativeNoA1CC BY-SA 4.0. About 600 identity descriptors; a subset ships with vfairness.
CrowS-Pairs, StereoSetGenerativeNoA2 or A3 (they compare sentence likelihoods)CC BY-SA 4.0. Not usable at A1.
Bias in BiosGenerative / text classifierNoA1 to A3MIT. Biographies with gender and profession labels; good for probing a model's own representations.
HELM raw resultsGenerativeYes, many models' outputsA1 with gold answers, no model run neededPer-instance outputs, including BBQ, for many models.
τ²-bench / τ-benchAgentNoA1 with tracesMIT. Customer-service tasks with simulated users: vary the user's persona and compare success.
AgentFairBenchAgent, multi-agentNoA1Apache-2.0. Built for agent fairness (hiring, lending, triage). New in 2026 and not yet replicated: treat its results as preliminary.
SotopiaMulti-agentYes, recorded episodesA1, no model run neededSocial simulations between agents with persona profiles and goal scores.
MASTMulti-agentYes, 1,642 tracesA1CC BY 4.0. Seven frameworks, labelled failure modes. No demographic axis: use it to test trace loading, not fairness.

Trace formats your agent framework may already write: OpenTelemetry GenAI semantic conventions, OpenInference, Langfuse and LangSmith exports.

How these files were made

  • scripts/build_test_objects.py builds the synthetic and Adult packs and writes the manifest. It refuses to build if the UCI origin file has changed.
  • scripts/build_live_test_objects.py builds the LLM, agent and multi-agent packs against a local model and records its digest. discrim-eval is fetched at a pinned revision and checked by sha256.
  • scripts/verify_test_objects.py runs every file through vfairness and writes the results shown above.

Names in the agent files are a standard correspondence-testing cue (Bertrand and Mullainathan, 2004). They signal a group to the model and say nothing about any real person. All personal details in the synthetic files are invented.