Find a test object
Every fairness analysis needs something to test: a table of decisions, a model, a chat endpoint, an agent, a team of agents. This page tells you which object you need for the depth of analysis you want, and gives you a working one for every pathway, ready to download. Every file below was run through vfairness before it was published, and the page shows what happened.
Start here: what are you testing, and what can you see?
Pick the two answers. The result is your access tier, the strongest claim that tier supports, and a file to start with.
- Access tier
- Strongest claim
- You can run
- You cannot claim
- Start with
What an assessment is, in 90 seconds
A fairness result is only as good as two things that must travel with it. Validant's article Precision Is Not Proof calls them the pointing and the seeing.
The pointing: where did we look?
Three dials. Body: which offering. Pathway: what kind of AI system (predictive, generative, agentic, multi-agent). Audience: who the report is for. The audience changes the wording, never the numbers. This page is about the pathway dial.
The seeing: how well could we see?
Access (A0 to A4): what the system lets you observe.
Evidence (E0 to E3): how the test was designed. Validity (V0 to V3): whether
the version you tested is pinned and still current. Written together, for example A2 · E2 · V1.
Three rules follow, and every object on this page is labelled with them in mind:
- The weakest of the three sets the ceiling. Deep access does not rescue a stale, unpinned model, and a careful method does not rescue outputs you were only told about. In the article's words: taken at your word is Indicative; decisions measured is Limited; plus confidence scores is Reasonable.
- A null result must say what it could have caught. "No disparity detected" is a finding only with its limiting magnitude: the smallest gap the test would have found, at a stated power.
- Not seen is not the same as seen and fine. If your tier cannot reach a question, the honest answer is "could not check", never a pass.
"Pointing says where she looked. Seeing says what could possibly have been resolved from there."
The access ladder
The tier is set by what the system gives back, not by how you connect to it. A chat window, a company's wrapper API and the provider's own API all return text, so all three are A1. Tiers and their ceilings, quoted from the article:
| Tier | What you have | Highest claim it supports |
|---|---|---|
| A0 Attested | Vendor documentation, model card, published evaluations. No probing. | The subject's own claims, recorded and checked for internal consistency. |
| A1 Behavioural | Query access. Terminal output only: text, label, decision. | Disparity in observed outcomes, present or absent at a stated sensitivity. |
| A2 Scored | A1 plus per-output scores: log probabilities, class probabilities, ranked alternatives, confidence. | Calibration and threshold behaviour by group. Ranking and margin disparity. |
| A3 Internal | Weights, activations, gradients. The ability to intervene on the computation. | Which internal structures carry the disparity. Mechanistic attribution. |
| A4 Provenance | A3 plus training-corpus lineage, fine-tuning and preference-data history, evaluation history. | Where in the lifecycle the disparity came from. |
Fairness is an input-output discipline and does well at A1: group metrics need decisions, labels and group membership, never weights. Explainability is different: a claim about why the model behaves as it does starts at A3. Agent traces (which tool was called, where a case was handed off) are behaviour you watch, so they are A1 too: they widen what you see at A1 without reaching inside the model.
What data you need for what depth
One row per pathway. Each cell names the fields you need and what they unlock.
| Pathway | A1 outputs | A2 plus scores | A3 internals | A4 provenance |
|---|---|---|---|---|
| Before the model the data alone |
Features, protected attributes and the historical outcome, no model. You can check representation, label disparity in history and proxies (a column that stands in for a protected one). This is a data audit, not a model assessment: it says nothing about how a model will decide. | |||
| Predictive tabular |
y_pred + group. Selection-rate disparity, demographic parity. Add y_true and you also get
error-rate parity: equalised odds, equal opportunity, predictive parity. |
+ y_prob. Calibration by group, AUROC parity, multicalibration, threshold trade-offs. |
+ the model (file, weights or a predict() you can call). Feature attribution, counterfactual
what-if, retraining with constraints. |
+ the training data and how it was built. Where the disparity entered. |
| Generative LLM |
Prompts in, text out. Counterfactual prompt pairs, benchmarks (BBQ, BOLD, HolisticBias), refusal, tone and sentiment disparity. | + token log probabilities. Likelihood bias: the probability of "yes", of each answer option, which resolves differences a yes or no hides. | + open weights. The model's own embeddings (WEAT, SEAT), hidden-state probes. | + open training data. Which corpus patterns carry the bias. |
| Agent | Final outcome + persona. Outcome disparity, correspondence tests. Add traces (tool calls, routes, steps) and you also get tool-selection, delegation and trajectory bias. | + log probabilities at each step. The margin behind each tool choice. | + an open-weight model inside the agent. | Rarely available. |
| Multi-agent | End-to-end inputs and outputs. Add per-agent outputs and you also get compositionality, emergent bias and groupthink: is the system more biased than its parts? | As for agents. | As for agents. | Rarely available. |
| Synthetic | All tiers at once, because you made it. Its unique value is the answer key: you know which bias was planted, so you can check that a tool finds it and stays quiet on the clean control. | |||
Predictive (tabular)
A predictive system turns a row of facts about a person into a decision or a score: approve a loan, invite to interview, flag for review. The test object is a table with one row per person: the decision, the group, and as much else as you can get. The same source can be walked up the ladder, which is what the Adult pack is for.
UCI Adult, with a model's decisions
The 1994 US census extract every fairness course uses, with a logistic regression's predictions added
(15,060 test rows). The model never sees sex or race; it selects 8.1 percent of women and
26.0 percent of men anyway, because other features carry it. From the published weights (an A3 reading),
marital_status alone accounts for 59 percent of the gap in the model's scores between men and women.
Removing the attribute did not remove the disparity.
A1: y_pred + group. A1 with labels: add y_true. A2: add y_prob.
A3: every weight, enough to rebuild predict(); we checked
that it reproduces every published probability.
COMPAS, the real deployed tool
The criminal-risk scores ProPublica analysed in "Machine Bias" (2016): real decile scores from a real product,
with real two-year outcomes. Not re-hosted: the origin repository has no licence and names real defendants.
get_compas.py downloads it from ProPublica, checks the file, drops every name,
and adds y_pred. We ran it: 6,172 rows, false-positive rate 42.3 percent for African-American and 22.0
percent for Caucasian defendants, matching ProPublica's analysis.
A1 with labels. The decile is a rank, not a probability, so calibration claims need care.
Our own synthetic files (lending, hiring, healthcare, recruitment) are in the catalogue. We measured what each contains before writing a word about it; the descriptions say what is in the data, not what was intended.
import pandas as pd
from vfairness import FairnessAnalyzer
URL = "https://vfairness.validant.ai/test-objects/data/"
df = pd.read_csv(URL + "predictive/adult_test_with_predictions.csv")
# A1 with labels: decisions, true outcomes, groups
a1 = FairnessAnalyzer(df.y_true, df.y_pred, df.sex).get_report(include_ci=False)
# A2: the same, plus the model's probabilities
a2 = FairnessAnalyzer(df.y_true, df.y_pred, df.sex, y_prob=df.y_prob).get_report(include_ci=False)
print(sorted(set(a2["metrics"]) - set(a1["metrics"])))
# ['auroc_parity', 'calibration_difference', 'integrated_calibration_index', 'multicalibration']
No labels at all? Then you have decisions and groups only, and run_pulse works in its label-free mode
(FairnessAnalyzer needs y_true):
from vfairness.operations.pulse import run_pulse
decisions_only = df[["sex", "race", "age", "y_pred"]]
result = run_pulse(decisions_only, {"source_kind": "predictive", "protected_attributes": ["sex", "race"]})
print(result["data"]["verdict"]["headline"])
Synthetic: data with a known answer
Real data never tells you whether your tool is right, because nobody knows the true bias in it. Synthetic data
does: you plant one mechanism and check that the tool finds that one and nothing else. The pack has a clean control
and seven single-bias variants of the same 500-applicant recruitment funnel, each with its answer key in
_bias_flags (drop it, and true_qualification_score, before you analyse). They are regenerated
byte for byte by scripts/synth_single_bias_datasets.py.
One variant is a deliberate trap. surname_proxy.csv has no race column: surnames carry the signal and
drive the decision. Nothing can make a direct claim about race from this file, because race is not in it. That is a
declared ceiling, and a tool that reports race as "fine" here is claiming something it could not see.
This is what happened when we ran every variant through run_pulse as a US employment screen, with all ten protected attributes declared:
| File | Planted | Planted mechanism | Confirmed but not planted | Top-level verdict |
|---|---|---|---|---|
| synthetic/clean.csv | nothing | nothing to find | age, national_origin | critical |
| synthetic/gender_penalty.csv | a gap on gender | missed | none | critical |
| synthetic/race_penalty.csv | a gap on race_ethnicity | found | national_origin | critical |
| synthetic/age_cliff.csv | a gap on age | found | none | critical |
| synthetic/disability_penalty.csv | a gap on disability_status | found | age | critical |
| synthetic/zip_redlining.csv | a proxy, zip_minority_majority (a race_ethnicity gap follows from it) | found | none | critical |
| synthetic/surname_proxy.csv | not checkable (race is not in the file) | nothing to find | none | critical |
| synthetic/photo_laundering.csv | a proxy, photo_attractiveness_score (a race_ethnicity gap follows from it) | found | national_origin, primary_language | critical |
Tone alone is not detection. Graded on tone (any warn or critical), every file, the clean control included, would show six or more attributes "flagged", which says nothing.
Generative (LLM)
An LLM is tested by what it says. Change only the person in the prompt (a name, an age, a stated gender) and compare what comes back. The prompts can come from a benchmark; the answers can come from you running the model, or from someone who already did.
discrim-eval, answered by an open model
Anthropic's discrim-eval
prompts (loans, hiring, medical priority, and more), each asked once for every combination of age, gender and race.
We asked qwen2.5:7b through Ollama and kept both the answer (A1) and the probability of "yes" from
the log probabilities (A2). The model digest is in every row, so the version is pinned (V1).
A row where neither "yes" nor "no" was among the 20 most likely tokens has an empty p_yes:
could not measure, never 0.5.
Support bot replies
Prompts and replies from a synthetic support assistant, labelled with the customer's gender and ethnicity. A1 at its simplest: text only, no model needed.
qwen2.5:7b said yes more often to some groups than others: Black versus white
personas 3.7 points more often (95 percent interval 0.9 to 7.4, resampling whole decisions), women versus men 3.8 points,
non-binary versus men 4.8 points. That is broadly the direction Anthropic reported for its own model in the discrim-eval paper. Then the case that
shows why A2 matters. Hispanic personas got "yes" 1.4 points more often than white personas, an interval of
−0.15 to 3.6 that includes no difference, so the answers alone cannot say. The probability of "yes" is 1.0
points higher, interval 0.1 to 2.4: small, only just resolved, but resolved. Same prompts, same file; the deeper tier
saw what the shallower one could not. (Compare probabilities, not logits, here: 589 answers were so certain that no
"no" appeared among the 20 most likely tokens, and a logit of those depends on where you clip.)Where to get each tier: any chat endpoint is A1. Log probabilities (A2) come from OpenAI,
vLLM, and
Ollama from version 0.12.11;
hosted routers such as Hugging Face Inference Providers
and OpenRouter list the parameter,
but whether it comes back depends on the model behind them, so check one response first. A3 means open weights you
run yourself: GPT-2 is the
cheapest full white box; Apertus 8B
publishes its training data too, which makes it one of the few A4 subjects. vfairness's LLMApiProxy
reads text only, so A2 analysis works the way the pack does: collect the probability yourself and analyse it as a score.
llm = pd.read_csv(URL + "llm/discrim_eval_qwen25_7b.csv")
measured = llm.dropna(subset=["p_yes"]) # A2: rows where p_yes could be read
print(len(llm) - len(measured), "rows could not be measured")
print(measured.groupby("race")["p_yes"].mean().round(3))
print(llm.groupby("race")["answer"].apply(lambda a: a.str.lower().str.startswith("yes").mean()).round(3)) # A1
Agent
An agent decides by acting: which tool it calls, whom it hands a case to, how many steps it takes. The test object is one row per episode with the group and the action, or the raw trace your framework already writes (OpenTelemetry, Langfuse, LangSmith). vfairness flattens raw traces into that table for you.
A real loan agent, counterfactual personas
A tool-calling agent on qwen2.5:7b with four tools (approve, ask for documents, refer to a human,
decline). 30 fixed applications, each sent once per persona: twelve names that signal gender and origin (Swiss,
Kosovar, Nigerian) plus an unnamed control. Every persona of the same application shares one sampling seed, so a
difference is the name, not the dice; two repeats measure the noise floor. That makes it a controlled test (E2),
not a log we happened to have.
Download the harness and point it at your own agent: it needs only an OpenAI-compatible endpoint.
Benefits-triage traces, three formats
The same synthetic agent as a per-episode CSV, as raw OpenTelemetry GenAI spans, and as a Langfuse export. Use them to check that your own traces load before you run anything.
runs = pd.read_csv(URL + "agent/loan_agent_persona_runs.csv")
named = runs[runs.group != "control"]
result = run_pulse(named, {"source_kind": "agent", "protected_attributes": ["group"]})
agent = result["data"]["agent"]
print(agent["sampleAdequacy"]["perGroup"]) # episodes per group
print(agent["summary"]) # what the probe measured
Print the probe's own summary, not only the top-level verdict. On this file the probe measures all six groups (120 episodes each) and finds no tool-selection bias, while the top-level verdict reads "Insufficient assessable data". That wording is wrong: it reports a measured null as a could-not-check. It is a known defect in how Pulse words a clean agent result, and the table below shows both readings until it is fixed.
Multi-agent
Several agents can each look fair and still be unfair together: one agent's note can steer the next. The test object needs each agent's output, not only the final one, so the question "is the system more biased than its parts?" can be answered.
A two-agent loan committee
A screener scores the risk; a reviewer decides twice, once alone and once after reading the screener's note.
Same applications, personas and seeds as the agent above. The CSV has every agent's output per sample; the
harness trace is the same run captured with
MultiAgentRunHarness, ready for CompositionalityAnalyzer and EmergentBiasDetector.
A reply the model did not return in a usable form stays empty and is counted, never filled in.
committee = pd.read_csv(URL + "multi_agent/loan_committee_runs.csv")
named = committee[committee.group != "control"]
shift = (named.reviewer_after_risk - named.reviewer_alone_risk).groupby(named.persona_origin).mean()
print(shift.round(2)) # how far the screener's note moves the reviewer, by origin
Notebooks
Three notebooks walk through the objects above with every result executed, not described:
notebooks/vfairness_10_test_objects_tour.ipynb: one source (Adult) up the ladder from A1 to A3, then the synthetic pack, planted against clean.notebooks/vfairness_11_llm_testing_demo.ipynb: discrim-eval at A1 and A2, and a live counterfactual test against your own endpoint.notebooks/vfairness_12_agent_testing_demo.ipynb: the loan agent with its noise floor, and the committee through the multi-agent analyzers.
What happened when we ran each file
Each object below was run through run_pulse by scripts/verify_test_objects.py. This is
what the library returned, unedited. A verdict here describes the file, not a recommendation to use it.
Read the verdict as a triage screen, not as the size of the gap. Its line depends on the domain and jurisdiction you pass: in lending, severity needs a statistically confirmed gap with a meaningful effect size, and the four-fifths ratio (0.80) is shown as a guide. The lending file has real gaps (women approved 9.2 points less often, Black applicants 6.0 points less), but their ratios are 0.82 and 0.88, so it can read "within fairness budget". Read the measured description in the catalogue next to the tone, and decide your own threshold before you look.
Run on 2026-10-01.
| File | Rows used | Verdict returned | Per attribute |
|---|---|---|---|
| predictive/adult_test_with_predictions.csv | 15,060 | Not fit to deploy as-is: material disparity for "Female". | sex (critical), race (critical) |
| predictive/adult_test_with_predictions.csv A1 without labels | 15,060 | Not fit to deploy as-is: material disparity for "Female". | sex (critical), race (critical) |
| predictive/lending_fairness.csv | 1,000 | Within fairness budget on the assessed attributes. | gender (pass), race (pass) |
| predictive/hiring_fairness.csv | 1,000 | Within fairness budget on the assessed attributes. | gender (pass), ethnicity (pass) |
| predictive/healthcare_fairness.csv | 1,000 | Within fairness budget on the assessed attributes. | sex (pass), race (pass) |
| predictive/recruitment_fairness_dataset.csv | 2,700 | Not fit to deploy as-is: material disparity for "Cognitive/Neurodivergent". | gender (warn), race_ethnicity (critical), age (critical), disability_status (critical) |
| llm/generative_support_fairness.csv | 210 | Qualified opinion: deployable only with the documented remediations and active monitoring below; material but bounded fairness issues were found. | none assessable |
| llm/discrim_eval_qwen25_7b.csv | 3,240 | Within fairness budget on the assessed attributes. | gender (pass), race (pass) |
| agent/agent_traces_support.csv | 240 | Adverse opinion: material fairness defects make this system unfit to deploy as-is. Trajectory gap on 'steps': man averages 6.97 vs 5.10 for woman (steps per episode; permutation test, BH-adjusted p 0.002997, Cohen's d 1.32); Episodes for 'woman' lean on memory-carried context in 100% of cases vs 60% for 'nonbinary' (chi-square p 8.176e-09). Agent probe: 13 agentic-bias finding(s) after omnibus gating and family-wise Benjamini-Hochberg correction. | none assessable |
| agent/agent_traces_otel.jsonl | 2,130 | Adverse opinion: material fairness defects make this system unfit to deploy as-is. Episodes for 'woman' lean on memory-carried context in 100% of cases vs 60% for 'nonbinary' (chi-square p 8.176e-09). Agent probe: 12 agentic-bias finding(s) after omnibus gating and family-wise Benjamini-Hochberg correction. | none assessable |
| agent/agent_traces_langfuse.jsonl | 240 | Adverse opinion: material fairness defects make this system unfit to deploy as-is. Trajectory gap on 'steps': woman averages 6.37 vs 5.30 for nonbinary (steps per episode; permutation test, BH-adjusted p 0.002997, Cohen's d 0.85); Episodes for 'woman' lean on memory-carried context in 100% of cases vs 60% for 'nonbinary' (chi-square p 8.176e-09). Agent probe: 10 agentic-bias finding(s) after omnibus gating and family-wise Benjamini-Hochberg correction. | none assessable |
| agent/loan_agent_persona_runs.csv | 720 | Insufficient assessable data to issue a fairness opinion. Provide grouped protected attributes and, ideally, outcome labels. Agent probe: No significant tool/action selection bias detected across groups (omnibus-gated, family-wise corrected). 6 per-tool test(s) could not fire at any data (too few uses of that tool to reach p<0.05) and are reported as not assessed, not as absence of bias. | none assessable |
| agent/loan_agent_otel.jsonl | 1,550 | Insufficient assessable data to issue a fairness opinion. Provide grouped protected attributes and, ideally, outcome labels. Agent probe: No significant tool/action selection bias detected across groups (omnibus-gated, family-wise corrected). 7 per-tool test(s) could not fire at any data (too few uses of that tool to reach p<0.05) and are reported as not assessed, not as absence of bias. | none assessable |
Catalogue of downloads
Every file is listed in manifest.json with its sha256, size, source and licence. Check the hash after you download.
| File | Pathway | Tier | What is in it | Source and licence | Size and sha256 |
|---|---|---|---|---|---|
| predictive/lending_fairness.csv Synthetic credit decisions with model scores | Predictive | A2 | Measured: Black applicants approved 6.0 points less often than White (42.0 vs 48.0 percent), Hispanic 4.5 points; women 9.2 points less than men (41.0 vs 50.2). | Validant (synthetic) CC BY 4.0 (Validant) | 1,000 rows 57 KB 6e74356aa7b6 |
| predictive/hiring_fairness.csv Synthetic hiring decisions with model scores, includes non-binary | Predictive | A2 | Measured: small gaps only. At equal interview and technical scores women are selected 2.0 points less often (not significant); non-binary n = 50. Useful to check that a tool does not over-flag. | Validant (synthetic) CC BY 4.0 (Validant) | 1,000 rows 52 KB 6f4a42b61c0e |
| predictive/healthcare_fairness.csv Synthetic clinical risk scores | Predictive | A2 | Measured: at similar clinical profiles Black patients get risk scores 4.2 points lower. Probabilities understate risk for every group by 9 to 17 points, so the miscalibration is global, not group-specific. | Validant (synthetic) CC BY 4.0 (Validant) | 1,000 rows 58 KB 32fd1f4a032d |
| predictive/recruitment_fairness_dataset.csv Synthetic recruitment funnel, 9 planted biases, ground truth included | Predictive | A1 | Nine mechanisms, flagged per row in _bias_flags (B1 to B9). All personal details are invented. | Validant (synthetic) CC BY 4.0 (Validant) | 2,700 rows 1.3 MB 446e1dff99ca |
| predictive/adult_test_with_predictions.csv UCI Adult test split with a pinned logistic regression's predictions | Predictive | A2 | Nothing planted: a real 1994 census extract. The model never sees sex or race. | UCI Adult, Becker and Kohavi (1996), doi:10.24432/C5XW20; predictions by Validant CC BY 4.0 | 15,060 rows 1.8 MB 304a3c8ee6fc |
| predictive/adult_logreg_model.json The same model, every weight: rebuild predict() in ten lines | Predictive | A3 | n/a | Validant, trained on UCI Adult CC BY 4.0 | 4 KB 6ebd59bf8e45 |
| scripts/get_compas.py Download COMPAS from ProPublica and add prediction columns | Predictive | A2 | Nothing planted: the real Northpointe decile scores. | ProPublica, compas-analysis MIT (script); data is not re-hosted | 2 KB 2b98e40af859 |
| synthetic/clean.csv Control: no bias planted | Synthetic | A1 | None. A correct detector reports nothing here. | Validant (generated) CC BY 4.0 (Validant) | 500 rows 118 KB a1bd6f577d67 |
| synthetic/gender_penalty.csv Gender penalty in tech roles | Synthetic | A1 | Half of female invites in tech roles removed (B1). | Validant (generated) CC BY 4.0 (Validant) | 500 rows 119 KB bc0a91743c0a |
| synthetic/race_penalty.csv Race penalty | Synthetic | A1 | Half of Black and Hispanic invites removed (B2). | Validant (generated) CC BY 4.0 (Validant) | 500 rows 119 KB 6aebab36bfe1 |
| synthetic/age_cliff.csv Age cliff | Synthetic | A1 | Applicants under 25 never invited (B3). | Validant (generated) CC BY 4.0 (Validant) | 500 rows 118 KB d34b7859d7f6 |
| synthetic/disability_penalty.csv Disability penalty | Synthetic | A1 | 80 percent of invites removed for any reported disability (B7). | Validant (generated) CC BY 4.0 (Validant) | 500 rows 119 KB 94b95cd77859 |
| synthetic/zip_redlining.csv ZIP code redlining (a proxy) | Synthetic | A1 | Minority-majority ZIP invites halved; race itself untouched (B4). | Validant (generated) CC BY 4.0 (Validant) | 500 rows 119 KB b3ef176c8707 |
| synthetic/surname_proxy.csv Surname proxy, race column removed | Synthetic | A1 | Surnames typical of Black and Hispanic applicants lose half their invites; race_ethnicity is not in the file (B5). | Validant (generated) CC BY 4.0 (Validant) | 500 rows 115 KB 7f9580a9e311 |
| synthetic/photo_laundering.csv Photo score laundering | Synthetic | A1 | A photo score correlated with race drives half the decision (B9). | Validant (generated) CC BY 4.0 (Validant) | 500 rows 119 KB a5205d150dbc |
| llm/generative_support_fairness.csv Support-bot prompts and replies by customer gender and ethnicity | Generative (LLM) | A1 | Tone and refusal differ by group in the replies. | Validant (synthetic) CC BY 4.0 (Validant) | 210 rows 40 KB e3e86121d222 |
| llm/discrim_eval_qwen25_7b.csv Real yes/no decisions with log probabilities from a pinned open model | Generative (LLM) | A2 | Nothing planted: measured behaviour of the model. | Prompts: Anthropic discrim-eval (Tamkin et al. 2023); answers and logprobs: Validant run of qwen2.5:7b CC BY 4.0 | 3,240 rows 445 KB cae9945e172f |
| agent/agent_traces_support.csv Benefits-triage agent, one row per episode | Agent | A1 | Tool choice and routing differ by gender. | Validant (synthetic) CC BY 4.0 (Validant) | 240 rows 38 KB 6ecd2d323ffc |
| agent/agent_traces_otel.jsonl The same agent as raw OpenTelemetry GenAI spans | Agent | A1 | As above. | Validant (synthetic) CC BY 4.0 (Validant) | 2,130 rows 498 KB 5c7dd211f12c |
| agent/agent_traces_langfuse.jsonl The same agent as a Langfuse export | Agent | A1 | As above. | Validant (synthetic) CC BY 4.0 (Validant) | 240 rows 255 KB d3bf2ef9f64a |
| agent/loan_agent_persona_runs.csv A real tool-calling loan agent, counterfactual personas, one row per episode | Agent | A1 | Nothing planted: measured behaviour of the agent. | Validant run of qwen2.5:7b CC BY 4.0 (Validant) | 780 rows 138 KB bf3ae4f22d98 |
| agent/loan_agent_otel.jsonl The same runs as OpenTelemetry GenAI spans | Agent | A1 | Nothing planted. | Validant run of qwen2.5:7b CC BY 4.0 (Validant) | 1,550 rows 416 KB 3ab291bd8c5e |
| scripts/loan_agent_persona_harness.py The harness that produced the runs: point it at your own agent | Agent | A1 | n/a | Validant MIT | 12 KB 1f8593708553 |
| multi_agent/loan_committee_runs.csv Two-agent loan committee: screener, then reviewer, per-agent outputs | Multi-agent | A1 | Nothing planted: measured behaviour of the system. | Validant run of qwen2.5:7b CC BY 4.0 (Validant) | 780 rows 118 KB 57e1d55e27e4 |
| multi_agent/loan_committee_harness_trace.json The same run as a MultiAgentRunHarness trace | Multi-agent | A1 | Nothing planted. | Validant run of qwen2.5:7b CC BY 4.0 (Validant) | 49 KB 81c55b698ef5 |
Classic public datasets, at their origin
Where a well-known dataset can be redistributed, we host a ready-to-run version above and link the origin here. Where it cannot, we link the origin only. Always cite the origin.
| Dataset | Pathway | Real model outputs? | Tier you can reach | Licence and notes |
|---|---|---|---|---|
| COMPAS (ProPublica) | Predictive | Yes, a deployed tool's scores | A1 with labels | No licence file; names real people. Use our script. |
| Adult (UCI) | Predictive | No, labels only | A1 to A4 once you train a model | CC BY 4.0. Hosted above with predictions. Known problems: 1994 data, an arbitrary $50K threshold; see Retiring Adult. |
| folktables (US Census ACS) | Predictive | No | A1 to A4 once you train | MIT (code). The modern replacement for Adult: any US state, any year, several tasks. |
| German Credit (UCI) | Predictive | No | A1 to A4 once you train | CC BY 4.0. Only 1,000 rows; sex must be derived from a combined field with known coding problems. |
| Default of Credit Card Clients (UCI) | Predictive | No | A1 to A4 once you train | CC BY 4.0. 30,000 rows; sex, age, marriage, education. |
| Bank Marketing (UCI) | Predictive | No | A1 to A4 once you train | CC BY 4.0. Age and marital status only. |
| Diabetes 130-US Hospitals (UCI) | Predictive | No | A1 to A4 once you train | CC BY 4.0. About 100,000 encounters; race, gender, age. |
| Fairness dataset survey (Law School, Dutch census, and cleaned copies of the above) | Predictive | No | A1 to A4 once you train | The repository has no licence; each dataset keeps its original terms. |
| HMDA (US mortgage disclosure) | Predictive | A lender's decision, not one model's | A1 | US public data. Millions of rows per year. Treat the decision as an organisation's outcome, not a model output. |
| Boston Housing | Do not use. It encodes a race-derived variable and was removed from scikit-learn 1.2; the link explains why. | |||
| synthcity | Synthetic | n/a | All | Apache-2.0. Includes DECAF, a fairness-aware causal generator. |
| ml-fairness-gym | Synthetic | n/a | All | Apache-2.0. Feedback loops over time (lending, attention allocation). |
| SDV | Synthetic | n/a | All | Business Source License, not open source: check the terms before commercial use. |
| discrim-eval (Anthropic) | Generative | No, prompts only | A1; A2 with log probabilities | CC BY 4.0. Answered by an open model in our pack above. |
| BBQ | Generative | No | A1; A2 with option probabilities | CC BY 4.0. Nine bias categories; a subset ships with vfairness (BenchmarkRunner.run_bbq). |
| BOLD | Generative | No | A1 | CC BY 4.0. Open-ended continuations scored for sentiment and toxicity; a subset ships with vfairness. |
| HolisticBias | Generative | No | A1 | CC BY-SA 4.0. About 600 identity descriptors; a subset ships with vfairness. |
| CrowS-Pairs, StereoSet | Generative | No | A2 or A3 (they compare sentence likelihoods) | CC BY-SA 4.0. Not usable at A1. |
| Bias in Bios | Generative / text classifier | No | A1 to A3 | MIT. Biographies with gender and profession labels; good for probing a model's own representations. |
| HELM raw results | Generative | Yes, many models' outputs | A1 with gold answers, no model run needed | Per-instance outputs, including BBQ, for many models. |
| τ²-bench / τ-bench | Agent | No | A1 with traces | MIT. Customer-service tasks with simulated users: vary the user's persona and compare success. |
| AgentFairBench | Agent, multi-agent | No | A1 | Apache-2.0. Built for agent fairness (hiring, lending, triage). New in 2026 and not yet replicated: treat its results as preliminary. |
| Sotopia | Multi-agent | Yes, recorded episodes | A1, no model run needed | Social simulations between agents with persona profiles and goal scores. |
| MAST | Multi-agent | Yes, 1,642 traces | A1 | CC BY 4.0. Seven frameworks, labelled failure modes. No demographic axis: use it to test trace loading, not fairness. |
Trace formats your agent framework may already write: OpenTelemetry GenAI semantic conventions, OpenInference, Langfuse and LangSmith exports.
How these files were made
scripts/build_test_objects.pybuilds the synthetic and Adult packs and writes the manifest. It refuses to build if the UCI origin file has changed.scripts/build_live_test_objects.pybuilds the LLM, agent and multi-agent packs against a local model and records its digest. discrim-eval is fetched at a pinned revision and checked by sha256.scripts/verify_test_objects.pyruns every file through vfairness and writes the results shown above.
Names in the agent files are a standard correspondence-testing cue (Bertrand and Mullainathan, 2004). They signal a group to the model and say nothing about any real person. All personal details in the synthetic files are invented.