Two ways in
Choose your path.
Select the documentation view that matches your role.
Business User Guide
For AI Product Managers, Compliance Officers, and Business Stakeholders responsible for delivering fair AI products.
View Business GuideDeveloper Documentation
For Data Scientists, ML Engineers, and Developers who need the API reference, code examples, and integration guides.
View API ReferenceThe pipeline
Why vfairness?
A full-pipeline fairness library: detect, mitigate, calibrate, and monitor. Select a stage to filter the capabilities below.
What’s inside
Capabilities that cover every stage of the fairness lifecycle. Open any one for the detail.
Full ML Pipeline
A 6-stage pipeline plus 8 specialized analysis surfaces, from data preprocessing to multi-agent testing.
Detail
A 6-stage pipeline plus 8 specialized analysis surfaces and a cross-cutting egress guard, covering the entire fairness lifecycle: data preprocessing, feature engineering, training-time interventions, post-processing calibration, evaluation, CI/CD operations, LLM testing, agent testing, and multi-agent testing. The LLM surface plugs in production scorers (VADER sentiment, alt-profanity-check toxicity) through the [llm] extra, and warns loudly when it falls back to its keyword placeholder.
Statistical Rigor
Bootstrap and Bayesian confidence intervals, permutation tests, effect sizes and multiple-testing corrections.
Detail
Built-in bootstrap & Bayesian confidence intervals, permutation testing, effect sizes (Cohen's d, odds ratio), and multiple testing corrections. Publication-ready statistical validation.
Training-Time Interventions
12 fairness-aware loss functions, adversarial debiasing, constraint-based training and scikit-learn compatible wrappers.
Detail
12 fairness-aware loss functions, adversarial debiasing, counterfactual losses, constraint-based training (Exponentiated Gradient, Grid Search), 5 regularizers, and scikit-learn compatible FairClassifier/FairRegressor wrappers.
Post-Processing & Calibration
5 calibration methods, group-aware calibrators, threshold optimization and trade-off diagnostics.
Detail
5 calibration methods (Platt, Isotonic, Beta, Temperature, Histogram), group-aware calibrators, threshold optimization, prediction reweighting, and impossibility theorem trade-off diagnostics.
FairExplAIner
A human-readable explanation, severity and recommendation for every metric.
Detail
Human-readable explanations for every metric. Understand what numbers mean, get severity assessments, and receive actionable recommendations, with no fairness PhD required.
Regulatory Compliance
Historical discrimination patterns with risk levels and precedents, and legal rule packs for seven jurisdictions.
Detail
43 historical discrimination patterns (as of 2026-08-28; run len(HISTORICAL_RISK_PATTERNS) for the current figure), each with a risk level, the affected groups and the precedent behind it. Seven of them cite EU AI Act Art. 5 prohibited practices and eight cite Annex III high-risk systems, four naming the EU AI Act penalty ceiling (up to €35 million or 7% of global turnover) in their recommendations. Separately, the legal rule packs cover seven jurisdictions: us-federal, us-ca, eu, uk, de, ch and br.
Intersectional Analysis
Hidden disparities at group intersections: auto-discovered attributes, proxy variables and subgroup audits.
Detail
Detect hidden disparities at group intersections. Auto-discover protected attributes, identify proxy variables via correlation & mutual information, and audit subgroups for fairness gerrymandering.
MLOps & CI/CD
MLflow, pytest assertions, training callbacks, fairness gates and drift detection for production.
Detail
MLflow integration, pytest assertions, training callbacks for Keras/PyTorch/sklearn, DataBiasValidator, ModelFairnessGate, real-time drift detection, adaptive alert thresholds, and prioritized alert routing for production monitoring.
Privacy-Preserving Reporting
A three-tier privacy scheme, on by default: small groups suppressed, mid-size groups noised, large groups exact.
Detail
Built-in three-tier privacy scheme, on by default in MetricsStore: groups under 10 are suppressed by k-anonymity, groups of 10 to 50 receive ε-differential-privacy Laplace noise, groups above 50 report exact values. Every row is labelled with the tier it came from, and the library says plainly that at the default ε = 1.0 the noisy tier is close to uninformative: treat a privacy_level == "noisy" row as withheld unless you have knowingly raised ε.
Clean, Unified API
One analyzer pattern across all modules, for classification, regression and ranking metrics.
Detail
Consistent analyzer pattern across all modules: FairnessAnalyzer, BiasDetector, CalibrationAnalyzer, FairnessTrainingAnalyzer, ThresholdAnalyzer. Classification, regression, and ranking metrics with sensible defaults.
Publication-Ready SVG Reports
Templated, self-contained SVG reports: dashboards, bias audits, calibration diagrams, drift reports and more.
Detail
A library of templated SVG visualizations: fairness dashboards, bias audits, calibration diagrams, drift reports, Pareto frontiers, and more. Beautiful, self-contained vector graphics ready for papers, presentations, and stakeholder reports.
Fairness A/B Testing
Experiments with per-intersection power analysis, sequential testing and Pareto trade-off optimization.
Detail
Full experiment framework with per-intersection power analysis, sequential testing (SPRT) for early stopping, and Pareto frontier optimization across fairness/accuracy trade-offs.
Auto-Discovery
Finds protected attributes, proxy variables, violations and intersectional subgroups without manual configuration.
Detail
Automatically detect protected attributes, identify proxy variables, scan for fairness violations, and discover intersectional subgroups, so you can run a full audit without manual configuration.
LLM Fairness Testing
Bias tests for LLMs without training data: counterfactual prompts, BBQ, BOLD, HolisticBias and DecodingTrust.
Detail
Test LLMs for bias without training data. Counterfactual prompt testing (9 strategies including persona-based), standardized benchmarks (BBQ, BOLD, HolisticBias), the 8-dimensional DecodingTrust suite (Wang et al. 2023, NeurIPS), output analysis, and non-determinism management for foundation models.
Agent Fairness Testing
Bias tests for AI agents: tool selection, RAG retrieval, action outcomes and delegation patterns.
Detail
Test AI agents for bias in tool selection, RAG retrieval, action outcomes, and delegation patterns. Includes correspondence testing and temporal trajectory tracking across agent lifecycles.
Multi-Agent Fairness Testing
Six analyzers for bias that only appears when agents interact, plus a framework-agnostic run harness.
Detail
Detect emergent bias in multi-agent systems. Six analyzers: compositionality, groupthink/echo-chamber, emergent amplification, adversarial collusion (Khan et al. 2023), demographic-conditional delegation routing (Bertrand & Mullainathan 2004), and turn-by-turn negotiation drift (Bianchi et al. 2024), plus a framework-agnostic MultiAgentRunHarness for autogen / crewai / langgraph.
Audit Trail on Every Result
Structured logging and audit metadata on every result, with a methodology version stamped into every report.
Detail
Structured logging via Python logging, audit-trail metadata (ISO 8601 UTC timestamp, library version, parameters) on every result, a methodology_version stamped into every report, to_dict()/to_json() serialization, progress callbacks for batch operations, and configurable thresholds. That is the record-keeping material an obligation such as EU AI Act Art. 12 asks for; conformity is a judgement about a deployment, which no library can make on its own.
From the Signal
What a fairness check is up against.
Three drawings from validant.ai’s Signal articles, and the part of vfairness that answers each one.
“Each column is a fairness check the model must pass before a decision; the figure walking the corridor is the model under audit.”
In vfairness35+ metrics, each with a three-state verdict
Bias is the Foundation →
“The single tall bar is the position where the harm lives; the average smooths it out of view.”
In vfairnessIntersectional subgroups, measured one by one
The Wrong Question, Asked at Scale →
“Bias enters at every stage of the lifecycle, and a feedback loop carries it back into the world.”
In vfairnessDetect, mitigate, calibrate and monitor
Bias is the Foundation →A real run
Ten subgroups. One it will not guess.
Eleven lines on the UCI Adult census data, grouped by race and sex. vfairness reports every rate it measured, the largest gap against its limit, and the one subgroup too small to judge: it is named and left out of the verdict, not counted as fair.
- Provenance in the verdict. How many rows were assessed, excluded and withheld, in the summary line itself.
- Three states, not two. Measured, failed, or could not check. Never a silent zero.
import pandas as pd
from vfairness import FairnessAnalyzer
df = pd.read_csv("adult_test_with_predictions.csv")
analyzer = FairnessAnalyzer(
y_true=df.y_true, y_pred=df.y_pred,
sensitive_attr=df.race + " / " + df.sex,
min_group_size=50,
)
report = analyzer.get_report()
print(report["assessment"]["summary"])
Try it
Move the threshold. Watch the verdict.
A model scores 15,060 people from the UCI Adult census data, and everyone above the threshold is approved. Each number below is what vfairness reports at that threshold.
actually earned over 50Kdid notapprovedbar height: share of the group, square-root scale
Every number is vfairness output on the UCI Adult test split, precomputed for each threshold by scripts/build_landing_demos.py; the page only looks them up. The data · GroupThresholdOptimizer
Integrates with your ML stack
Quick start
From install to a verdict.
Install from PyPI, then run your first audit.
# vfairness 0.1.0 (beta), Python 3.11 or newer
pip install vfairness
from vfairness import FairnessAnalyzer
# Create analyzer with your model predictions
analyzer = FairnessAnalyzer(
y_true=actual_outcomes,
y_pred=model_predictions,
sensitive_attr=demographic_groups,
fair_explainer=True # Enable human-readable explanations
)
# Generate comprehensive report with confidence intervals
report = analyzer.get_report(include_ci=True, n_bootstrap=5000)
# The verdict, the rows it was computed over, and anything it could not check
print(report['assessment']['summary'])
# Access metrics
print(f"Demographic Parity: {report['metrics']['demographic_parity_difference']:.3f}")
print(f"Equal Opportunity: {report['metrics']['equal_opportunity_difference']:.3f}")
# Get explanations
for metric, explanation in report['explanations']['metrics'].items():
print(f"\n{metric}: {explanation['severity']}")
print(f" {explanation['evaluation']}")
print(f" {explanation['recommendation']}")
0/5 metrics within thresholds (data provenance: 1200 of 1200 rows assessed, 0 excluded, missing_strategy='exclude')
Demographic Parity: 0.255
Equal Opportunity: 0.280
demographic_parity_difference: critical
Critical. The difference of 0.2553 indicates severe disparity that requires immediate investigation and remediation.
URGENT: Review deployment decisions for this model. ...
The data provenance clause is on every summary, and the count is over the metrics that were actually graded. A metric or a group that could not be measured is named separately and never folded into that count. See Three States, Not Two.
15 sub-packages
Library architecture.
15 top-level sub-packages: a 6-stage fairness pipeline, 8 specialized analysis surfaces and 1 cross-cutting infrastructure package. The main areas follow the ML fairness pipeline from data to production.
vfairness.preprocessing
vfairness.in_processing
vfairness.post_processing
vfairness.evaluation
vfairness.operations.monitoring
vfairness.operations.reporting
vfairness.operations.experimentation
vfairness.operations.cicd
vfairness.llm
vfairness.agents
vfairness.multi_agent
Preprocessing"] --> B["2 · Training-Time
Interventions"] B --> C["3 · Prediction-Time
Interventions"] C --> D["4 · Evaluation &
Measurement"] D --> E["5 · Monitoring"] D --> F["6 · Reporting &
Dashboards"] D --> G["7 · Experimentation"] D --> H["8 · Workflow
Integration"] D --> I["9 · LLM Fairness
Testing"] I --> J["10 · Agent Fairness
Testing"] J --> K["11 · Multi-Agent
Testing"]
AIF360 · Fairlearn · Aequitas
How vfairness compares.
Other open-source fairness libraries are excellent at what they do. vfairness focuses on an end-to-end, audit-grade workflow that spans data and modelling, production monitoring, and regulatory evidence.
| Feature | vfairness | AIF360 | Fairlearn | Aequitas |
|---|---|---|---|---|
| Group fairness metrics | Extensive | 70+ | Yes | Yes |
| Bias mitigation (pre / in / post-processing) | Full | Full | Yes | Audit-only |
| Prediction-time interventions (calibration, thresholds) | Full | Yes | Yes | No |
| Data & preprocessing bias auditing | Full | Yes | Partial | Full |
| Fairness explainability (SHAP / IG / counterfactual) | Yes | Partial | No | No |
| Confidence intervals & small-sample (Bayesian) | Built-in | Manual | Manual | Manual |
| Auto-discovery (protected attributes & proxies) | Yes | No | No | No |
| Ranking / recommender fairness | Yes | No | No | No |
| Production monitoring & drift gates (CI/CD) | Built-in | No | No | No |
| pytest fairness assertions | Built-in | No | No | No |
| MLflow / experiment logging | Native | Manual | Manual | Manual |
| LLM & agent fairness testing | Yes | No | No | No |
| Regulatory compliance mapping | 7 jurisdictions | No | No | No |
Compared against each library's out-of-the-box capabilities from public documentation (reviewed June 2026). AIF360 and Fairlearn both provide mature, well-validated mitigation algorithms, and Aequitas is a focused bias-audit toolkit; all are actively developed and may add capabilities over time.
What it does not do
Known limitations.
Transparency about what vfairness does and does not do.
vfairness is in beta (v0.1.0) published 2026-10-02. APIs may change before the stable v1.0.0 release. The version number says beta, and the beta gate agrees: scripts/release_gate.py reports BETA READY, with all eight beta conditions met and no known defect open. The stricter 1.0 gate is not met yet: a second, independent examiner still has to confirm part of the graded code. Quality and Hardening carries the standing figures.
vfairness==0.1.0, because the API is not frozen until 1.0.0.
Population benchmarks for representation analysis must be user-provided. Defaults are US Census 2020 only.
Historical pattern and protected attribute detection uses column name keyword matching, not NLP or semantic analysis.
Groups below min_group_size (30) are excluded from the disparity metrics, so they get no verdict. Minority and intersectional groups are most affected. The exclusion is reported, not silent: a UserWarning names each dropped group and its size, and the report lists it under insufficient_evidence_groups with a reason, outside the pass/fail count.
insufficient_evidence_groups before quoting the headline.
Demographic parity, equalized odds, and calibration cannot all be satisfied simultaneously (Impossibility Theorem). Choose your metric before analysis.
HOLC redlining data covers ~40 US cities only. No European, Swiss, or APAC geographic discrimination data is included.
Detailed limitation callouts appear throughout the documentation, marked with the Limitation badge.
Read next
Resources.
Getting Started Guide
Installation, first fairness analysis, and basic concepts.
Business User Guide
Comprehensive guide for product managers and compliance officers.
API Reference
Reference for the library’s classes, functions, parameters and return types. It does not yet cover every public code unit.
Core Concepts
Deep dives into fairness definitions, metrics, and methodology.
SVG Templates
Interactive gallery of fairness visualisations, dashboards, and report templates.
