Start · Getting started

Getting Started with vfairness

A full-pipeline Python library for measuring, mitigating, and monitoring fairness in machine learning systems.

vfairness covers the entire ML fairness lifecycle: from pre-training bias auditing and fairness-aware model training, through group-specific calibration and post-processing, to production monitoring, stakeholder reporting, and controlled experimentation. Every stage produces auditable artifacts (SVG visualizations, structured reports, and machine-readable metrics) designed for regulatory environments including the EU AI Act.

~141k
Lines of Code
181,328 total across 227 files. Built on numpy, pandas, scipy, scikit-learn, requests.
Hover or focus for the split
15
Sub-Packages
Plus 44 SVG templates from the rendering package.
Hover or focus for what each covers
215
Capabilities
Registry rows naming 222 objects.
Hover or focus for the full list
475
Functions
Across 204 modules.
Hover or focus for all 475
392
Classes
Across 204 modules.
Hover or focus for all 392
691
Methods
Across 204 modules.
Hover or focus for all 691

1,558 public functions, classes and methods across 204 modules. A capability is one dispatch key in the registry, not a fourth kind of object: the 215 capability rows name 222 of those same functions, classes and methods, so the capability tile is not one to add to the other three. One module, vfairness.mcp.server, could not be imported (ImportError), so whatever it defines is not in these counts. An uncounted module understates the surface, so it is reported rather than dropped. How much of this surface has been tested is reported on Quality and Hardening, in these same four nouns.

Installation

bash
# Core (numpy, pandas, scipy, scikit-learn, requests)
pip install vfairness

# With visualization support (matplotlib, seaborn)
pip install "vfairness[viz]"

# With interactive dashboards (Plotly + Dash)
pip install "vfairness[interactive]"

# Plotly charts only
pip install "vfairness[dashboard]"

# Experiment tracking (MLflow, Weights & Biases)
pip install "vfairness[mlops]"

# Full installation (all optional dependencies)
pip install "vfairness[all]"

0.1.0 was published on 2026-10-02 and installs with pip install vfairness. The version number says beta, and the beta gate agreed when it was cut: scripts/release_gate.py reports BETA READY, with all eight beta conditions met. The stricter 1.0 gate is not met yet: a second, independent examiner still has to confirm part of the graded code. The API is not frozen until 1.0.0. See the Releasing for how versions will ship, and Quality and Hardening for what has been graded, what has not, and which defects are still open.

python
import vfairness
print(f"vfairness {vfairness.__version__}")

Library Structure

vfairness is organized into 15 top-level sub-packages: a 6-stage fairness pipeline (preprocessing, in_processing, post_processing, evaluation, operations, rendering), 8 specialized analysis surfaces (llm, agents, multi_agent, xai, vision, legal, mcp, validity) and 1 cross-cutting infrastructure package (net, the egress guard). The areas below walk that pipeline from data to production.

1
Data & Preprocessing vfairness.preprocessing
Detect historical discrimination patterns, find proxy variables, audit representation, and apply fairness-aware feature transforms before training.
2
Training-Time Interventions vfairness.in_processing
Train models with fairness constraints using reductions, Lagrangian methods, and adversarial debiasing. Compare methods and visualize accuracy/fairness trade-offs.
3
Prediction-Time Interventions vfairness.post_processing
Optimize group-specific decision thresholds and reweight predictions to equalize outcomes without retraining. Includes CI/CD deployment gates.
4
Evaluation & Measurement vfairness.evaluation
Compute 35+ fairness metrics for classification, regression, and ranking tasks. Add bootstrap confidence intervals, effect sizes, and multiple-testing corrections.
5
Monitoring vfairness.operations.monitoring
Track fairness in production with sliding-window metrics, multi-scale drift detection (wavelet + KS), adaptive alert thresholds, and temporal trend analysis.
6
Reporting & Dashboards vfairness.operations.reporting
Generate multi-tier reports (executive, engineer, auditor) with natural language summaries. Build interactive Plotly dashboards with progressive disclosure.
7
Experimentation vfairness.operations.experimentation
Run fairness-aware A/B tests with intersectional analysis, per-subgroup power analysis, SPRT early stopping, Pareto optimization, and causal decomposition.
8
Workflow Integration vfairness.operations.cicd
Embed fairness into CI/CD with deployment gates, hierarchical intersectional checking, pytest plugin, pre-commit hooks, MLflow/W&B logging, and PR report cards.
9
LLM Fairness Testing vfairness.llm
Test LLMs for bias without training data. Counterfactual prompt testing, benchmark evaluation (BBQ, BOLD), output analysis, and non-determinism management.
10
Agent Fairness Testing vfairness.agents
Test AI agents for bias in tool selection, RAG retrieval, action outcomes, and delegation patterns. Includes correspondence testing and temporal trajectory tracking.
11
Multi-Agent Fairness Testing vfairness.multi_agent
Detect emergent bias in multi-agent systems. Tests non-compositionality, groupthink convergence, and coalition formation.

Step 0: Find Something to Test

Every analysis starts with a test object: a table of decisions, a model, a chat endpoint, an agent or a team of agents. How deep the analysis can go depends on what that object lets you see, not on how you connect to it. Decisions alone support claims about outcomes; probabilities add calibration; only the model itself supports claims about why.

You are testingWith outputs only (A1)With scores (A2)Starter file
A predictive modeldecisions + groups (+ labels for error rates)+ probabilities: calibrationAdult with predictions
An LLMprompts in, text out+ log probabilities: likelihood biasdiscrim-eval answers
An agentoutcomes, plus tool calls if you have traces+ log probabilities per steploan agent runs
Several agentseach agent's output, not only the lastas for agentsloan committee runs
Your detector itselfsynthetic data with a planted bias and a clean controlplanted and clean

Find a test object explains every pathway, the data each depth needs, and links each file to its source. Every file there was run through vfairness before it was published.

Quick Start

Measure fairness in a loan approval model in under 30 lines. The output below is what this script prints, captured from a run against 0.1.0.

python
import numpy as np
from vfairness import FairnessAnalyzer, demographic_parity_difference

# Sample data
y_true = np.array([1, 0, 1, 0, 1, 0, 1, 0, 1, 1] * 100)
y_pred = np.array([1, 0, 1, 1, 1, 0, 0, 0, 1, 1] * 100)
gender = np.array(['M','M','M','M','M','F','F','F','F','F'] * 100)

# Single metric
dp = demographic_parity_difference(y_true, y_pred, gender)
print(f"Demographic Parity Difference: {dp:.3f}")

# Full analyzer with a confidence interval.
# random_state pins the bootstrap, so the interval is reproducible.
analyzer = FairnessAnalyzer(y_true, y_pred, gender, min_group_size=30)
result = analyzer.demographic_parity_difference(
    include_ci=True, n_bootstrap=5000, confidence_level=0.95, random_state=0
)
lo, hi = result.confidence_interval
print(f"DP Diff: {result.value:.3f}  95% CI: [{lo:.3f}, {hi:.3f}]")

# All metrics at once
for name, value in analyzer.compute_all_metrics().items():
    print(f"  {name}: {value:.3f}")

# The verdict, and the rows it was computed over
print(analyzer.get_report()["assessment"]["summary"])
output
Demographic Parity Difference: 0.400
DP Diff: 0.400  95% CI: [0.346, 0.456]
  demographic_parity_difference: 0.400
  equalized_odds_difference: 0.500
  equal_opportunity_difference: 0.333
  demographic_parity_ratio: 0.500
  predictive_parity_difference: 0.250
0/5 metrics within thresholds (data provenance: 1000 of 1000 rows assessed, 0 excluded for missing values, 0 in 0 group(s) withheld by the group-size floor, missing_strategy='exclude')

The data provenance clause is part of every summary. It states how many rows reached the metrics and how many were dropped on the way, so a verdict can never be read without knowing what it was computed over.

Not every report metric gets an interval Limitation

get_report(include_ci=True) fills metrics_with_ci for three of the five default metrics: demographic_parity_difference, equalized_odds_difference and equal_opportunity_difference. demographic_parity_ratio and predictive_parity_difference come back as plain point estimates there, with no interval and no interval-derived verdict.

Do: for those two, call disparate_impact_ratio_with_ci and predictive_parity_difference_with_ci directly. This matters most for the four-fifths ratio, which is the statistic a regulator is most likely to ask about.
Choose Your Metric First Limitation

Fairness metrics are mathematically incompatible: you cannot satisfy demographic parity, equalized odds, and calibration simultaneously (Kleinberg et al., 2016). Decide which definition matches your use case before running analysis. See The Impossibility Theorem.

Three States, Not Two

Sooner or later a run will tell you it could not check something. That is not a failure of the library and it is not a fairness violation. It is the third state, and it is deliberate: a group or a metric that was never measured gets no number, no verdict, no colour, and no place in any count that implies it was measured. Reading these outputs is the one thing worth learning before anything else on this page.

  • Assessed, within tolerance: measured, and the evidence supports the pass.
  • Assessed, exceeds tolerance: measured, and the evidence supports the breach.
  • Could not check: not measured. Never folded into either of the other two.

A group too small to assess

Add 12 nonbinary applicants to the 1,000 rows above. They are below the default min_group_size of 30, so they cannot carry a verdict. They are not dropped in silence either: a UserWarning names them, and they appear in the report as insufficient evidence.

python
import numpy as np
from vfairness import FairnessAnalyzer

# The same 1,000 rows as above, plus 12 nonbinary applicants.
y_true = np.array([1, 0, 1, 0, 1, 0, 1, 0, 1, 1] * 100)
y_pred = np.array([1, 0, 1, 1, 1, 0, 0, 0, 1, 1] * 100)
gender = np.array(['M','M','M','M','M','F','F','F','F','F'] * 100)

y_true = np.append(y_true, [1, 0] * 6)
y_pred = np.append(y_pred, [1, 1] * 6)
gender = np.append(gender, ['NB'] * 12)

assessment = FairnessAnalyzer(y_true, y_pred, gender).get_report()["assessment"]

print(assessment["summary"])

group = assessment["insufficient_evidence_groups"][0]
print(f"{group['group']}: n={group['n']}, verdict={group['verdict']}")
print(group["reason"])
output
UserWarning: Excluding 1 group(s) below min_group_size=30: {'NB': 12}. That leaves
12 of 1012 rows (1.2 percent) out of this metric. The metric is computed over the
remaining groups only, so it measures nothing about the excluded group(s): they are
could not check, not a measured pass. To include them, lower min_group_size.

0/5 metrics within thresholds (1 group(s) with insufficient evidence, excluded from the verdict: ['NB']) (data provenance: 1012 of 1012 rows assessed, 0 excluded for missing values, 12 in 1 group(s) withheld by the group-size floor, missing_strategy='exclude')
NB: n=12, verdict=insufficient_evidence
n=12 is below min_group_size=30; this group is excluded from the disparity metrics, so there is insufficient evidence to assess it. Underpowered (n<30): the rate is noisy; treat as indicative only.

Note what the summary does not say. It does not report 5/5, and it does not report a violation for NB. It reports the count over the metrics it graded, and it names the group it could not reach.

A run with nothing left to compare

A disparity metric needs at least two groups. When filtering leaves only one, there is no comparison to make, so the whole report is marked not assessable and the score is None rather than a number.

python
import numpy as np
from vfairness import FairnessAnalyzer

# 400 men and 3 women: only one group is large enough to measure, so there is
# nothing to compare it against.
gender = np.array(['M'] * 400 + ['F'] * 3)
y_true = np.append(np.random.default_rng(1).integers(0, 2, 400), [1, 0, 1])
y_pred = np.append(np.random.default_rng(2).integers(0, 2, 400), [1, 1, 0])

assessment = FairnessAnalyzer(y_true, y_pred, gender).get_report()["assessment"]

print("assessable:    ", assessment["assessable"])
print("fairness_score:", assessment["fairness_score"])
print("not assessable:", [m["metric"] for m in assessment["not_assessable_metrics"]])
print(assessment["summary"])
output
assessable:     False
fairness_score: None
not assessable: ['demographic_parity_difference', 'demographic_parity_ratio', 'equalized_odds_difference', 'equal_opportunity_difference', 'predictive_parity_difference']
NOT ASSESSABLE: 1 valid group(s) after filtering (need at least 2): disparity metrics are vacuous and do not certify fairness. (1 group(s) with insufficient evidence, excluded from the verdict: ['F']) (5 metric(s) not assessable, excluded from the verdict) (data provenance: 403 of 403 rows assessed, 0 excluded for missing values, 3 in 1 group(s) withheld by the group-size floor, missing_strategy='exclude')

fairness_score is None, not 0.0. A zero would be read as a measured, maximally unfair result, which is the opposite of what happened here. The same rule governs the release gates: assert_fairness refuses a metric whose value is NaN with "NOT MEASURABLE ... so the metric was never compared against threshold (fail closed)", and ModelFairnessGate blocks rather than approving a check that never ran. The charts follow it too: an empty reliability diagram renders "NOT ASSESSABLE: no data was supplied, so nothing was measured" and prints its ECE as N/A, instead of a reassuring 0.000.

Production Features

Every test result carries audit-trail metadata: an ISO 8601 UTC timestamp, the library version that produced it, and the parameters it ran with. That is the raw material for a record-keeping obligation such as EU AI Act Art. 12; whether a given deployment satisfies that obligation is a question for its operator, not something a library can establish.

python
from vfairness.llm import OutputAnalyzer

analyzer = OutputAnalyzer(alpha=0.05)
result = analyzer.analyze_sentiment(texts_a, texts_b)

# Every result has audit metadata
print(result.metadata.timestamp)        # ISO 8601 UTC
print(result.metadata.library_version)  # current vfairness.__version__
print(result.metadata.parameters)       # {"alpha": 0.05, "metric": "sentiment", ...}

# Serialize for database storage
data = result.to_dict()    # Plain dict
json_str = result.to_json()  # JSON string

# Structured logging (configure once)
import logging
logging.basicConfig(level=logging.INFO)
analyzer.analyze_all(texts_a, texts_b)
output
INFO:vfairness.llm.output_analysis:analyze_all: group_a=group_a (18 samples), group_b=group_b (18 samples), correction=benjamini_hochberg
WARNING:vfairness.llm.output_analysis:Sample size (18, 18) below recommended minimum of 25 for metric 'semantic_quality'
WARNING:vfairness.llm.output_analysis:Sample size (18, 18) below recommended minimum of 25 for metric 'sentiment'
...
INFO:vfairness.llm.output_analysis:LLM judge scorer not configured, skipping analyze_llm_judge
INFO:vfairness.llm.output_analysis:analyze_all complete: 11 metrics evaluated

The sample-size warning is emitted once per metric, so a thin run announces the weakness of every number it is about to hand you.

Production Scorers

The sentiment and toxicity scorers delegate to third-party packages that ship in the [llm] extra: pip install "vfairness[llm]". The rest are implemented in vfairness itself, and the counts below are read straight off those implementations.

MetricScorerWhere it comes fromDetails
SentimentVADER[llm] extraDelegated to vaderSentiment. Lexicon-based, with negation and intensity handling.
Toxicityalt-profanity-check[llm] extraDelegated to alt-profanity-check, a linear-SVM classifier. Its accuracy is the upstream project's to report, not ours.
RefusalPattern-basedin vfairness63 patterns in 5 categories (hard, soft, partial, conditional, policy), weighted scoring.
HelpfulnessMulti-signal heuristicin vfairness6 quality signals: length, vocabulary, structure, specificity, engagement, deflection.
StereotypeCurated word listsin vfairness61 terms plus 14 phrase patterns (gender, racial, age, religious).
Without the extra, you get a placeholder Limitation

If vaderSentiment is not installed, the sentiment scorer falls back to a 30-word keyword list and raises a PlaceholderScorerWarning saying so: "Using keyword-based sentiment scorer (30 words only). This is a LOW-ACCURACY placeholder." The warning is the signal that a number came from the fallback and not from VADER. Do not treat a scored run as a production result until the extra is installed and the warning is gone.

Progress Callbacks

Long-running batch operations support progress callbacks for UI integration:

python
from tqdm import tqdm
from vfairness.llm import LLMApiProxy

proxy = LLMApiProxy(endpoint_url=..., api_format="openai", model_name=...)

# With tqdm progress bar
pbar = tqdm(total=100)
proxy.send_batch(
    prompts=my_prompts,
    n_runs=25,
    progress_callback=lambda current, total: pbar.update(1)
)
pbar.close()

Bias Detection

Audit training data for representation bias, proxy variables, and historical discrimination patterns before any model is trained.

python
import pandas as pd
from vfairness import BiasDetector, detect_historical_patterns, identify_proxy_variables

# Full audit
detector = BiasDetector(
    df,
    protected_attributes=['gender', 'age'],
    outcome_column='approved',
    benchmarks={'gender': {'M': 0.49, 'F': 0.51}}
)
report = detector.full_audit()
print(report.summary())

# Historical discrimination patterns (scans all column names)
for r in detect_historical_patterns(df):
    print(f"  {r.feature}: {r.risk_level.value} - {r.pattern_type}")

# Proxy variable detection
for p in identify_proxy_variables(df, protected_attributes=['gender'], correlation_threshold=0.3):
    print(f"  {p.feature} -> {p.protected_attribute}  (r={p.correlation:.3f})")

On a 600-row frame with a zip_code column, the historical scan reports:

output
Overall Risk Score: 50.0%

SUMMARY OF FINDINGS
Historical Pattern Issues: 1
Representation Bias Issues: 2
Statistical Disparities: 0
Proxy Variables Identified: 0

TOP RECOMMENDATIONS
1. CRITICAL: Review and consider removing high-risk features: zip_code
...

  zip_code: high - Geographic Redlining

Pass a benchmarks entry for every protected attribute. If one is missing the detector says so out loud, with a UserWarning naming the attribute and the fallback it used: "No benchmark provided for 'age'; falling back to the built-in approximate US Census 2020 distribution." A representation finding against a default benchmark is a finding about US Census 2020, not about your population.

Calibration

Ensure that a predicted probability of 70% means the same risk for every demographic group. Detect and correct calibration disparities.

python
from vfairness.post_processing.calibration import CalibrationAnalyzer

analyzer = CalibrationAnalyzer(
    y_true=labels, y_prob=model_probabilities,
    protected_attr=gender, attribute_name='gender'
)

report = analyzer.full_analysis(context='lending')
print(f"Well Calibrated: {report.is_well_calibrated}")
print(f"Significant Disparity: {report.has_significant_disparity}")
print(f"Overall ECE: {report.overall_metrics['ece']:.4f}")

# Apply group-specific recalibration
analyzer.fit_calibrator(method='isotonic')
calibrated_probs = analyzer.transform(y_prob_test, gender_test)

Explanations

Generate context-aware, stakeholder-ready explanations for every metric and audit result.

python
from vfairness import FairnessAnalyzer

# FairExplAIner mode for per-metric explanations
analyzer = FairnessAnalyzer(y_true, y_pred, gender, fair_explainer=True)
report = analyzer.get_report(include_ci=True)

for name, expl in report['explanations']['metrics'].items():
    print(f"{name}: {expl['severity']} - {expl['evaluation']}")

# FairnessExplainer for cross-module explanations
from vfairness.explainer import FairnessExplainer

explanation = FairnessExplainer.explain(audit_report)  # Takes a vfairness report object (BiasAuditReport, CalibrationReport, GateDecision, ...)
print(explanation.severity)           # 'info' | 'low' | 'medium' | 'high' | 'critical'
print(explanation.recommendations)    # Actionable next steps

Visualization

A library of SVG templates for card-based dashboards, plus Matplotlib and Plotly outputs. SVG rendering needs the [rendering] extra (Jinja2); the output SVG is self-contained. Call vfairness.rendering.list_templates() for the current set. As of 2026-08-28 it returns 46 names: 44 renderable charts plus two internal partials, _shared_defs and _could_not_check. The second one is the third state made visible: it is the panel a chart renders in place of a plot when it measured nothing.

python
from vfairness import classification_fairness_report, plot_fairness_metrics, create_fairness_dashboard
from vfairness.rendering import radar_chart_to_svg, fairness_detailed_report_to_svg

report = classification_fairness_report(y_true, y_pred, gender, include_ci=True)

# Matplotlib (static): returns a Matplotlib Axes
ax = plot_fairness_metrics(report, style='academic', show_thresholds=True)
ax.figure.savefig('fairness_metrics.png', dpi=300, bbox_inches='tight')

# Plotly (interactive)
dashboard = create_fairness_dashboard(report, style='modern')
dashboard.write_html('fairness_dashboard.html')

# SVG templates (needs the [rendering] extra: Jinja2)
svg = radar_chart_to_svg(metrics_data, save_path='radar.svg')

plot_fairness_metrics needs the [viz] extra (Matplotlib); create_fairness_dashboard needs the [dashboard] extra (Plotly); the SVG templates need the [rendering] extra (Jinja2). Unlike the Matplotlib and Plotly charts, the SVG templates render to self-contained SVG with no runtime, browser, or JavaScript dependency.

See the SVG Gallery for every template with live previews.

Monitoring

Track fairness in production with real-time metrics, multi-scale drift detection, and adaptive alerting.

python
import pandas as pd
from vfairness.operations.monitoring import FairnessMonitor, FairnessDriftDetector

# Real-time monitoring: each batch is a DataFrame with
# 'prediction', 'label', and one 'group_<attr>' column per protected attribute
batch_df = pd.DataFrame({
    'prediction': predictions,
    'label': labels,
    'group_gender': gender,
})

monitor = FairnessMonitor(window_size=1000)
window = monitor.update_and_check(batch_df)
if window.any_alert:
    print(f"Alert! Flagged metrics: {[k for k, v in window.alerts.items() if v]}")

# Multi-scale drift detection: set a known-good baseline once,
# then compare each new stretch of metric history against it
detector = FairnessDriftDetector()
detector.set_baseline(pd.Series(reference_metrics))
drift = detector.check_drift(pd.Series(current_metrics), metric="demographic_parity")
print(f"Drift detected: {drift.drift_detected}  (score: {drift.overall_drift_score:.3f})")

Reporting

Transform metrics into stakeholder-ready intelligence with multi-tier reports and interactive dashboards.

python
from vfairness.operations.reporting import MetricsStore, FairnessDashboard, ReportGenerator

store = MetricsStore()
store.ingest_from_monitor(monitor)

# Interactive Plotly dashboard
dashboard = FairnessDashboard(store)
fig = dashboard.create_executive_view()
fig.write_html("dashboard.html")

# Automated multi-format report (HTML, PDF, JSON)
gen = ReportGenerator(store, dashboard)
report = gen.generate_executive_report()
report.save("executive_report.html")

create_executive_view needs Plotly (the [dashboard] extra); MetricsStore and ReportGenerator run on the core install.

Experimentation

Run fairness-aware A/B tests with intersectional analysis and automated deployment recommendations.

python
from vfairness.operations.experimentation import (
    FairnessExperiment, FairnessPowerAnalyzer, ExperimentAnalysis
)

exp = FairnessExperiment(
    control_data=df_ctrl, treatment_data=df_treat,
    protected_attributes=['gender', 'race'], outcome_column='approved',
)
result = exp.run_full_analysis()

# Per-intersection power analysis
power = FairnessPowerAnalyzer(exp)
print(power.get_power_summary())

# Multi-objective analysis and deployment decision
analysis = ExperimentAnalysis(result, experiment=exp)
rec = analysis.decision_recommendation()
print(f"Decision: {rec.decision}")
output
Decision: RecommendationDecision.EXTEND_EXPERIMENT

rec.decision is a RecommendationDecision enum with four members: DEPLOY_TREATMENT, KEEP_CONTROL, EXTEND_EXPERIMENT and INVESTIGATE_FURTHER. The last two are the three-state discipline in experiment form: an experiment that has not yet separated the arms returns "keep running", never a deploy recommendation resting on an underpowered comparison. FairnessPowerAnalyzer reports is_powered per intersection so you can see which cell is holding the decision open.

Workflow Integration

Embed fairness checks into every stage of the development lifecycle, from experiment tracking and version control to CI/CD gates and collaborative review.

python
from vfairness.operations.cicd import (
    ModelFairnessGate, HierarchicalGateConfig, FairnessReportCard
)

# Hierarchical gate: overall → single-attribute → intersectional.
# Name only metrics this gate computes: demographic_parity_difference,
# equalized_odds_difference, false_positive_rate_difference,
# predictive_parity_difference, disparate_impact_ratio. Any other name is
# treated as a check that did not happen, and the gate BLOCKS on it.
gate = ModelFairnessGate(
    metrics=['demographic_parity_difference', 'equalized_odds_difference'],
    thresholds={'demographic_parity_difference': 0.1, 'equalized_odds_difference': 0.1}
)
config = HierarchicalGateConfig(
    check_intersections=True, intersection_depth=2
)
decision = gate.evaluate_hierarchical(
    y_true, y_pred,
    protected_attrs={'gender': gender, 'race': race},  # arrays, keyed by attribute
    hierarchical_config=config
)

# Generate PR-ready report card
card = FairnessReportCard(decision, model_name='loan-approval-v2.1')
print(card.to_markdown())  # Paste into PR comment
output
## Fairness Report Card: loan-approval-v2.1

**Status**: `BLOCKED`
**Result**: Deployment blocked

### Hierarchy Summary

| Level | Status | Failures |
|-------|--------|----------|
| overall | BLOCKED | 2 |
| attr:gender | BLOCKED | 2 |
| attr:race | APPROVED | 0 |
| intersection:gender_x_race | BLOCKED | 2 |

#### overall: Failed Metrics

- **demographic_parity_difference**: 0.2183 (threshold: 0.1000)
- **equalized_odds_difference**: 0.1576 (threshold: 0.1000)

#### intersection:gender_x_race: Failed Metrics

- **demographic_parity_difference**: 0.2916 (threshold: 0.1200)
- **equalized_odds_difference**: 0.2537 (threshold: 0.1200)

Every failed row carries the value it was measured at and the threshold it was compared to, so a reviewer can tell a breach from a check that never ran. A metric the gate could not compute is reported as a blocking reason in words, never as a numeric failure.

python
from vfairness import log_fairness_to_wandb, auto_log_fairness

# W&B experiment tracking
log_fairness_to_wandb(analyzer, prefix='fairness')

# Auto-log decorator: the wrapped function returns (y_pred, y_true, sensitive_attr)
@auto_log_fairness(backend='mlflow')
def evaluate_model(X, y, sensitive):
    model = LogisticRegression().fit(X, y)
    return model.predict(X), y, sensitive  # fairness metrics logged automatically

Experiment tracking needs the [mlops] extra (pip install "vfairness[mlops]" for MLflow and Weights & Biases). Without it, log_fairness_to_wandb raises an ImportError, and the auto_log_fairness decorator warns and skips logging.

See the Workflow Integration Guide for CI/CD YAML templates, pre-commit hooks, pytest plugin setup, and end-to-end examples.

Next Steps