Start · Getting started
Getting Started with vfairness
A full-pipeline Python library for measuring, mitigating, and monitoring fairness in machine learning systems.
vfairness covers the entire ML fairness lifecycle: from pre-training bias auditing and fairness-aware model training, through group-specific calibration and post-processing, to production monitoring, stakeholder reporting, and controlled experimentation. Every stage produces auditable artifacts (SVG visualizations, structured reports, and machine-readable metrics) designed for regulatory environments including the EU AI Act.
1,558 public functions, classes and methods across
204 modules.
A capability is one dispatch key in the registry, not a fourth kind of object:
the 215 capability rows name
222 of those same functions, classes and
methods, so the capability tile is not one to add to the other three.
One module, vfairness.mcp.server, could not be imported (ImportError), so whatever it defines is not in these counts.
An uncounted module understates the surface, so it is reported rather than dropped.
How much of this surface has been tested is reported on
Quality and Hardening, in these same four nouns.
Installation
# Core (numpy, pandas, scipy, scikit-learn, requests)
pip install vfairness
# With visualization support (matplotlib, seaborn)
pip install "vfairness[viz]"
# With interactive dashboards (Plotly + Dash)
pip install "vfairness[interactive]"
# Plotly charts only
pip install "vfairness[dashboard]"
# Experiment tracking (MLflow, Weights & Biases)
pip install "vfairness[mlops]"
# Full installation (all optional dependencies)
pip install "vfairness[all]"
0.1.0 was published on 2026-10-02 and installs with pip install vfairness. The version number says beta, and the beta gate agreed when it was cut: scripts/release_gate.py reports BETA READY, with all eight beta conditions met. The stricter 1.0 gate is not met yet: a second, independent examiner still has to confirm part of the graded code. The API is not frozen until 1.0.0. See the Releasing for how versions will ship, and Quality and Hardening for what has been graded, what has not, and which defects are still open.
import vfairness
print(f"vfairness {vfairness.__version__}")
Library Structure
vfairness is organized into 15 top-level sub-packages: a 6-stage fairness pipeline (preprocessing, in_processing, post_processing, evaluation, operations, rendering), 8 specialized analysis surfaces (llm, agents, multi_agent, xai, vision, legal, mcp, validity) and 1 cross-cutting infrastructure package (net, the egress guard). The areas below walk that pipeline from data to production.
vfairness.preprocessingvfairness.in_processingvfairness.post_processingvfairness.evaluationvfairness.operations.monitoringvfairness.operations.reportingvfairness.operations.experimentationvfairness.operations.cicdvfairness.llmvfairness.agentsvfairness.multi_agentnotebooks/vfairness_0_library_validation.ipynb validates the core functions across preprocessing, training, calibration, evaluation and operations. It ships with the source in the public repository.
Step 0: Find Something to Test
Every analysis starts with a test object: a table of decisions, a model, a chat endpoint, an agent or a team of agents. How deep the analysis can go depends on what that object lets you see, not on how you connect to it. Decisions alone support claims about outcomes; probabilities add calibration; only the model itself supports claims about why.
| You are testing | With outputs only (A1) | With scores (A2) | Starter file |
|---|---|---|---|
| A predictive model | decisions + groups (+ labels for error rates) | + probabilities: calibration | Adult with predictions |
| An LLM | prompts in, text out | + log probabilities: likelihood bias | discrim-eval answers |
| An agent | outcomes, plus tool calls if you have traces | + log probabilities per step | loan agent runs |
| Several agents | each agent's output, not only the last | as for agents | loan committee runs |
| Your detector itself | synthetic data with a planted bias and a clean control | planted and clean | |
Find a test object explains every pathway, the data each depth needs, and links each file to its source. Every file there was run through vfairness before it was published.
Quick Start
Measure fairness in a loan approval model in under 30 lines. The output below is what this script prints, captured from a run against 0.1.0.
import numpy as np
from vfairness import FairnessAnalyzer, demographic_parity_difference
# Sample data
y_true = np.array([1, 0, 1, 0, 1, 0, 1, 0, 1, 1] * 100)
y_pred = np.array([1, 0, 1, 1, 1, 0, 0, 0, 1, 1] * 100)
gender = np.array(['M','M','M','M','M','F','F','F','F','F'] * 100)
# Single metric
dp = demographic_parity_difference(y_true, y_pred, gender)
print(f"Demographic Parity Difference: {dp:.3f}")
# Full analyzer with a confidence interval.
# random_state pins the bootstrap, so the interval is reproducible.
analyzer = FairnessAnalyzer(y_true, y_pred, gender, min_group_size=30)
result = analyzer.demographic_parity_difference(
include_ci=True, n_bootstrap=5000, confidence_level=0.95, random_state=0
)
lo, hi = result.confidence_interval
print(f"DP Diff: {result.value:.3f} 95% CI: [{lo:.3f}, {hi:.3f}]")
# All metrics at once
for name, value in analyzer.compute_all_metrics().items():
print(f" {name}: {value:.3f}")
# The verdict, and the rows it was computed over
print(analyzer.get_report()["assessment"]["summary"])
Demographic Parity Difference: 0.400
DP Diff: 0.400 95% CI: [0.346, 0.456]
demographic_parity_difference: 0.400
equalized_odds_difference: 0.500
equal_opportunity_difference: 0.333
demographic_parity_ratio: 0.500
predictive_parity_difference: 0.250
0/5 metrics within thresholds (data provenance: 1000 of 1000 rows assessed, 0 excluded for missing values, 0 in 0 group(s) withheld by the group-size floor, missing_strategy='exclude')
The data provenance clause is part of every summary. It states how many rows reached the metrics and how many were dropped on the way, so a verdict can never be read without knowing what it was computed over.
get_report(include_ci=True) fills metrics_with_ci for three of the five default metrics: demographic_parity_difference, equalized_odds_difference and equal_opportunity_difference. demographic_parity_ratio and predictive_parity_difference come back as plain point estimates there, with no interval and no interval-derived verdict.
disparate_impact_ratio_with_ci and predictive_parity_difference_with_ci directly. This matters most for the four-fifths ratio, which is the statistic a regulator is most likely to ask about.
Fairness metrics are mathematically incompatible: you cannot satisfy demographic parity, equalized odds, and calibration simultaneously (Kleinberg et al., 2016). Decide which definition matches your use case before running analysis. See The Impossibility Theorem.
Three States, Not Two
Sooner or later a run will tell you it could not check something. That is not a failure of the library and it is not a fairness violation. It is the third state, and it is deliberate: a group or a metric that was never measured gets no number, no verdict, no colour, and no place in any count that implies it was measured. Reading these outputs is the one thing worth learning before anything else on this page.
- Assessed, within tolerance: measured, and the evidence supports the pass.
- Assessed, exceeds tolerance: measured, and the evidence supports the breach.
- Could not check: not measured. Never folded into either of the other two.
A group too small to assess
Add 12 nonbinary applicants to the 1,000 rows above. They are below the default min_group_size of 30, so they cannot carry a verdict. They are not dropped in silence either: a UserWarning names them, and they appear in the report as insufficient evidence.
import numpy as np
from vfairness import FairnessAnalyzer
# The same 1,000 rows as above, plus 12 nonbinary applicants.
y_true = np.array([1, 0, 1, 0, 1, 0, 1, 0, 1, 1] * 100)
y_pred = np.array([1, 0, 1, 1, 1, 0, 0, 0, 1, 1] * 100)
gender = np.array(['M','M','M','M','M','F','F','F','F','F'] * 100)
y_true = np.append(y_true, [1, 0] * 6)
y_pred = np.append(y_pred, [1, 1] * 6)
gender = np.append(gender, ['NB'] * 12)
assessment = FairnessAnalyzer(y_true, y_pred, gender).get_report()["assessment"]
print(assessment["summary"])
group = assessment["insufficient_evidence_groups"][0]
print(f"{group['group']}: n={group['n']}, verdict={group['verdict']}")
print(group["reason"])
UserWarning: Excluding 1 group(s) below min_group_size=30: {'NB': 12}. That leaves
12 of 1012 rows (1.2 percent) out of this metric. The metric is computed over the
remaining groups only, so it measures nothing about the excluded group(s): they are
could not check, not a measured pass. To include them, lower min_group_size.
0/5 metrics within thresholds (1 group(s) with insufficient evidence, excluded from the verdict: ['NB']) (data provenance: 1012 of 1012 rows assessed, 0 excluded for missing values, 12 in 1 group(s) withheld by the group-size floor, missing_strategy='exclude')
NB: n=12, verdict=insufficient_evidence
n=12 is below min_group_size=30; this group is excluded from the disparity metrics, so there is insufficient evidence to assess it. Underpowered (n<30): the rate is noisy; treat as indicative only.
Note what the summary does not say. It does not report 5/5, and it does not report a violation for NB. It reports the count over the metrics it graded, and it names the group it could not reach.
A run with nothing left to compare
A disparity metric needs at least two groups. When filtering leaves only one, there is no comparison to make, so the whole report is marked not assessable and the score is None rather than a number.
import numpy as np
from vfairness import FairnessAnalyzer
# 400 men and 3 women: only one group is large enough to measure, so there is
# nothing to compare it against.
gender = np.array(['M'] * 400 + ['F'] * 3)
y_true = np.append(np.random.default_rng(1).integers(0, 2, 400), [1, 0, 1])
y_pred = np.append(np.random.default_rng(2).integers(0, 2, 400), [1, 1, 0])
assessment = FairnessAnalyzer(y_true, y_pred, gender).get_report()["assessment"]
print("assessable: ", assessment["assessable"])
print("fairness_score:", assessment["fairness_score"])
print("not assessable:", [m["metric"] for m in assessment["not_assessable_metrics"]])
print(assessment["summary"])
assessable: False
fairness_score: None
not assessable: ['demographic_parity_difference', 'demographic_parity_ratio', 'equalized_odds_difference', 'equal_opportunity_difference', 'predictive_parity_difference']
NOT ASSESSABLE: 1 valid group(s) after filtering (need at least 2): disparity metrics are vacuous and do not certify fairness. (1 group(s) with insufficient evidence, excluded from the verdict: ['F']) (5 metric(s) not assessable, excluded from the verdict) (data provenance: 403 of 403 rows assessed, 0 excluded for missing values, 3 in 1 group(s) withheld by the group-size floor, missing_strategy='exclude')
fairness_score is None, not 0.0. A zero would be
read as a measured, maximally unfair result, which is the opposite of what happened
here. The same rule governs the release gates: assert_fairness refuses a
metric whose value is NaN with "NOT MEASURABLE ... so the metric was never
compared against threshold (fail closed)", and ModelFairnessGate
blocks rather than approving a check that never ran. The charts follow it too: an
empty reliability diagram renders "NOT ASSESSABLE: no data was supplied, so
nothing was measured" and prints its ECE as N/A, instead of a
reassuring 0.000.
Production Features
Every test result carries audit-trail metadata: an ISO 8601 UTC timestamp, the library version that produced it, and the parameters it ran with. That is the raw material for a record-keeping obligation such as EU AI Act Art. 12; whether a given deployment satisfies that obligation is a question for its operator, not something a library can establish.
from vfairness.llm import OutputAnalyzer
analyzer = OutputAnalyzer(alpha=0.05)
result = analyzer.analyze_sentiment(texts_a, texts_b)
# Every result has audit metadata
print(result.metadata.timestamp) # ISO 8601 UTC
print(result.metadata.library_version) # current vfairness.__version__
print(result.metadata.parameters) # {"alpha": 0.05, "metric": "sentiment", ...}
# Serialize for database storage
data = result.to_dict() # Plain dict
json_str = result.to_json() # JSON string
# Structured logging (configure once)
import logging
logging.basicConfig(level=logging.INFO)
analyzer.analyze_all(texts_a, texts_b)
INFO:vfairness.llm.output_analysis:analyze_all: group_a=group_a (18 samples), group_b=group_b (18 samples), correction=benjamini_hochberg
WARNING:vfairness.llm.output_analysis:Sample size (18, 18) below recommended minimum of 25 for metric 'semantic_quality'
WARNING:vfairness.llm.output_analysis:Sample size (18, 18) below recommended minimum of 25 for metric 'sentiment'
...
INFO:vfairness.llm.output_analysis:LLM judge scorer not configured, skipping analyze_llm_judge
INFO:vfairness.llm.output_analysis:analyze_all complete: 11 metrics evaluated
The sample-size warning is emitted once per metric, so a thin run announces the weakness of every number it is about to hand you.
Production Scorers
The sentiment and toxicity scorers delegate to third-party packages that ship in the [llm] extra: pip install "vfairness[llm]". The rest are implemented in vfairness itself, and the counts below are read straight off those implementations.
| Metric | Scorer | Where it comes from | Details |
|---|---|---|---|
| Sentiment | VADER | [llm] extra | Delegated to vaderSentiment. Lexicon-based, with negation and intensity handling. |
| Toxicity | alt-profanity-check | [llm] extra | Delegated to alt-profanity-check, a linear-SVM classifier. Its accuracy is the upstream project's to report, not ours. |
| Refusal | Pattern-based | in vfairness | 63 patterns in 5 categories (hard, soft, partial, conditional, policy), weighted scoring. |
| Helpfulness | Multi-signal heuristic | in vfairness | 6 quality signals: length, vocabulary, structure, specificity, engagement, deflection. |
| Stereotype | Curated word lists | in vfairness | 61 terms plus 14 phrase patterns (gender, racial, age, religious). |
If vaderSentiment is not installed, the sentiment scorer falls back to a 30-word keyword list and raises a PlaceholderScorerWarning saying so: "Using keyword-based sentiment scorer (30 words only). This is a LOW-ACCURACY placeholder." The warning is the signal that a number came from the fallback and not from VADER. Do not treat a scored run as a production result until the extra is installed and the warning is gone.
Progress Callbacks
Long-running batch operations support progress callbacks for UI integration:
from tqdm import tqdm
from vfairness.llm import LLMApiProxy
proxy = LLMApiProxy(endpoint_url=..., api_format="openai", model_name=...)
# With tqdm progress bar
pbar = tqdm(total=100)
proxy.send_batch(
prompts=my_prompts,
n_runs=25,
progress_callback=lambda current, total: pbar.update(1)
)
pbar.close()
Bias Detection
Audit training data for representation bias, proxy variables, and historical discrimination patterns before any model is trained.
import pandas as pd
from vfairness import BiasDetector, detect_historical_patterns, identify_proxy_variables
# Full audit
detector = BiasDetector(
df,
protected_attributes=['gender', 'age'],
outcome_column='approved',
benchmarks={'gender': {'M': 0.49, 'F': 0.51}}
)
report = detector.full_audit()
print(report.summary())
# Historical discrimination patterns (scans all column names)
for r in detect_historical_patterns(df):
print(f" {r.feature}: {r.risk_level.value} - {r.pattern_type}")
# Proxy variable detection
for p in identify_proxy_variables(df, protected_attributes=['gender'], correlation_threshold=0.3):
print(f" {p.feature} -> {p.protected_attribute} (r={p.correlation:.3f})")
On a 600-row frame with a zip_code column, the historical scan reports:
Overall Risk Score: 50.0%
SUMMARY OF FINDINGS
Historical Pattern Issues: 1
Representation Bias Issues: 2
Statistical Disparities: 0
Proxy Variables Identified: 0
TOP RECOMMENDATIONS
1. CRITICAL: Review and consider removing high-risk features: zip_code
...
zip_code: high - Geographic Redlining
Pass a benchmarks entry for every protected attribute. If one is missing the detector says so out loud, with a UserWarning naming the attribute and the fallback it used: "No benchmark provided for 'age'; falling back to the built-in approximate US Census 2020 distribution." A representation finding against a default benchmark is a finding about US Census 2020, not about your population.
Calibration
Ensure that a predicted probability of 70% means the same risk for every demographic group. Detect and correct calibration disparities.
from vfairness.post_processing.calibration import CalibrationAnalyzer
analyzer = CalibrationAnalyzer(
y_true=labels, y_prob=model_probabilities,
protected_attr=gender, attribute_name='gender'
)
report = analyzer.full_analysis(context='lending')
print(f"Well Calibrated: {report.is_well_calibrated}")
print(f"Significant Disparity: {report.has_significant_disparity}")
print(f"Overall ECE: {report.overall_metrics['ece']:.4f}")
# Apply group-specific recalibration
analyzer.fit_calibrator(method='isotonic')
calibrated_probs = analyzer.transform(y_prob_test, gender_test)
Explanations
Generate context-aware, stakeholder-ready explanations for every metric and audit result.
from vfairness import FairnessAnalyzer
# FairExplAIner mode for per-metric explanations
analyzer = FairnessAnalyzer(y_true, y_pred, gender, fair_explainer=True)
report = analyzer.get_report(include_ci=True)
for name, expl in report['explanations']['metrics'].items():
print(f"{name}: {expl['severity']} - {expl['evaluation']}")
# FairnessExplainer for cross-module explanations
from vfairness.explainer import FairnessExplainer
explanation = FairnessExplainer.explain(audit_report) # Takes a vfairness report object (BiasAuditReport, CalibrationReport, GateDecision, ...)
print(explanation.severity) # 'info' | 'low' | 'medium' | 'high' | 'critical'
print(explanation.recommendations) # Actionable next steps
Visualization
A library of SVG templates for card-based dashboards, plus Matplotlib and Plotly outputs. SVG rendering needs the [rendering] extra (Jinja2); the output SVG is self-contained. Call vfairness.rendering.list_templates() for the current set. As of 2026-08-28 it returns 46 names: 44 renderable charts plus two internal partials, _shared_defs and _could_not_check. The second one is the third state made visible: it is the panel a chart renders in place of a plot when it measured nothing.
from vfairness import classification_fairness_report, plot_fairness_metrics, create_fairness_dashboard
from vfairness.rendering import radar_chart_to_svg, fairness_detailed_report_to_svg
report = classification_fairness_report(y_true, y_pred, gender, include_ci=True)
# Matplotlib (static): returns a Matplotlib Axes
ax = plot_fairness_metrics(report, style='academic', show_thresholds=True)
ax.figure.savefig('fairness_metrics.png', dpi=300, bbox_inches='tight')
# Plotly (interactive)
dashboard = create_fairness_dashboard(report, style='modern')
dashboard.write_html('fairness_dashboard.html')
# SVG templates (needs the [rendering] extra: Jinja2)
svg = radar_chart_to_svg(metrics_data, save_path='radar.svg')
plot_fairness_metrics needs the [viz] extra (Matplotlib); create_fairness_dashboard needs the [dashboard] extra (Plotly); the SVG templates need the [rendering] extra (Jinja2). Unlike the Matplotlib and Plotly charts, the SVG templates render to self-contained SVG with no runtime, browser, or JavaScript dependency.
See the SVG Gallery for every template with live previews.
Monitoring
Track fairness in production with real-time metrics, multi-scale drift detection, and adaptive alerting.
import pandas as pd
from vfairness.operations.monitoring import FairnessMonitor, FairnessDriftDetector
# Real-time monitoring: each batch is a DataFrame with
# 'prediction', 'label', and one 'group_<attr>' column per protected attribute
batch_df = pd.DataFrame({
'prediction': predictions,
'label': labels,
'group_gender': gender,
})
monitor = FairnessMonitor(window_size=1000)
window = monitor.update_and_check(batch_df)
if window.any_alert:
print(f"Alert! Flagged metrics: {[k for k, v in window.alerts.items() if v]}")
# Multi-scale drift detection: set a known-good baseline once,
# then compare each new stretch of metric history against it
detector = FairnessDriftDetector()
detector.set_baseline(pd.Series(reference_metrics))
drift = detector.check_drift(pd.Series(current_metrics), metric="demographic_parity")
print(f"Drift detected: {drift.drift_detected} (score: {drift.overall_drift_score:.3f})")
Reporting
Transform metrics into stakeholder-ready intelligence with multi-tier reports and interactive dashboards.
from vfairness.operations.reporting import MetricsStore, FairnessDashboard, ReportGenerator
store = MetricsStore()
store.ingest_from_monitor(monitor)
# Interactive Plotly dashboard
dashboard = FairnessDashboard(store)
fig = dashboard.create_executive_view()
fig.write_html("dashboard.html")
# Automated multi-format report (HTML, PDF, JSON)
gen = ReportGenerator(store, dashboard)
report = gen.generate_executive_report()
report.save("executive_report.html")
create_executive_view needs Plotly (the [dashboard] extra); MetricsStore and ReportGenerator run on the core install.
Experimentation
Run fairness-aware A/B tests with intersectional analysis and automated deployment recommendations.
from vfairness.operations.experimentation import (
FairnessExperiment, FairnessPowerAnalyzer, ExperimentAnalysis
)
exp = FairnessExperiment(
control_data=df_ctrl, treatment_data=df_treat,
protected_attributes=['gender', 'race'], outcome_column='approved',
)
result = exp.run_full_analysis()
# Per-intersection power analysis
power = FairnessPowerAnalyzer(exp)
print(power.get_power_summary())
# Multi-objective analysis and deployment decision
analysis = ExperimentAnalysis(result, experiment=exp)
rec = analysis.decision_recommendation()
print(f"Decision: {rec.decision}")
Decision: RecommendationDecision.EXTEND_EXPERIMENT
rec.decision is a RecommendationDecision enum with four members: DEPLOY_TREATMENT, KEEP_CONTROL, EXTEND_EXPERIMENT and INVESTIGATE_FURTHER. The last two are the three-state discipline in experiment form: an experiment that has not yet separated the arms returns "keep running", never a deploy recommendation resting on an underpowered comparison. FairnessPowerAnalyzer reports is_powered per intersection so you can see which cell is holding the decision open.
Workflow Integration
Embed fairness checks into every stage of the development lifecycle, from experiment tracking and version control to CI/CD gates and collaborative review.
from vfairness.operations.cicd import (
ModelFairnessGate, HierarchicalGateConfig, FairnessReportCard
)
# Hierarchical gate: overall → single-attribute → intersectional.
# Name only metrics this gate computes: demographic_parity_difference,
# equalized_odds_difference, false_positive_rate_difference,
# predictive_parity_difference, disparate_impact_ratio. Any other name is
# treated as a check that did not happen, and the gate BLOCKS on it.
gate = ModelFairnessGate(
metrics=['demographic_parity_difference', 'equalized_odds_difference'],
thresholds={'demographic_parity_difference': 0.1, 'equalized_odds_difference': 0.1}
)
config = HierarchicalGateConfig(
check_intersections=True, intersection_depth=2
)
decision = gate.evaluate_hierarchical(
y_true, y_pred,
protected_attrs={'gender': gender, 'race': race}, # arrays, keyed by attribute
hierarchical_config=config
)
# Generate PR-ready report card
card = FairnessReportCard(decision, model_name='loan-approval-v2.1')
print(card.to_markdown()) # Paste into PR comment
## Fairness Report Card: loan-approval-v2.1
**Status**: `BLOCKED`
**Result**: Deployment blocked
### Hierarchy Summary
| Level | Status | Failures |
|-------|--------|----------|
| overall | BLOCKED | 2 |
| attr:gender | BLOCKED | 2 |
| attr:race | APPROVED | 0 |
| intersection:gender_x_race | BLOCKED | 2 |
#### overall: Failed Metrics
- **demographic_parity_difference**: 0.2183 (threshold: 0.1000)
- **equalized_odds_difference**: 0.1576 (threshold: 0.1000)
#### intersection:gender_x_race: Failed Metrics
- **demographic_parity_difference**: 0.2916 (threshold: 0.1200)
- **equalized_odds_difference**: 0.2537 (threshold: 0.1200)
Every failed row carries the value it was measured at and the threshold it was compared to, so a reviewer can tell a breach from a check that never ran. A metric the gate could not compute is reported as a blocking reason in words, never as a numeric failure.
from vfairness import log_fairness_to_wandb, auto_log_fairness
# W&B experiment tracking
log_fairness_to_wandb(analyzer, prefix='fairness')
# Auto-log decorator: the wrapped function returns (y_pred, y_true, sensitive_attr)
@auto_log_fairness(backend='mlflow')
def evaluate_model(X, y, sensitive):
model = LogisticRegression().fit(X, y)
return model.predict(X), y, sensitive # fairness metrics logged automatically
Experiment tracking needs the [mlops] extra (pip install "vfairness[mlops]" for MLflow and Weights & Biases). Without it, log_fairness_to_wandb raises an ImportError, and the auto_log_fairness decorator warns and skips logging.
See the Workflow Integration Guide for CI/CD YAML templates, pre-commit hooks, pytest plugin setup, and end-to-end examples.
Next Steps
API Reference
Reference for the library’s classes, functions, parameters and return types. It does not yet cover every public code unit.
Business Guide
Regulatory context, decision frameworks, and EU AI Act compliance guidance.
Concepts
Fairness definitions, the impossibility theorem, causal fairness, and statistical methodology.
SVG Gallery
Browse the full set of visualization templates with live previews and code examples.
Workflow Integration
CI/CD gates, pytest plugin, pre-commit hooks, MLOps logging, and PR report cards.