One toolkit, two review workflows
Synthetic demo of the Responsible AI Toolkit. Generated from commit local run on 2026-09-27 00:38 UTC.
All data here is synthetic, generated with fixed seeds. These pages show what the software did on that data. They are not a model validation, a deployment, or evidence of use by any institution.
The same audit, policy, and human-review components handle consumer lending decisions at a community bank and property insurance quotes at a small insurer. Only the configuration differs: the data, rule settings, and reviewer roles.
Consumer lending decisions
12 model recommendations. 3 without a configured concern, 9 routed to a reviewer, who overrode the model in no cases.
Small-business property insurance quotes
12 model recommendations. 3 without a configured concern, 9 routed to a reviewer, who overrode the model in no cases.
Ongoing model monitoring
Month-by-month drift monitoring
The lending model's scores over six months, compared with its validation period. 1 month crossed the escalation limit and went to a model-risk reviewer.
What stayed the same and what changed
Built from each run's configuration and results. All runs use the same toolkit code.
| Consumer lending decisions | Small-business property insurance quotes | |
|---|---|---|
| Audit logging | AuditLogger, unchanged | AuditLogger, unchanged |
| Policy checks | PolicyEngine, unchanged; 2 built-in rules | PolicyEngine, unchanged; 2 built-in rules and 1 custom rule |
| Human review | HITLOrchestrator, unchanged; reviewers with role “lending” | HITLOrchestrator, unchanged; reviewers with role “insurance” |
| Decision support | Shared Python assess; versioned domain profile; independent of confidence | Shared Python assess; versioned domain profile; independent of confidence |
| Confidence threshold | 0.60 | 0.70 |
| Required inputs | debt_to_income, credit_history_months | prior_claims, last_inspection_year |
| Outcome | 3 without a configured concern, 9 routed for review, 6 still open; no customer outcome changed | 3 without a configured concern, 9 routed for review, 5 still open; no customer outcome changed |
| Reviewer controls | 4 of 4 as expected | 4 of 4 as expected |
| Audit chain | 51 entries verified; edit detected | 64 entries verified; edit detected |
Scenario-specific code in this demo is limited to the synthetic data, the rule settings, one custom rule per scenario where needed (written with the toolkit's Policy class), the reviewer roles, and a fixed rule standing in for reviewer judgment.
These are illustrative configurations on synthetic data. Running the same components under different settings shows how they are configured. It is not evidence of the effort needed to adapt them to a real institution's models, data, and procedures, which requires a separate, documented evaluation.
Reproduce these runs
python -m pip install -e ".[dev]" python examples/model_review_demo.py demo_output
The decisions, controls, and audit results match on every run. Timestamps, entry IDs, and hashes differ.