One closed loop from strategic behavior to auditable evidence
Three adaptive policies interact inside a constrained AML environment. Executed actions and delayed outcomes update an online world model, then locked held-out tests measure whether the workflow generalizes to unseen behavior shifts.
Transactions are generated under baseline, volume, timing, and topology regimes.
Three PPO actor-critics adapt; the bank also learns penalties for operational constraints.
A hard budget limits simulated reviews; ground-truth labels arrive later.
An action-conditioned GRU ensemble predicts future states and missed-illicit risk.
Paired seeds, multi-step rollouts, constraint checks, and event provenance test the full workflow.
Uniform Online MLE predicts the environment 20 simulated steps ahead more accurately than Static MLE: 0.105 versus 0.148 MAE.
Approved illicit value versus Bank Only: 28.59 fewer synthetic-value units per episode; 95% interval: 15.91 to 41.26.
Lightweight vectorized CPU generator across 25K accounts; not full-workflow or production throughput. 50 automated tests cover model and workflow logic.
Select a held-out behavior regime
Each slice was excluded from training and used only for locked testing. The selector filters stored results; it does not retrain a model.
Full Workflow + Uniform Online World Model
What changes between the four workflows
Fast and interpretable, but cannot adapt when transaction behavior changes.
Respects review capacity, while attacker and customer behavior remain fixed.
Captures strategic interaction, but does not predict future environment states.
Adds an online world model for uncertainty-guided scenario search.
Average approved illicit value per episode ↓
What each number means: the average illicit transaction value that this workflow approved and allowed to execute in one test episode.
Full workflow vs bank only
What this number means: the average reduction in approved illicit value per episode, calculated as Bank Only minus Full Workflow. It is not a percentage.
This is the fixed overall comparison across all held-out regimes; it does not change with the selector above.
Detection recall versus false-positive rate
Each point is one workflow. The upper-left direction is preferable: more illicit activity detected with fewer legitimate transactions flagged.
Uniform online updates reduce prediction error
The key comparison is the deployed Uniform Online MLE against the non-updating Static MLE. Lower error means the model tracks changing behavior more accurately.
Example transaction audit trail
Three representative events from one formal policy-aware ecosystem run (seed 509): a capacity override that preserves a legitimate transaction, an illicit transaction sent to simulated review, and a low-risk legitimate transaction approved under a timing shift. This audit example is separate from the Uniform Online workflow metrics above and does not change with the evaluation selector.
Each event card shows the requested and executed transaction amount in synthetic-value units, risk score, attacker and customer behavior, and the bank action before and after the hard review-capacity rule.
Review is simulated; no human reviewer participates. When the review queue is within capacity, an illicit reviewed event is detected with the configured 78% probability. If proposed reviews exceed capacity, lower-priority reviews fall back to approval; the capacity rule does not convert reviews into declines.
This is a synthetic/shadow evaluation system.
Public AMLSim marginals do not identify causal effects of bank actions, customer responses, or attacker adaptation. Production use requires institutional action logs, temporal replay, off-policy evaluation, monitoring, human review, and model-risk approval.