AnyFluxion

Cybersecurity in Finance

A synthetic/shadow testing system for measuring how bank decisions perform when attackers adapt, customers respond, labels arrive late, and review capacity is finite.

How the system works

One closed loop from strategic behavior to auditable evidence

Three adaptive policies interact inside a constrained AML environment. Executed actions and delayed outcomes update an online world model, then locked held-out tests measure whether the workflow generalizes to unseen behavior shifts.

Synthetic/shadow testingNo live bank decisionsNo human reviewer in this simulation
Long-range prediction
29% lower error

Uniform Online MLE predicts the environment 20 simulated steps ahead more accurately than Static MLE: 0.105 versus 0.148 MAE.

Decision quality
1.73% less

Approved illicit value versus Bank Only: 28.59 fewer synthetic-value units per episode; 95% interval: 15.91 to 41.26.

Simulator throughput
~245K events/s

Lightweight vectorized CPU generator across 25K accounts; not full-workflow or production throughput. 50 automated tests cover model and workflow logic.

Explore the formal results

Select a held-out behavior regime

Each slice was excluded from training and used only for locked testing. The selector filters stored results; it does not retrain a model.

Held-out regimeA behavior shift excluded from training and used only for evaluation.
H=20 rolloutA prediction made 20 simulated steps into the future.
Uniform Online MLEThe world model is updated continuously from collected observations.
Constraint pass rateShare of runs meeting miss, false-positive, review, and latency limits.
Selected workflow

Full Workflow + Uniform Online World Model

Constraint compliance
Recall ↑
—
Share of illicit events detected
False-positive rate ↓
—
Legitimate events incorrectly flagged
Approved illicit value / episode ↓
—
Illicit value allowed to execute
Workflow variants

What changes between the four workflows

Static ThresholdFixed decision rules

Fast and interpretable, but cannot adapt when transaction behavior changes.

Budgeted Bank OnlyLearns bank decisions

Respects review capacity, while attacker and customer behavior remain fixed.

Multi-agent, No World ModelAll three policies adapt

Captures strategic interaction, but does not predict future environment states.

Full Workflow + Uniform Online World ModelAdaptation plus future-state prediction

Adds an online world model for uncertainty-guided scenario search.

Decision workflows

Average approved illicit value per episode ↓

What each number means: the average illicit transaction value that this workflow approved and allowed to execute in one test episode.

Overall paired result

Full workflow vs bank only

What this number means: the average reduction in approved illicit value per episode, calculated as Bank Only minus Full Workflow. It is not a percentage.

This is the fixed overall comparison across all held-out regimes; it does not change with the selector above.

Overall result: the Full Workflow reduces approved illicit value by 1.73% versus the operational Bank Only baseline and delivers the highest recall with the lowest false-positive rate among the learned workflows. Results vary by held-out regime, so the evidence supports this measured trade-off rather than claiming one method wins every metric.
Operating trade-off

Detection recall versus false-positive rate

Each point is one workflow. The upper-left direction is preferable: more illicit activity detected with fewer legitimate transactions flagged.

World-model evidence

Uniform online updates reduce prediction error

The key comparison is the deployed Uniform Online MLE against the non-updating Static MLE. Lower error means the model tracks changing behavior more accurately.

H=20 rollout error—vs. static model
State RMSE—vs. static model
Miss-rate MAE—vs. static model
Auditability

Example transaction audit trail

Three representative events from one formal policy-aware ecosystem run (seed 509): a capacity override that preserves a legitimate transaction, an illicit transaction sent to simulated review, and a low-risk legitimate transaction approved under a timing shift. This audit example is separate from the Uniform Online workflow metrics above and does not change with the evaluation selector.

Each event card shows the requested and executed transaction amount in synthetic-value units, risk score, attacker and customer behavior, and the bank action before and after the hard review-capacity rule.

Review is simulated; no human reviewer participates. When the review queue is within capacity, an illicit reviewed event is detected with the configured 78% probability. If proposed reviews exceed capacity, lower-priority reviews fall back to approval; the capacity rule does not convert reviews into declines.

Requested transaction amountExecuted transaction amountAmounts use synthetic-value units; risk bars use green / amber / red bands.
Reality boundary

This is a synthetic/shadow evaluation system.

Public AMLSim marginals do not identify causal effects of bank actions, customer responses, or attacker adaptation. Production use requires institutional action logs, temporal replay, off-policy evaluation, monitoring, human review, and model-risk approval.