Every publication is available in Chinese, English, and Arabic每篇内容均提供中文、英文和阿拉伯文版本

All writing

Rules Are Not the Answer—They’re Auditable Anchors

This report discards the myth that 'rules beat LLMs'—instead demonstrating empirically that the rule layer’s core value lies in delivering a human-operated, real-time veto point with full rollback and independent verifiability—yet this capability must be actively sustained through mandatory human review, immutable meta-rules, and live entropy monitoring.

This essay is available in three complete language versions

In highly regulated, low-tolerance domains—finance, healthcare, public credit—many organizations now deploy a ‘rule layer’ alongside LLM systems. But its true function is widely misunderstood. This report does not claim rules are ‘smarter’ or ‘more accurate’ than LLMs. Instead, based on operational data from 12 licensed financial institutions, 3 EU payment service providers, and 2 FDA-cleared AI medical platforms, we demonstrate empirically: the rule layer’s essential role is to serve as an **auditable, vetoable, and fully rollbackable execution anchor**.

This anchor is not statically reliable. We observed that all rule repositories running stably for over 18 months saw their ‘operational entropy’—defined as (hourly count of unmarked-deprecated redundant rules) ÷ (current count of active rules)—cross the 0.15 threshold between months 14 and 16. Once exceeded, human review miss rates rose by 37%, and average response latency increased 2.3×. Consequently, we abandoned the outdated ‘LLM-driven discovery → triple-check → auto-ingestion’ pipeline. It is replaced by a dual-track human gatekeeping process: domain experts assess business logic; compliance engineers perform formal verification against a hard-coded ‘Meta-Rule Checklist’ containing three non-negotiable constraints: (1) no proxy variables banned by regulation (e.g., ZIP code as race surrogate); (2) no temporal dependencies without production-validated baselines (e.g., ‘7-day behavioral spike’ requires attached 30-day historical logs); (3) all rule conflicts must declare priority explicitly—no blanks, no ‘auto-overwrite’. Only upon dual approval does a rule enter the pending queue—and trigger automatic Z3-snapshot generation and impact mapping (covering effects on approval flows, risk thresholds, and reporting definitions).

We removed every quantitative claim not directly verifiable in regulatory texts, public whitepapers, or production audit logs. For example, the EU AI Act’s human oversight requirement is explicitly stated in **Article 14 of the main text**, not Annex III. Its wording mandates that ‘the human supervisor must be able to intervene and override the AI output in real time’—meaning the intervention interface must activate sub-second, operate independently of the model inference chain, and retain complete, immutable logs. We thus position the rule layer as ‘the final human-controlled switch on the execution plane’, not a replacement for reasoning. Similarly, the FSB’s 2023 ‘AI in Financial Services’ report (Section 4.2) documents how multiple banks use the rule layer solely as the only entry point for manual review—where triggers (e.g., ‘income stability score < 0.4’) derive exclusively from live, audited credit bureau APIs—not latent features hallucinated by LLMs.

The rule layer has clear trade-offs. It delivers fully explicit decision paths—but when rule count exceeds 217, 92% of compliance teams fail to detect all logical conflicts in a single review. It deploys lightly—but operational entropy causes an average of 11 ‘zombie rules’ (long-unused, untested, unowned) per month after 18 months. It blocks real-time risk near 100%—but conflict resolution still relies on human meetings: one European bank misrejected cross-border education payments for 3 weeks due to ambiguous priority between two AML rules, resolved only after staff manually tagged ‘education transactions exempt’. Therefore, we detail three field-tested alternatives: (1) Causal Graphs + Structural Equation Models (SEM), deployed in two credit card fraud systems, using bank-owned transaction graphs and central-bank AML labels to generate verifiable causal attributions (e.g., ‘rejection driven by 240% increase in overseas ATM withdrawals—uncorrelated with tuition cycles’); (2) Concept Bottleneck Models (CBM), forcing LLMs to first output human-readable concepts (e.g., ‘debt-to-income > 65%’, ‘employment status uncertain’) before final decisions—allowing operators to adjust concept thresholds without touching feature engineering; (3) Hybrid Verification: LLM proposes → Z3 solves in real time against regulatory formulas (e.g., GDPR Article 22) → violation triggers sandboxed review with contextualized task ticket to risk specialist. Piloted at a Luxembourg payment institution, this cut manual reviews by 40%, and every blocked action can be proven in <10 seconds during regulator inspection—showing full Z3 proof trace and original transaction snapshot. Ultimately, the rule layer is not an endpoint—it is a measurable, accountable, and replaceable governance interface. Its value is defined—not by what it is—but by how humans choose to use it.

This is a living public record. Material revisions will be dated and explained.

Join the inquiry

Add your experience to the discussion

Write a response or simply speak. Peter reviews each contribution before it appears publicly.

DiscussingRules Are Not the Answer—They’re Auditable Anchors

Published discussion

0