Small, single-purpose applications are spreading quickly across financial institutions and the authorities that supervise them. A tool that reconciles safeguarded balances, or ages inspection evidence, or decomposes a provision movement, can now be built in days rather than quarters. The obvious next question is how much artificial intelligence should sit inside them. The answer is less than the enthusiasm suggests, and in a narrower place than most designs assume.
Why the question is different for micro apps
Enterprise model governance was built for a small number of large, slow-moving models: credit scorecards, capital models, pricing engines. Each is validated deliberately because each is expensive and consequential. Micro apps invert every one of those properties. They are numerous, cheap, quick to build and easy to copy. That is precisely what makes them useful, and precisely what makes an unexamined AI component dangerous. A single unvalidated model in a capital calculation is a known problem with a known remedy. Forty small tools, each quietly calling a model to produce a number that someone treats as a finding, is a governance problem that nobody has assigned an owner to.
The proliferation risk that has always attended end-user computing does not disappear when the spreadsheet becomes a web application. It changes shape. Adding generative capability to that population without a boundary rule multiplies it.
The boundary rule
The rule that holds up in practice is simple to state. Artificial intelligence is used to read, not to rule.
It converts unstructured material into structured, checkable facts. It does not decide what those facts mean. The logic that turns facts into a rating, a classification, a breach determination or a penalty band stays deterministic, written down and inspectable.
Three reasons make this more than a stylistic preference. A supervisory or risk conclusion must be reproducible when it is re-run months later during a challenge, and a generative model drifts silently across versions. It must be explainable to the firm or the committee that receives it, and “the model produced this” is not an explanation. And it must be owned by a named person, which it cannot be if the reasoning happened somewhere nobody can cross-examine.
Applied consistently, this rule also settles the model-risk question that otherwise consumes disproportionate effort. A tool that extracts and organises is not a model in the governance sense. A tool that scores is, whether or not any machine learning is involved. Validation effort should follow that line, not the presence or absence of the word “AI” in the design document.
Where it genuinely helps
Within that boundary, four uses earn their place. Extraction pulls fields from credit files, application forms and returns, with each extracted value shown alongside the source it came from. Classification assigns pre-defined categories, chosen by the analyst rather than invented by the model, to complaints, incidents and exception rationales. Clustering surfaces candidate patterns across a population for a person to test. Drafting produces first-pass text that the author rewrites and owns.
What these share is that a human can check the output against a source in seconds. That is the practical test for whether an AI component belongs in a micro app: if verification takes longer than doing the work manually, the component is decorative.
An earlier stocktake of prudential supervisory technology found that more than half of the tools surveyed analysed mainly qualitative information (Beerman, Prenio & Zamil, 2021). This is where AI matters most in micro apps, and it is a useful corrective to the assumption that these tools are essentially numerical. The bottleneck in most review processes is not calculation. It is reading.
Deployment, and the constraint that helps
Confidential supervisory and customer information imposes a real limit on where inference can run. Three patterns work. Some tools need no model at all, and deterministic text handling over structured extracts is sufficient for most reconciliation, ageing and register work. Others can run a small model locally, inside the institutional perimeter, accepting reduced capability in exchange for a clean data boundary. The remainder can be human-mediated, with the analyst processing material in a separate approved environment and bringing the structured result back into the offline tool.
The choice should be made and recorded per tool, not settled once for a platform and assumed thereafter.
Two controls make the rest defensible. Log the input, the task template, the model and version, the raw output, the edits made and the person who accepted it. And store confirmed output as data rather than regenerating it when the tool reopens. An application that produces different text each time it is opened cannot support a conclusion.
Micro apps are valuable because they sit close to a specific decision and can be changed quickly when the decision changes. That closeness is exactly why the reasoning inside them should be visible. AI expands what these tools can take as input, which is a genuine gain when the input is a folder of documents rather than a table of numbers. It should not be allowed to expand what they conclude.The output remains an input to judgement, and the judgement remains the analyst’s.
References
Beerman, K., Prenio, J., & Zamil, R. (2021). Suptech tools for prudential supervision and their use during the pandemic (FSI Insights on policy implementation No. 37). Bank for International Settlements.
Prenio, J. (2024). Peering through the hype: Assessing suptech tools’ transition from experimentation to supervision (FSI Insights on policy implementation). Bank for International Settlements.




Leave a Reply