Skip to content
Caramel.
← All products

Product · Foundation models · Banking

Learn the customer once, answer many questions

One representation of the customer, learned once and shared by churn, adoption, fraud and recommendation.

Origin
A leading bank in Argentina
Industry
Banking and financial services
Stage
Feasibility study, with runnable demos

The problem

In a bank, every question about the customer is answered with a pipeline of its own. One team builds features for churn, another for fraud, another for product adoption. Each starts again from the same sources, the same joins and the same counts under a different name, and the next case takes weeks or a quarter to ship.

What gets lost on the way is order. A customer who checks the balance and then transfers is not the same as one who does the reverse, but a monthly count records them identically. The signal is in the sequence, and the sequence is exactly what counts flatten.

The answer

A foundation model of the customer inverts the order of the work. Instead of building features for one question, a backbone is pre-trained on the raw events — type, value and time — with no labels and no decision up front about what it is for. What it learns is a general-purpose representation: one vector per customer that sums up their history.

The answers come out of that vector. Churn, product adoption, propensity to trade dollars, fraud, recommendation: each case is a linear head on the same representation, not a new pipeline. The economics move. The cost sits in the backbone and is paid once; every case after that is marginal, and what remains is installed capability that plugs into the rest of the ecosystem.

The reference is PRAGMA, the foundation model Revolut and NVIDIA trained on 24 billion events. Pragma Criollo evaluates whether that mechanism is viable with the data and the scale of a bank in the region, and what it would take to put it into production.

What we built

An interactive demo centre, built to answer one concrete question from a leading bank in Argentina: is it viable to train a foundation model of its own on its customers' transactional behaviour? It is not a mockup. The backbone is trained and running, the machine learning engine executes in the browser, and the feasibility plan can be walked end to end.

What we proved

3
Encoder levelsProfile, event and history, with temporal attention
96
Dimensions per customerThe vector every use case comes out of
76 %
Masked accuracyChance sits at 0.8 %
<1 s
Adapting a new caseA linear head on the frozen embedding
Holdout AUC on synthetic data.
Task20-feature baselineBackbone
Churn0.7070.677
Product adoption0.7430.710
Propensity to trade dollars0.6930.687
Order of events0.5340.993

On the three linear tasks the backbone matches a baseline of 20 hand-built features, and that parity is the expected ceiling: the synthetic labels are born of the very features the baseline gets for free.

The difference shows up on the fourth, designed so that counts land at chance. The events are the same and the only thing that changes is temporal precedence: the baseline stays at 0.534 and the backbone reaches 0.993. The mechanism captures structure that manual feature engineering cannot see. With only 50 labels, a probe on the frozen embedding already reads that pattern, so the representation transfers.

The technical surface

The integration contract is a single vector per customer. Every new use case is a linear head on that vector, not a new pipeline. These are the three surfaces, and the line between what runs today and what is planned is not blurred.

Implemented

The live engine

pragma-lab.js, in the browser and with no backend

generate(n, seed) → Usuario[]
Generates n reproducible synthetic customers.
trainAll(users, seed) → {models, aucs, rates}
Trains churn, adoption and dollars with a 25 % holdout.
predict(model, feats) → [0, 1]
Probability over a vector of 20 features.
drivers(model, feats, k) → {f, dir}[]
The k most influential factors, for explainability.
pca2(rows) → [x, y][]
2D projection for visualising embeddings.
save() / load() → bool / obj
Persists the model: the bridge from the lab to the assistant.
Implemented

The backbone repository

PyTorch, on the command line

synthetic_generator.py
A key–value–time JSONL corpus, with no PII.
train.py
Pre-training of the backbone with masked modelling.
evaluate.py
A probe on frozen embeddings against the baseline.
predict.py
The 96-dimension vector and the decision heads.

Checks M1–M3 green, a tokenizer with a round-trip test, and a runs/ folder with versioned checkpoints and metrics.

Roadmap

How the bank consumes it

Planned in the roadmap, not delivered yet

An inference and embeddings service, batch and online.
A model registry versioning the backbone and every head.
CI/CD for fine-tuning and deploying downstream heads.
Drift monitoring per use case.
Change propagation: one data change affects N cases.

Honesty by design

Everything runs on synthetic data, without a single real record or any personal information. That fixes what the numbers prove: the mechanism end to end — generate, train, predict, explain — and not business performance. Differential uplift is what the next phase looks for, with the bank's real data.

PRAGMA's published results — +130 % on credit scoring, +79 % on engagement, +67 % on fraud recall, +41 % on recommendation — are Revolut's. They are neither ours nor the bank's: they serve as a reference for magnitude, not as a committed projection. Anti-money-laundering is explicitly out of scope, because it requires looking at relationships between customers and not only at each customer's history.

Saying all this does not weaken the argument; it is what makes a decision possible. A committee that knows what is proven and what is not can approve the next phase with its eyes open.

How we do this at Caramel

  • Advise

    The feasibility study and its gates: 97 tasks across four phases, three Go/No-Go decisions and a verdict that stands on evidence rather than enthusiasm.

  • Build

    The tokenization, the encoders, the backbone, the synthetic corpus, the honest baseline and the demo centre. Feasibility as code: the plan, the data, the model and the executive deliverables all come out of the same repository and are regenerated from it.

  • Run

    What it takes to live in production: model risk management, MLOps, retraining and drift monitoring case by case.

Architect-led · Engineer-built · Production-proven

Stack

Shall we test it on your data?

Request a demo