When a financial document needs to be approved or rejected, there usually isn't one model making the decision.
In the check / payment-slip systems I worked on, the pipeline looked more like this:
1. Multimodal models + OCR to extract fields
2. YOLO to detect regions and crop signatures
3. Fine-tuned signature verification models, outperforming generic models on bank-specific documents
4. A business-rule engine for decisions the model should not be making
Some lessons were particularly clear:
- A model can read the amount. A rule should verify that the amount in numbers matches the amount in words.
- A model can detect a signature. A specialized model can verify whether it matches an authorized signer.
- Front/back consistency, required fields, thresholds, and institution-specific requirements belong in deterministic logic.
- For critical fields, specialized crops + specialized models can outperform an end-to-end VLM.
And when document formats differ significantly between institutions, having separate rule branches isn't necessarily technical debt. Sometimes it's simply the correct representation of the business.
The interesting part starts when you scale this.
On Databricks, I would treat document understanding, model serving, monitoring, and business rules as one governed production path?not as a notebook demo connected to a webhook.
If you're designing IDP on Databricks / Mosaic AI, I'm interested in how you're drawing the line between what the model should decide and what should remain deterministic.
Frank