When people ask why we chose underwriting as the problem to solve, the answer starts with understanding what AI is actually good at. Underwriting is a partial art form built on top of guidelines that are reasonably well defined but still leave real ambiguity, and some of that ambiguity can genuinely be resolved by looking at historical data. What’s interesting is that underwriting isn’t confined to one player in the loan lifecycle. A correspondent buying and onboarding closed loans is essentially re-underwriting against the original file. A retail lender is underwriting from origination forward. Even a company buying loans for servicing is assessing risk in a way that functions as its own form of underwriting. It’s this dense, text-heavy task that recurs constantly across the entire lifecycle of a loan, which makes it uniquely well suited to a single AI approach that can serve many different kinds of customers at once. It’s also the core structural blocker in this business. You can generate all the inbound volume you want, but none of it matters if your underwriters can only process two and a half loans a day. Solve that constraint and you’ve solved the thing actually limiting growth for a lot of lenders, not just made an existing process marginally faster.
The honest state of frontier models today is that they’re genuinely excellent at parsing massive amounts of text, synthesizing it, and summarizing it. What they’re not naturally good at is doing that with high precision when ambiguity shows up, which is exactly the problem underwriting hands you constantly. The way we’ve approached solving that is by grounding the model in historical data rather than asking it to reason from guidelines alone. And there’s a real ceiling here worth understanding. These models can technically attempt almost any task you throw at them, but the real constraint is how much relevant context you feed them at once. Give a model too much at one time and you actually increase the risk of hallucination, because it’s holding more than it can meaningfully reconcile. I like to compare it to sitting through a single focused hour of class versus an unbroken eight hour session and then being quizzed on material from the first hour. What you absorbed early on doesn’t hold up as well by hour eight. Managing that context window deliberately, giving a model exactly what it needs for a specific task rather than everything you have, matters more right now than raw compute power. The chips themselves are already quite capable. Companies pushing AI onto edge devices and autonomous systems without a cloud connection are proof of that. The harder constraint is how intelligently you scope what a model actually sees.
That’s the philosophy behind what we mean at Balerion by intelligence before underwriting. Underwriting is arguably the single most expensive part of loan manufacturing, and underwriters are constantly burdened with problems that should never have reached their desk in the first place, a missing document, an income calculation the loan officer got wrong, a borrower who simply isn’t qualified yet. Our focus is catching that upstream, before a file ever lands with an underwriter, so that when it does arrive, it’s genuinely clean and doesn’t require constant back and forth between the borrower, the loan officer, the processor, and the underwriter. We surface that with a simple red, amber, green status. Red means something needs to be resolved directly with the borrower. Green means the AI is confident enough in its assessment that it may not require additional human review. Amber is the category I think matters most, because it’s the model honestly saying it isn’t sure, and building that threshold for genuine uncertainty into the system is far more valuable than a model that’s confidently wrong. Knowing when to hand something to a human is as important as knowing when not to.
When I talk to the people actually buying this technology, usually a COO or chief lending officer, the conversation almost always comes back to cost, but not in the way people assume. It’s rarely just the headline cost per loan. It’s the accumulated cost of maintaining twenty or thirty different vendor relationships, and the very real opportunity cost of being locked into a two or three year contract with one of them while better technology emerges elsewhere. What we’ve tried to build is closer to an orchestrator, something that can plug into the loan lifecycle starting at underwriting and expand in either direction from there, rather than one more point solution that adds to the pile. The number one question we hear, phrased a dozen different ways, is essentially how do I take advantage of AI without dramatically increasing my cost per loan in the short term while I figure out the long-term payoff.
That cost sensitivity is exactly why the way we partner with lenders looks different from a typical vendor relationship. We don’t hand over an API and a UI and hope operating costs improve on their own. We spend real time, usually a week or two at the start of any partnership, actually inside a lender’s LOS observing their process and understanding their specific bottlenecks before we tailor anything. There’s usually an eighty percent overlap in what any two lenders need from a system like ours, but that remaining twenty percent of customization is where a real partnership either earns its keep or falls apart. We back that up by running trial testing against hundreds of actual loan files, shadow underwriting our output against what real underwriters concluded on the same files, so lenders aren’t taking our accuracy claims on faith.
None of that works without pairing the right kind of technical talent with real domain experience, and that combination is deliberate. On one side we need people who genuinely understand reinforcement learning, post-training, and how to build and scale agents, often coming from equally regulated, document-heavy industries like construction and insurance where a lot of the same complexity exists. On the other side, mortgage is its own tightly closed world. You’re either genuinely inside it or you’re clearly outside it, and understanding why this domain is so complex, why a borrower with seasonal or non-traditional income doesn’t fit neatly into a standard box, requires people who’ve actually lived through that complexity, not just studied it from the outside. Marrying those two backgrounds is what lets us build something that’s technically sound and genuinely usable inside a real lending operation, rather than a technically impressive product that quietly ignores how underwriting actually gets done.
Looking ahead, one of the things I’m most excited about is publishing what I believe will be the mortgage industry’s first real lending evaluation, several hundred fully labeled loan files benchmarked against how different frontier and open source models actually perform on them. I think that kind of transparency is overdue in an industry where everyone claims high accuracy but few are willing to show their work against a shared standard. If this space is going to earn the trust it’s asking lenders to extend, publishing honest, comparable results is a good place to start.