
When researchers asked three leading general-purpose AI models to flag bank-statement deposits that “could be of foreign origin,” the models tagged transactions tied to English-sounding names as suspect just 13.3% of the time. When the depositor’s name wasn’t English, that rate jumped to 77%. It’s one of several findings from a new Columbia University study that put AI mortgage assistants through a battery of real underwriting questions — and found the top models got the full, correct answer wrong roughly one time in four.
The research, “MortarBench: Evaluating Mortgage Loan Origination Agents,” was published on the preprint server arXiv by a team led by Matthew Toles, a Columbia University doctoral student, and Zhou Yu, a Columbia associate professor, working with mortgage technology company Tidalwave. It is the primary study behind a report published this week by Realtor.com. On the benchmark’s strictest measure — whether a model’s complete answer matched the known correct one — Gemini 3.1 Pro scored 77.1%, GPT-5.5 scored 76.8%, and Claude Sonnet 4.6 scored 51.4%, according to the study and Realtor.com’s reporting.
The timing matters. More than 80% of mortgage lenders were evaluating AI tools as of June, according to a survey by The Mortgage Collaborative, and 17% had already deployed AI in live production workflows. That adoption curve is a direct response to cost pressure: originating a retail mortgage cost lenders about $11,800 per loan in the second quarter of 2025, per Freddie Mac data cited in the study, while basic digital underwriting tools have saved roughly $1,700 per loan and cut production time by about five days. RealtyWire has previously covered how AI is reshaping real estate search and how brokerages are adopting AI operating platforms more broadly across the industry.
Where the models failed
To build MortarBench, researchers pulled real questions submitted to a mortgage assistant and narrowed them to the tasks loan officers handle most often during origination: whether payroll deposits match the employer listed on an application, which deposits are large enough to require scrutiny, and whether an account is jointly held with someone who isn’t applying for the loan.
The models’ biggest weakness was picking specific transactions out of long bank statements — what Zhou Yu described as finding “the needle in the haystack.” But the errors ran in both directions: when researchers manually reviewed Gemini’s incorrect answers on transaction-list questions, misclassification was the most common problem. In one case, the model counted a personal loan as a buy-now-pay-later transaction. Other errors included assuming all wire transfers were international, treating deposits from co-borrowers as automatically documented, and classifying a one-time housing payment as recurring.
Toles cautioned that the scores reflect “naive use of foundational models, largely equivalent to taking the application package, pasting it into ChatGPT, and asking it a bunch of questions about it,” and that major lenders typically build their own proprietary systems, which perform better. Tidalwave’s own mortgage-trained model, SOLO, scored 95% on yes-or-no compliance questions in a separate, earlier benchmark, according to the study’s authors. In the new work, the researchers also introduced a proposed fix, a confidence-calibration framework called CRIT, that lifted accuracy to 80.5% in testing.
What it means
The findings arrive as regulators sharpen their focus on AI oversight. Fannie Mae and Freddie Mac formalized new AI governance requirements for their seller-servicers in 2026, putting the burden on lenders to manage risk from AI systems — including tools built by outside vendors — and to disclose what AI they use and what safeguards are in place when asked.
That regulatory backdrop is separate from the study’s own findings, which are limited to a specific benchmark of general-purpose models answering origination questions without specialized training; MortarBench doesn’t test every product on the market, and the researchers themselves note that purpose-built, proprietary tools can score meaningfully higher. Still, the bias finding around non-English names raises a fair-lending concern that goes beyond simple accuracy, since flagging deposits based on how a name sounds rather than documented risk factors is not a permissible underwriting practice.
Toles argued the stakes extend past any one lender’s technology choice. “We don’t have visibility into what these models are doing, what companies are doing with them, and what the outcomes are,” he said, adding that if many mortgage companies converge on the same underlying models, systemic errors could compound across the industry. Diane Yu, Tidalwave’s co-founder and CEO, said borrowers should directly ask lenders whether their financial information is being passed to an outside large language model and whether that AI has undergone independent evaluation. “You should be very careful,” she said.
What to watch
Because MortarBench is open source, lenders, regulators and vendors can now run the same test against their own tools rather than each defining accuracy on their own terms — a shift that echoes broader benchmarking efforts already underway in mortgage technology. Watch for individual lenders and AI vendors to publish their own MortarBench scores, for Fannie Mae and Freddie Mac to detail how their new AI governance rules will be enforced, and for follow-up research testing whether purpose-built mortgage models close the bias gap as reliably as they close the accuracy gap.



