Hudson Labs CEO Kris Bennatti: Financial AI Still Gets the Numbers Wrong

In the latest episode of Invest with AI, Brett and Khe sat down with Kris Bennatti, CEO and co-founder of Hudson Labs, who spends a great deal of time stress-testing these models against real workflows. Brett's question: how precise can these tools really be when the models underneath are non-deterministic by nature?

A recent test of top models revealed that asking for a single metric will return the wrong number about 30% of the time when pulled more than 4-8 reporting periods in the past. This wasn't just true for edge cases, as models were failing the routine work of pulling key figures across a few years of filings. However, Hudson Labs just launched what Kris says no one else in financial AI has promised: a no-hallucination guarantee. Her read is that the fact that this guarantee is so rare tells you how much inaccuracy risk is sitting in the average financial AI tool.

The takeaway from the conversation is not that you should abandon using AI for finance. Kris's argument is that institutional-grade accuracy lives in the back end: the data, the retrieval, and the checks rather than the frontier models. The analyst who adapts their process with this in mind is the analyst who will get the most out of financial AI tools.

We get into:

  • Why frontier models still return the wrong number about 30% of the time past a few reporting periods, and why precision, not reasoning, is the problem improving the slowest

  • The no-hallucination guarantee Hudson just launched, restatement-adjusted across a five-year period, and why Kris says no one else in financial AI will make it

  • Where hallucination actually comes from: the model that does not know what it does not know once you push past its comfortable context limit

  • Why NotebookLM feels more accurate than a generalist model, which comes down to the search and not a bigger context window, and what a vector database does that just using Claude does not

  • Sentiment screening you cannot run anywhere else, like asking for the most stressed-out CEOs, and the metadata that makes it possible

  • Inside the forensic risk score: turning qualitative fraud signals, a CFO who quit, related-party contracts, off-balance-sheet debt, into a single number, with machine learning doing the weighting

  • The track record, where a score of 70 or higher means roughly a one-in-three chance of SEC enforcement

  • How MCP went from introducing hallucinations into results to something Kris now recommends, and why the guarantee does not apply through it

  • From prompts to skills: why prompting matters less every month, and how typing SSS gets you same-store sales

  • Integrate or die, Kris's 2026 thesis that the future analyst is coding, and why she would never buy a product without an API

  • The cost inversion nobody is pricing in: a dataset that runs about $100 on Hudson's architecture can cost $13,000 straight off an Opus API, and why today's AI is quietly inefficient

  • Why Kris thinks Anthropic will not be king forever, the open-source moment, and the banks building their own model routers

  • The soft guidance frontier models still miss because they key on verb tense, like when capex is now expected to be

If you work in equity research, at a hedge fund, or anywhere on the buy-side and you are trying to figure out which AI tools you can actually rely on, start here.

Chapters (Timestamps)

Timestamps:

[00:00] Intro
[00:40] Meet Kris Bennatti, CEO of Hudson Labs
[01:22] From Pre-LLM Filings to Founding Hudson Labs
[04:16] Why Precision Is Finance AI's Hardest Problem
[05:35] The Test: Wrong Numbers 30% of the Time
[06:29] Launching the No-Hallucination Guarantee
[10:20] Why NotebookLM Feels Better (It's the Search)
[12:01] Vector Databases vs. Just Using Claude
[13:36] Screening for the "Most Stressed-Out CEOs"
[14:16] Have MCP and Connectors Fixed Accuracy?
[19:50] From Prompts to Skills
[20:50] MCP vs. CLI
[23:04] Integrate or Die: The 2026 Business Model
[26:52] Where Generalist Models Catch Up (and Where They Won't)
[33:36] Inside the Forensic Risk Score
[38:47] The Track Record: Score 70+, a 1-in-3 Chance of SEC Action
[44:17] The Cost Problem: $100 for Hudson, $13K on Opus
[46:26] Fable, Co-work, and the New Token Economics
[51:20] The Guidance LLMs Still Miss (It's the Verb Tense)
[54:05] Where to Find Kris & Hudson Labs


Want to actually build these workflows yourself?

The AI Accelerator is Fundamental Edge's 6-month cohort for investors who want repeatable AI workflows. Learn More below:

https://www.fundamentedge.com/ai-accelerator

Follow Invest with AI on:

Spotify: https://open.spotify.com/show/033xcEEovVViS7hIYwNuGZ

Apple Podcasts: https://podcasts.apple.com/us/podcast/invest-with-ai/id1896918892

Next
Next

Fable Is Here, But Is It Actually Better? | Invest with AI Vibe Check