Invest with AI
Brett Caughran and Khe Hy lead a deep dive into AI for investing, joined by guests at the cutting edge of the field. Our goal is to be your Sherpa through a rapidly changing landscape by distilling what's working, what isn't working, and how you can leverage AI in your own process. Follow along as we tackle AI's biggest challenges and opportunities, one episode at a time.
Invest with AI
Hudson Labs CEO Kris Bennatti: Financial AI Still Gets the Numbers Wrong
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
Kris Bennatti and Hudson Labs are now giving a no-hallucination guarantee, something she says no other financial AI offers. The reason it's rare is the uncomfortable part: her team found state-of-the-art models return the wrong financial number about 30% of the time once you go past a few reporting periods.
Kris, Brett and Khe get into why precision is still a difficult problem, how a forensic risk score turns fraud signals into a hard number, and why she thinks 2026 is "integrate or die."
"I believe we're the only people on the market that are willing to do a no-hallucination guarantee, which does give you a bit of a sense of how much risk there is in the average financial AI software." — Kris Bennatti, CEO of Hudson Labs
-----------------------------------------------
Timestamps:
[00:00] Intro
[00:40] Meet Kris Bennatti, CEO of Hudson Labs
[01:22] From Pre-LLM Filings to Founding Hudson Labs
[04:16] Why Precision Is Finance AI's Hardest Problem
[05:35] The Test: Wrong Numbers 30% of the Time
[06:29] Launching the No-Hallucination Guarantee
[10:20] Why NotebookLM Feels Better (It's the Search)
[12:01] Vector Databases vs. Just Using Claude
[13:36] Screening for the "Most Stressed-Out CEOs"
[14:16] Have MCP and Connectors Fixed Accuracy?
[19:50] From Prompts to Skills
[20:50] MCP vs. CLI
[23:04] Integrate or Die: The 2026 Business Model
[26:52] Where Generalist Models Catch Up (and Where They Won't)
[33:36] Inside the Forensic Risk Score
[38:47] The Track Record: Score 70+, a 1-in-3 Chance of SEC Action
[44:17] The Cost Problem: $100 for Hudson, $13K on Opus
[46:26] Fable, Co-work, and the New Token Economics
[51:20] The Guidance LLMs Still Miss (It's the Verb Tense)
[54:05] Where to Find Kris & Hudson Labs
-----------------------------------------------
Want to actually build these workflows yourself?
The AI Accelerator is Fundamental Edge's 6-month cohort for investors who want repeatable AI workflows. Learn More below:
https://www.fundamentedge.com/ai-accelerator
Watch the full podcast series on our site: https://www.fundamentedge.com/invest-with-ai
Follow Invest with AI on:
Spotify: https://open.spotify.com/show/033xcEEovVViS7hIYwNuGZ
Apple Podcasts: https://podcasts.apple.com/us/podcast/invest-with-ai/id1896918892
For the foreseeable future, if you want to make sure that your numbers are definitely correct, you do need some pre-processing and some built-in checks.
SPEAKER_02How have you thought about that transition of the instructions from the investor to the portal?
SPEAKER_00You know, prompting is becoming a lot less important. And I believe we're the only people on the market that are willing to do a no-hallucination guarantee, which does give you a bit of a sense of how much risk there is in the average five-minute toll AI software.
SPEAKER_02All right. Chris was the co-instructor for the first iteration of the Fundamental Edge AI Accelerator. Has been, you know, someone who's really taught me a lot about deploying AI and investment research and is a founder and CEO of Hudson Labs. So thanks so much for being here, Chris. And uh maybe start just telling us a little bit about your background, how you got into AI and finance.
SPEAKER_00Awesome. Thanks so much for having me. As always, a pleasure. Uh yeah, I'm I'm Chris Bonotti. I've been around for a minute now. Uh I got my start way back when as a data scientist at a corporate governance advisory firm, uh, where I focused on sort of pre-LLM securities filing processing and was building my career as AI was making really big moves and LLM specifically in the research community. And this would be would have been back in say 2017, 2018, we started seeing auto-suggest on our phones. Uh, the research community was getting super excited about LLMs, and yet at my workplace, I was being hailed as you know, like the ultimate innovator running uh essentially all word-based counts on Canadian corporate disclosure. Um and saw that and uh, you know, sort of recognized the fact that oftentimes, at least historically in finance, uh, we're in corporate governance and in the corporate fields, we're many years behind consumer technology. Uh, and we were not adopting LLMs nearly as quickly as you know, iPhone, Gmail, all of these cool applications, quit my job, co-founded Hudson Labs with uh the LLM guy here in Canada who is SuHas Pi. If you haven't purchased his book, Designing LLM Applications, I highly recommend doing it. It's published with O'Reilly, it's now been translated into over 10 languages, including Polish. Uh, and uh SuHus does a lot of really cool uh training and onboarding around AI, agentic uh modeling right now. So do give him a follow and uh attend some of his seminars. Uh yeah, and then we we uh launched what was then bedrock AI, now Hudson Labs, and we focus on AI for institutional finance. Our big contributions to AI research have to do with materiality ranking uh and precision. So being able to uh maintain accuracy over very large contexts, uh, be able to pull numbers, uh, pull guidance, uh, pull sentiment in a way that's consistent, etc. We also uh started our business with the forensic risk score. Uh and uh for for many years now have been popular among the short community, among plaintiff class action lawyers and DNO insurers as well as buy-side firms.
SPEAKER_02Great, great. Maybe um I've always found um I always get like super excited on things, and then I call you and you're like, well, let's talk about the precision of uh of these elements, and precision is so critical for finance use cases. You know, you can sort of in so many AI things create these impressive looking demos, but if the numbers aren't at a level of precision, they're just not institutionally useful. How do you think about just the broad challenge of precision with large language models, which are non-deterministic by nature? And and what are a few of the key vectors that you've thought about for years about making these tools more precise?
SPEAKER_00Yeah, it's a really tricky problem and it's been a persistent problem. It's definitely getting better uh from a from a generalist perspective, but for the foreseeable future, if you want to make sure that your numbers are definitely correct, you do need some pre-processing and some built-in checks. It's pretty much impossible to just take a generalist model, apply it to finance, and be like, okay, great, these this is all going to be right. Um, for reference, we did a test on all of the state-of-the-art models a few months ago now. Uh so it's been before before 5.5, whatever the the Claude and uh Open AI models were a few months ago, and whenever you asked for a metric for more than four to eight periods, you'd get an incorrect number 30% of the time. Uh so when you think about hallucination right now, there's mostly a problem or the biggest problem when you start getting beyond comfortable context limits. Uh, and that's one of the big problems we focus on. Uh, it will continue to get better, but as you've seen so far, it's getting better much, much, much more slowly than things we've seen, like reasoning, uh, mathematical elements, all of those things. We did, as of yesterday, formally launch our no hallucination guarantee. Uh, so we're pretty feeling incredibly confident about our ability to not pull incorrect numbers, uh, restatement adjusted, error-free numbers across a five-year period. Uh, that's all done using AI architecture rather than just thinking about the last generative step. Um, and you know, it's it's taken us a while to um be able to do that five-year period, restatement adjusted perfectly, uh, but we're we're feeling really excited about that. And I believe we're the only people on the market that that are willing to do a no hallucination guarantee, uh, which does give you a bit of a sense of of how much risk there is in the average financial AI software.
SPEAKER_02What did what does that entail to gain the conviction to offer a guarantee like that? Like you said, pre pre-processing and built-in checks. Like, what is that? How hard is that problem? What are the big work streams to sort of take the general models and convert them to that hallucination-free layer?
SPEAKER_00Yeah, so uh if you want to think about hallucination, you I think it's chapter 11 of Suhaus's textbook. He talks about the the place where you end up with hallucination is when the model doesn't know the answer and it doesn't know it doesn't know the answer. That is the highest risk. Time when you're not feeding the model the information it needs and it doesn't know that. So, as you can imagine, when you start moving beyond these comfortable context limits, that's when you start seeing hallucination in a generalist set. How do we overcome that? We talked about pre-processing. So once you get towards more than say four, eight quarters, you need to start using search and retrieval to get the right information. And one of the reasons so many tools, I'm not going to name names, but have issues in that long time horizon, is because search and retrieval is really hard. If you're just looking for revenue and you're searching for it again across 20 press releases, you're going to be pulling in a lot of noise. And if you think about how complicated given getting a revenue right, depending on what you're asking for, can be, you know, maybe you're looking for a sub-segment nine-month. And instead you're going to be getting numbers in your retrieved element that are consolidated and three-month. So those sort of long horizon complex queries, in order to fix that problem, you do need some pre-processing where your embeddings need to have information about what is this, what time period is it related to. And that allows the model to be able to select the information it needs and actually use it. And when we talk about guardrails or checks, it's about making sure you have specific pathways that exist so that the model knows that you're asking for a KPI poll and it should be able to expect this information. And if it's not getting it, it just gives you an error.
SPEAKER_01Can I ask a uh uh lazy uh layperson question? How um if you think of the I definitely appreciate the context window challenge? How does this come like how does something like Notebook LM solve this where presumably it's it as a user that doesn't understand the technical in and outs of it, it feels like the context window is much larger, just like I can give it 500 files. It seems to be more accurate in the citation uh nature of it. It's not as intelligent, it's like pure retrieval. Um, but can you how does that compare like like what's happening under the hood in a notebook LM versus a generalist model versus a Hudson's uh labs um approach?
SPEAKER_00So one of the reasons that Notebook LM feels better is because it has a more intelligent search.
unknownYeah.
SPEAKER_00So it's gonna be, it's not finance specific, obviously. It's not numeric, it's not built for pulling KPIs. You are going to still, even with Notebook LLM, you are going to notebook LM. Uh, you are gonna run into that um sort of a quarter plus failure mode. I don't know if you've tried it. Uh so for the same reasons around search, but it is using a more advanced search mechanism. So it's not about it's not, it doesn't have a better context window, it just has better search.
SPEAKER_01Got it.
SPEAKER_00Yeah.
SPEAKER_01And so Hudson being finance experts, you can basically customize that search process way better than any than than a generalist approach.
SPEAKER_00So if you want to think about Hudson Labs versus NoBook LM versus just having a claude instance, what's happening? The back end is a big part of the difference. So whenever we take in an SEC filing, a transcript, etc., it's we use vector databases. So we're converting a sentence, a word, a paragraph to an embedding. Are you uh familiar with yeah, or a vector, yeah. So there's lots of different ways that you can do that, and you should be doing that. But if you're using Claude, your data isn't in a vector database, which is one of the way reasons that it can be so expensive. Uh and if you're using Notebook LM, it's uh it's essentially search rather than a vector database. Got it. But if you have a vector database or you have a more sophisticated backend, you can do a lot of really cool things in pre-processing that facilitate both accuracy, relevance, and completeness on the front end. Got it. So in addition to to be able to, if you have these things as embeddings, you can give metadata to the model in addition to containing better information about what that sentence is. And you can do cool things with your embeddings. For instance, the way that we store embeddings, we have meta declarations not just about is this uh a KPI or not, is it forward-looking historical, but also information about sentiment. Which is why I believe we're the only platform where you can do sentiment-based screening. So ask like who are the most stressed-out CEOs, those types of queries. Because if you're using that like LM and it's just using search, how does it search for stress? There's no keyword that you can use there. Uh, but if you have embeddings that have information about sentiment, you can do really cool searches.
SPEAKER_02This was a this was a huge problem for the institutional finance use cases of chatbots in default 25. And in all of my experimentation going into Chat GPT or Claude, almost anything quantitative was not helpful due to this sort of accuracy, like 60, 70% benchmark. Part of the way the ecosystem is aimed to solve that is through MCP and connectors into an agentic workspace. And so now I can connect into rather than web searching for a number, I can connect into the MCP of a DeLupa or a FACSET, et cetera, through my agencorkspace. Is that how do you think about sort of what you're doing and sort of the you you've aimed to solve this problem in the way you just explained? How would you evaluate the way that the industry has tried to solve this problem with that MCP and connector uh movement into agents?
SPEAKER_00Yeah, so I know the last time we talked about this, I had complained about the Claude MCP introducing hallucinations into Hudson Labs results, which had been very frustrating. Uh and I would say I've been very, very impressed with how MCPs have evolved since. They're a lot less brittle than they used to be. Uh, I you know, I think I think it's a good approach. I think um having connectors that deal with the vectorization, the pre-processing, uh, doing all of this back-end work for you and pulling that into your own universe is a good idea. Uh there's definitely some trade-offs if you're using MCP, if you're using Hudson Labs via MCP, it's gonna be a lot less efficient. There's sort of a one of the ways that Claude has become less brittle is by using more calls. Uh so it so it is a much less efficient way of using Hudson Labs, but it is much better than I don't know the last time we talked about this, Brett, but um but it it's it's a lot better than it was before. I mean, there's still some situations where it becomes harder to control. Uh, we we don't offer uh no hallucination guarantee if you are using Hudson Labs through MCP because it's it's non-deterministic, right? It's it's using another model on top of ours, and sometimes it can introduce things that we didn't expect or you didn't expect, but at this point, we feel that the benefits outweigh the risks, which wasn't true a year ago.
SPEAKER_02Interesting, interesting. Yeah. And that's principally like the the evolution has principally been been what? Just the the models getting better, the engineering around it. Like how has that brittleness problem been uh been uh mitigated?
SPEAKER_00Models getting better, more guidance around how to set up an MCP. I would say a lot of issues right now with usage of say any connector, they're not they're mostly solvable through better API API documentation, better MCP documentation, facilitating more endpoints. Uh so it's it's because there's a little bit more more documentation and more flexibility, it becomes easier to guide these MPC, MCP servers to make better use of your own own data.
SPEAKER_02And I hear, and I don't fully understand this, I hear some people say, like, oh, XYZ vendor has a their API is great, or their MCP is great, their MCP is not so good. Like, what is the skill in in sort of like for a neophyte, like what's the process of taking your data into an MCP and where is the skill in that? Like, how how do some vendors sort of do that well? How do some vendors fail in that in your in your assessment?
SPEAKER_00Like many things, I think it's a function of effort and attention to detail because it does take effort to break down your product and make it identically usable. You're translating essentially what is was originally a user interface into something that makes the most sense to an LLM, and what makes sense to an LLM doesn't always make sense to a human being. I think you know, because we have Sue Hoss, we're particularly good at making our our tools LLM readable. That is another thing we talked a little bit about, you know, what's different between Hudson Labs, what's going on on the back end, and one of the things that we do do differently is when you ask a question as a human being, we have a translation step and make your prompt more LLM friendly, particularly when we're talking about market search, where those search and retrieval steps and how you you ask that question really, really matters. And that applies to to MCP as well.
SPEAKER_02Yeah. How have you thought about you know, how have you thought about the evolution of prompts to to skills? I find myself really not prompting much anymore. I have a few sort of conversion prompts, but I find my day-to-day investment process has really centered around skills into an agentic workspace. How have you thought about that transit that transition of the instructions from the investor to the to the portal?
SPEAKER_00Yeah, the goal as a provider is to provide the thing that the customer wants with as little instruction as possible. So we think about that a lot. You know, if you just type in SSS, we will guess that you want same-store sales and bring that back to you. And that's not something that we were able to do or doing two years ago. So you're right, you know, prompting is becoming a lot less important.
SPEAKER_01Can I can I ask about on the MCP uh kind of uh creation and evolution? I'm gonna put my like power user open claw hat on, so it's not maybe directly relevant to financial services LLM, but there's a big push even in that community from MCPs to commit to CLIs. So, like even more like, and again, it's gonna be beyond my my technical understanding, but CLIs being more efficient and even more friendly for um agents to communicate with. I guess my question for you is like, is that is that true? Is that um how do you see that playing out? As an example, like in my own workspace, I've been using the Google CLI instead of the Google Connector. It's like so much more efficient and so much faster. I've been using a Notion CLI instead of the MCP server. Um, and the results are much better. Um, and so I'm wondering uh if that's something that you are all thinking about, or how you kind of think of that evolution or the relationship between. MCP and CLI.
SPEAKER_00It's not something that I spent a lot of time thinking about. I'm sure our technical team has. And the one thing I would also say is you'd be surprised by you know, we're all on Twitter a lot. We see sort of wow the coolest things that everyone's doing with AI. A lot of our customers don't even know what MCP is. So it's gonna be probably a minute before we prioritize doing a CLI given given it it's tough. I think it's there's a there's a huge range of where consumers are are at right now, and you've got to cater to to all of them, but we're we're okay with where we're at right now. Yeah.
SPEAKER_02How how have you thought about the in sort of business model perspective or about the around the agent path and sort of my experience having demoed many dozen of the the sort of chat bot finance chatbots last year? I sort of ultimately found like three or four that were really good that I liked, and Hudson Labs was one of those. And that was the good news. The bad news is it sort of remained a cumbersome exercise to log into Portrait and Hudson Labs and Alpha Sense and sort of trying to remember where to go to ask what question wasn't a great user experience contrasted to now an agentic workspace with my skills embedded. I don't have to sort of copy and paste prompts and my connectors. It's a much more delightful, seamless experience. Um and so how but how do you think about sort of shifting what what have you seen that as well amongst your user base? And how do you think about sort of shifting the Hudson Labs business model to align with that agent path?
SPEAKER_00Yeah, my perspective is in 2026, integrate or die. You have to be the future analyst, is coding, they're building their own tools. I'm relatively technical. I would never buy a product that doesn't have an API. That's going to become everyone going forward. So you have to offer integrations, you have to offer a product that works for people who are building their own dashboards, building their own skills. And we're doing that. We're improving all of that functionality and making sure that we have connectors that work for a variety of use cases. Um the other way we're thinking about it is we have this bifurcation in the market right now where there's a lot of users like you, Brett, who are really excited about their skills, being able to pull things from other all these different places. They're using Perplex V computer, doing all of the cool stuff. They love the MCP experience, and we have these big enterprise contracts with people who can only use Copilot. Uh, they've never tried any of this before, and they're looking to us to become the the place where where they collect information and bring bring information into us and and are thinking about us as that place where you go to collect everything and almost be the server itself. So we're prioritizing integrating with other people right now, um, but also thinking about how do we make sure uh that our users who don't have access to these more powerful tools can take advantage of Hudson Labs in a more extensible way so that we also have connectors coming in, not just going out.
SPEAKER_02Gotcha. So almost sort of a bifurcation where in some like I think everyone's searching for a single pane of glass right now. So in some instances, you would be that single pane of glass and you would have sort of ingressed MCP connectors in that. Is that is that is that right? And then you would also sort of have egress MCP into another single pane of glass.
SPEAKER_00Yeah, and not just MCP, API, everything, all of the integrations, yeah.
SPEAKER_02How how do you sort of think about the the durable advantage of of Hudson labs in that sort of you know, single pane of glass elsewhere? Like what are the key, we could talk about forensic risk score, obviously, or sort of a differentiated uh data. What what do you sort of like? I'm trying to sort of figure out what the institutional grade stack of connectors looks like, how many that is, uh, where do I get my alt data transcripts, et cetera? Like, where do you think, what do you see as the TAM or the market share of Hudson Labs in that institutional grade fundamental agent stack?
SPEAKER_00We're gaining popularities in model building and hope to corner that market with the no hallucination guarantee being able to pull all of the information you actually need. Obviously, there's lots of tools like Dalupa that do more out-of-the-box modeling for people who are doing it themselves. Um, but generally where we think about our moat versus others, so we've got the no hallucination elements, which is a precision element, which is something institutional users really care about, whether that's single company or multi-company. Um, but we do expect eventually generalist models in the next five years to catch up from a hallucination perspective. And we don't think that you know, 10 years from now we can still be competing on a precision element, particularly from a single company perspective. Uh, so really where we provide a lot of value is in situations where you need to be doing really good search and retrieval. Right now, that does apply to single company workflows because you can't get good results over you know a longer time period and pull those numbers accurately if you want them restatement adjusted, if you don't want them to be full of junk. Uh, but where we offer a ton of value now and in the future is being able to retrieve information across 10,000, 20,000 companies. Uh, I think we have 52 million data points in our our vector vector database at this point. And you can do things with Hudson Labs that you literally cannot do with any other platform, like pull out uh cool sentiment-based elements, be able to find the things you're looking for quickly using materiality ranking, using our uh our declarations around metadata, etc. So uh to some extent, mostly right now, big firms don't have sophisticated search and retrieval, they're using the out-of-the-box brute force methodology, and you mentioned the fact that they don't care about spend, and it's okay that these big firms are spending like $30,000 uh every couple of weeks trying to create information without a backend. But the problem is that it's not just about cost when we're thinking about search and retrieval and using a brute force mechanism across a large data set. It's also about accuracy. Um, you know, if you try to use a brute force methodology to, you know, I'm an accountant, look up accounting policy changes using Claude right now, you end up missing so much information because it's hard to find. Uh, so there's this big demand for how do we do not just single company deep dives, but how do we do AI-based screening, data set creation, high precision data set creation, uh, go the extra step uh in these multi-company workflows.
SPEAKER_01Is it fair to say that the brute, so the brute force, because presumably like with every new model release, the brute force powers gets get better. Is it uh if I'm hearing you correctly, brute force can give you like breath, like it might improve your breath capacity, like your breath capabilities, but you still have to manage the accuracy, like brute force is not gonna be as helpful to maintain that accuracy component.
SPEAKER_00Brute force is so what's a topic that you want to learn about right now? Okay.
SPEAKER_01What's a topic? Um, let's see. I'm trying to learn about uh data lakes and data warehouses, and like the distillation of information across different structures.
SPEAKER_00There's a lot of information out there on the web about data lakes. There's a lot of information in SEC filings. Let's say data lakes, there's 15 articles about not data lakes, but other types of databases that are competitors and don't have the word data lake in them. If you want to use the brute force approach right now, you're only going to be able to effectively search for things that are easily findable. Right?
unknownYeah.
SPEAKER_00So if you're thinking about getting a complete lay of the land using AI search, you need a way to discover information that is related but not easily searchable. And that's true for a lot of AI style search. One reason I often point to sentiment is because it becomes really clear in that that example. If you go, okay, find frustrated managers or find managers that are show evidence of deflection. That's not something that you can type in and search for. And if you think of the AI model, how is it going to find it if you can't search for it either? And you need some way of making that information findable so that when you pull that into the context window, there's the model has something to work with.
SPEAKER_02Yeah. Talk to us a little bit about um you know you're an accountant, the forensic accounting score, which is um part of the one of the delightful capabilities of AI is to take uh a previously highly cumbersome process like doing a deep forensic accounting analysis and sort of distill that into a score which has the quantitative elements, but also the unstructured data elements. Um this was sort of something that caught my attention very early on, and I've been a big fan of the the Hudson risk score approach. Can you sort of just crack that open for us and and walk through uh what you've been doing there?
SPEAKER_00Yeah. So just to take a quick step back, people have been trying to predict securities fraud since the beginning of time. And historically, uh, when we used to try to do fraud prediction, we would use generally things like day sales outstanding, how quickly is revenue growing, accruals, these financial metrics. Have you ever tried doing that, Brett?
SPEAKER_02Oh yeah.
SPEAKER_00Yeah, yeah. And the problem is that it does find fraudulent companies, but it also finds a ton of just normal companies. It's a very noisy way of predicting fraud. It's a the reason why when you read, you know, a Hindenburg report or when you're pitching a short to your boss, you're not going, these sales outstanding, increase by X percentage. You're saying, hey, look, the CFO quit. They're selling the product to their mom. There's millions of dollars of off-balance sheet debt. The the market thinks is is really excited about this contract, but the contract's actually with a related party, and the market doesn't know that. So those are the things that that create a good short seat thesis. So uh we take that information, integrity of the management team, turnover at the top, you know, the number of times they're changing their segment disclosure, um, aggressive accounting policies, off balance sheet risk, related party transactions, uh dependence, uh governance. Fun fact, uh dual class governance is highly predictive of both growth and fraud. Uh, all of these elements and using LLMs, we're able to take all of these previously very qualitative elements of risk that you know maybe you'd be able to look at and go, hey, that's sketchy, but we can convert that into a mathematical representation of risk. How important is this weird family related party relationship? And then use those quantifications to predict fraud. So uh our model looks at all of these different types of qualitative risks and gives says given that they restated once in the last three years, and the CFO uh quit and they changed their useful life estimate. What's the likelihood they'll be subject to an SEC enforcement action or uh litigation related to securities fraud?
SPEAKER_02How have you um and part of part of the reason I love that is sort of like heretofore like tracking segment you know disclosure changes over time at scale, there was just like no way to there was no way to do that, like all these little signals into a into a collective mosaic? Like, yes, you get this sort of footprint of like this feels like you know there's some obfuscation here, right? And there was no way to scale that scale that before before. Um how do you think about like the actual like algorithm of like putting that into a score? Like how do you how how have you thought about weighting those various components?
SPEAKER_00We let machine learning do the weighting. So we have a forensic model that classifies or represents the underlying risk. And then we take those and create variables, quantitative variables, and then we run a standard ML process where we have a big database of companies that were demonstrably fraudulent, either because they had an SEC investigation or a settled class action lawsuit related to fraud, and we let the machine learning model figure out the weights, which are often not the weights that uh a human short seller would apply, which is is somewhat interesting.
SPEAKER_02Interesting. How have you tracked how how long has that model been live for? And how do you track like efficacy? Is there any way to sort of distill that down into a batting average or hit rate or alpha, et cetera?
SPEAKER_00Yeah, so it's been a while since we've done a backtest versus uh share price. Um I believe there's a about like a 14-point differential uh in performance, so uh underperforming by about 15% for high-risk companies back back when we did it. But uh we've been doing this since 2019, uh, very popular among actuarial teams that do DNO insurers. So we uh do a lot of tracking based on uh SEA's securities class actions, which is the most important fraud indicator for them. Uh so high-risk companies are three times more likely to uh be subject to an SEA of any type of any kind. And every company that has a score of 70 or higher has about a one in three chance of SEC enforcement. So it's it's very high precision. Uh yeah. And of course, the the other two out of three also likely frauds, just not subject to to SEC enforcement.
SPEAKER_02Yeah, yeah. Yeah. The conversation we had one time is is um just how sort of natively good good LLM to Bennett identifying risk. And I think the sort of conclusion was there's a lot of evidence in the training corpus to sort of pattern recognize, like, am I am I remembering that conversation correctly? Like why why why have LLMs been sort of why why has this been sort of a strong native skill of of AI, the sort of risk dimension dimension?
SPEAKER_00We've had mixed results actually with out-of-the-box LLMs. Okay so some things that have worked well are find headwinds, those types of requests. Where we've seen issues is with more complex elements. So I've actually been somewhat disappointed with out-of-the-box LLMs and being able to understand, say, a big bank's risk or off-balance sheet risk in a way that's a bit more nuanced. So yeah, I think there's still a couple gaps in understanding that personally I would have expected to be fixed by now. Yeah, I I yeah, I think the reason for that though is that there is just not a lot of good forensic information on the web to learn from.
SPEAKER_02Okay.
SPEAKER_00Where there is a lot of good information about general risks, like operational risks, etc., but there's not a ton of good forensic accounting stuff out there up until two years ago, when we until we wrote a blog post on it, if you Googled what is a critical audit matter, it would tell you the wrong thing.
SPEAKER_02Yeah. There's sort of a weird circuit too. Like when I'll search for things about buy side and AI, like it'll bring my own blog post back to me. So you probably Yeah.
SPEAKER_00Yeah. No.
SPEAKER_02No, I want someone smarter than me telling me.
SPEAKER_00I don't want to read my own. Yeah.
SPEAKER_02I don't want to read my own blogs. But a lot of Hudson Labs stuff shows up there too. So that's funny, sort of uh getting your own documents fed back to you.
SPEAKER_00Um honestly, the reason we created a blog post about critical audit matters is because it was it was bothering me because I kept reading short reports that would reference the Google result for critical audit matter, which had clearly been written by AI. Uh, and there's all of these short reports going out being like, oh, you know, this company has a critical audit matter related to revenue, ergo fraud. And I'm going, you know, every single software company has a critical audit matter related to revenue. This is we need to we need to clear this up. So yeah, uh, I I do think there is still still a few gaps in understanding. Much, much better than it was a year ago, but in some of the more nuanced disclosure, there's I'm still finding some reasoning error.
SPEAKER_02Okay. Makes sense. Yeah. What um super super super interesting. Um, and I think sort of like a very low calorie way for investors to just make better decisions, like having some sort of like systematic like if I sort of think about like the agent path and the agent stack, like you know, some sort of like forensic accounting risk element in that agent stack is critical. Like if I if I can do this in a seamless way before I make a trading decision, having sort of a risk checklist, forensic risk score, like hey, this you know, there's a one in three chance of this thing being a fraud. Like, I'd like to know that before I make. The trading decision instead of after I make the trading decision, like odds of restatement risk, et cetera. Which those are sort of very painful 8Ks to sort of wake up to. Um, having the systematize in your decision dashboard just seems like an absolute no-brainer to me.
SPEAKER_00Agreed. And one of the things that's cool, we talked a little bit about cost and efficiency. And it's if you try to run Hudson Labs risk scores out of the box, even for us, uh, it costs hundreds of thousands of dollars to run the Hudson Labs Forensic Risk Score over a 10-year period. Um, which is a sort of somewhat interesting element of AI. And one of the reasons that we're now finding it's harder to support quant funds, which I think is a pretty just because you know, whenever you're updating something in that stack, if you have to rerun 10 years of data, uh it becomes harder to get to offer the state-of-the-art functionality to purely quant funds who need that back testing, which I think is a pretty interesting development in AI where we've all been thinking about like, oh, our LLM is amazing for quant funds. Uh, but we end up running into these cost constraints where we're really only able to offer state-of-the-art functionality to people who don't need 10 years of data.
SPEAKER_02Yeah. How's it? I haven't uh how's it work? I mean, so I haven't tried this yet, but you know, it's a conversation that's inspiring me to try this. So I'm a healthcare investor. Um, can I create like a risk skill where it's like a three-page risk checklist on a name, pulling in all of my underlying research on those companies, but also pulling in the Hudson Lab's forensic risk score, pulling from the 52 million, you know, data point vector database via MCP? Like, is that something where I could sort of hit just press that button and get a get sort of a custom-made risk risk report on any of my companies?
SPEAKER_00We should we should chat. Let's chat about uh trying to get you set up.
SPEAKER_01Cool. All right, sounds great. Can I ask about on this cost question? The um there's so much happening, right? So you have Fable that is 10x more expensive than Opus. You have all these, a lot of my clients have become heavy cowork users, and guess what? That's not a cheap harness. Um, turns out. Uh, then you have Chinese models, open weight models. How are you seeing this kind of cost? Like, and and also people care about cost, like people stopped caring about people didn't care about costs three months ago, and now a lot of people, maybe some of the mega funds excluded, uh, don't care. What's your take on the shifting token tokenics landscape as it relates to investment decision making?
SPEAKER_00Yeah, it's it's gonna be interesting. One thought I have is let's all max out our subscriptions in the short term because they're not gonna stick around. Um one thing that we tell our customers, which is a funny thing to be telling customers, is Hudson Lab subscriptions are profitable for Hudson Labs. So you can, you know, trust that we're not gonna randomly hike your prices. Um yeah, it's it's it's it's it's gonna be interesting and it and it does come back to this. We're creating these really extraordinary outputs in part by using very inefficient, very inefficient ways of using AI, very inefficient search, very inefficient, like using uh a million calls instead of one in order to get something you want. And I don't think the fact that it's so expensive and we're not seeing the cost matters. It just means that when this starts to collapse on us and our our prices start to go up, there's gonna be a lot more, a lot of people are gonna start thinking a lot more about AI infrastructure, AI architecture, and what's actually happening on the back end instead of just thinking about that top level LLM, the top level model. It's gonna be that's gonna be the next push, is we're gonna move beyond the generation step and really think about how are we storing data? Where are we getting data from? There's so many ways to make make all of this more efficient, which is one of the reasons why Hudson Labs can run data sets using long running agents, and it you know, it only costs us a hundred bucks. But if you just went to an Opus 4.8 API, you're gonna accidentally spend 13 grand. You know, that it it it all of these problems are are solvable. Uh they just require, they require more.
SPEAKER_01I wonder too, right? Because you mentioned like it's at the app layer, like this this cost decision is gonna have to be made at the individual contributor layer, the app layer, the model layer, the model selection layer. Like, do you think one? I mean, you you have a horse in the race, but do you think like one of the levers matters most on like optimizing against cost or optimizing for cost?
SPEAKER_00There's a reason open source tools like open code are having such a big moment right now. People who are thinking about cost don't want to get locked into one model provider, it doesn't make sense. Um open code is a are they good friends of ours, and and and they're getting huge because developers are frustrated with all if you're you're you're using plot code, only getting to use anthropic models. Anthropic probably won't be king forever. So yeah, if if you're thinking about cost, it's a good idea to make sure that you can you can switch providers, you're not locking in to one model provider in perpetuity. We know a lot of bigger banks, bigger enterprise customers that have built their own routers rather than relying on somebody else's MCP so they do have more control. Uh, when we think about AI and when we're going, it's yeah, it uh it's all about being able to control your own outcomes and make your own choices.
SPEAKER_02We covered a lot of ground, Chris. Any other you know, core debates that you're having with clients or peers or friends or um sort of big questions that you think still need to be resolved that we should talk about?
SPEAKER_00One thing that's somewhat interesting is uh have you tried using any of the big models for extracting soft guidance?
SPEAKER_02No, tell me about it.
SPEAKER_00Or do it, yeah. So we built a guidance model two years ago now that we thought would only have a moat for about nine months. The reason we bought built the guidance model is because LLMs are brilliant, but when they think about tense, they focus when they think about whether or not something's future or past, they focus on the tense of the verb. So we talked about this before, Brett, where you know American Eagle frequently says CapEx is now expected to be. And if you run state-of-the-art models and say, get me all management guidance, it's missing a bunch of guidance. Um and that appears to still be true, which I find absolutely fascinating, especially because so many of those big model providers have been investing in finance-specific training. Now that I've said it on this podcast, it will probably be uh fixed tomorrow, but all all 10 listeners to the Invest with AI. All 10 listeners, yeah, yeah.
SPEAKER_02Although it's probably being scraped into many many agents, I'm sure, probably. Um that's interesting.
SPEAKER_01Um yeah.
SPEAKER_02Yeah, fast fascinating. Okay, great. Um, yeah, why why do you think like I I hear this a lot like the Opus model of agents have more been more specifically finance trained than the Chat GPT? Dick, do you have any sense of like why that's been the case? Why has that been your claw's more finance? Yeah, like is that just the intention of the training process?
SPEAKER_00Like I think a lot of the big firms are now specifically focusing more on finance because A, they realized it was a major gap before, and B realizing it that this is a consumer base actually willing to pay.
SPEAKER_02Yeah.
SPEAKER_00Yeah.
unknownYeah.
SPEAKER_02I think it was a little bit of a galvanizing moment when Anthropic did their a finance day and showed that finance was the second largest vertical behind tech.
SPEAKER_00Um perplexity as well. It's yeah, obviously.
SPEAKER_02Finance AI is definitely having a moment. We even have a we even have a specific finance and investing podcast out now. Invest with AI, available on all podcast platforms, wherever you get your podcasts, like and subscribe, as the kids say. So uh thank you so much for for being with us, Chris. This has been really uh really helpful. And uh part of what we're trying to sort of uh do in this podcast is uh sort of like a flashlight on a dark evening, just see three feet in front of us. I don't really know, I don't think any of us know what what uh what lies around the bend, but um we hope to continue to have you on to sort of see a little bit further down the road and understand at least understand where we're at in the moment. So thanks for thanks for everything you've you've shared. And what's the best way for people to get in contact with you if they want to learn more about Hudson Labs?
SPEAKER_00Hudson-labs.com. Come give us a view. And I'm Chris Binotti, you can find me on LinkedIn, X everywhere.
SPEAKER_02Great. Well, thanks so much for the time today.
unknownThank you.
SPEAKER_02Thank you, Chris. I'll bother you soon with uh more questions, I'm sure.
SPEAKER_00I can't wait.