Selling Signals - the Data Monetisation Podcast
Selling Signals is the podcast for anyone building, selling, or buying data, with a focus on commercialising data in the investor ecosystem.
Each episode brings together industry insiders to share real, first-hand experience from the front lines of data sales. We unpack what actually works when turning raw data into revenue, whilst exploring other data buying silos to break down the walls between them.
Selling Signals delivers practical lessons to help data teams sell better and build stronger, more commercial data businesses.
Selling Signals - the Data Monetisation Podcast
Freeman Lewin: Agentic Data Sourcing
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
What happens when the person sourcing your data is no longer a person?
In this episode of Selling Signals, we’re joined by Freeman Lewin, co-founder of Brickroad. Freeman began his career in private equity and later worked as an attorney negotiating data and AI contracts.
Brickroad uses agents to discover new data suppliers and automate what Freeman calls the “messy middle” of procurement. Its agents search thousands of sources and infer key details about a dataset. They then surface the suppliers most worth investigating.
We also discuss how vendors should think about monetising their data with both frontier AI labs and smaller specialist AI teams.
Welcome to Telling Signals, the podcast focused on how businesses actually monetize and sell data. Each episode, we interview an industry insider to hear their experiences and lessons learned.
SPEAKER_01The spirit is powered by Valcus, the company that transforms your data into investment-ready intelligent products.
SPEAKER_00If you enjoyed the episode, please subscribe wherever you get your podcast.
SPEAKER_01Today we're joined by Freeman Lewin, founder of BrickRoad. Before becoming a founder, Freeman spent years negotiating data and AI contracts with some of the world's leading AI labs, giving a unique perspective on how AI companies buy data. Today he's building BrickRoad to reinvent how data sets are discovered and evaluated by both humans and AI agents. We'll explore what that means for a data vendor, hedge fund, and whether AI agents will soon become the buyers themselves. Freeman, welcome to the pod.
SPEAKER_02Thanks for having me. Appreciate you.
SPEAKER_01Maybe the first question to jump into what gap did you see in the market that led you to build BrickRoad? And maybe give a bit of background of what BrickRoad actually is.
SPEAKER_02Yeah, thanks for the question. So my background, just to start off, I started my career in private equity, was an attorney for uh for a good amount of time. And then about three and a half years ago, I stopped practicing lawful time and uh started building companies with the idea that we could uh transform the internet or the economics of the internet through selling data. Um so I've been in the data space for a good amount of time. Um and one thing that we noticed pretty quickly was that across both you know the alt data set sector, but also when selling to hedge fund when selling to AI labs and and um and corporates, the reality is that this space is dominated by an N times M problem. That means that every buyer negotiates bilaterally with every vendor. Um and every fund runs the same diligence on the same data set. Every vendor answers the same DDQ every you know, hundreds of times. And we saw that the whole thing was running on conferences, referrals, and inbound emails. We noticed at the same time that the ICP for who was buying data is changing pretty rapidly. Whereas it was once was you know the silver-haired guys having the three martini lunches. Um that's hyperbole, but but close enough. Um the data strategists and the people buying data in the last year or two have increasingly become more engineers and researchers and people my age that don't really want to jump on a call with anybody. And they will pay a premium for that opportunity to connect with the data set um in their own time. So at the same time, there's a a couple at the same time, um the that ecosystem works when there are only a handful of vendors, but we noticed that more and more vendors were coming online, especially when you think about data being sold to AI labs and unstructured data. And the the the the data that most both across the hedge funds and the and the AI labs that was most valuable was the data that was being produced as a byproduct of a company's operations. So just to recap, we saw a massive opportunity with this enormous explosion in possible data vendors and the use of data, but the rails upon which data was being procured remained very much on a on a human-to-human basis. So we built Brick Road with the idea that we can transform or make programmatic some of the hard some some of the more uh rote pieces of the data procurement process. We focus on three areas, um, and and that's through what we call a multiplexer technology, which is automation, composition, and smart routing. So in each area there there's some some nuance, but automation is both discovery, sourcing, sample acquisition, uh DDQ, uh DDQ filling out or fulfillment, um, and and some of those initial pieces. Uh composition is when you have many data sets, um that many data sets that can be floated to one. So we automatically compose data, multiple data sets into one single one. Um and routing is is you know the best par the best correlation is with open router. So open router routes uh models to the best possible, routes users through the best possible model. Um we route data uh based on on uh our and buyers' needs. Um we give them the best data through our system.
SPEAKER_01Your first point there around consumption, uh, from what I understand for BitRoad is is you're sort of sending agents out uh to different websites uh et cetera to collect information on what data assets could be there. Do you feel like you're maybe adding to the problem and sort of bringing more and more providers to the space, or is that sort of solving a different problem of trying to find differentiated data sets?
SPEAKER_02Yeah, it's a good question. I think the idea here is to find differentiated data sets and and and our discovery tool is just to just to be clear, our discovery tool predicts when a company will come online will likely have data, likely has data as a byproduct of uh their operations. Um for a lot of funds, the existing paradigm uh of going to you know whether it is uh directory directory A or B or C or marketplace, um they their position is that that access to that data has a decaying element. As soon as it is on those websites, um the their position is that everybody else probably has it and it is less valuable. So our goal with data discovery is to predict when it is at or when a company or potential supplier uh is about to come online and enable the the our clients to go out and start those conversations before anybody else. Um whether we are adding to the problem or or not, I think you know the end solution is to is to speed up the data procurement process as as a whole. Um and the first step is to get people comfortable with using agents just generally uh to do procurement.
SPEAKER_00So you mentioned both um discovery um and sorting, and then all the way to collecting a sample. Obviously, as you said, at the moment that's broadly done by humans. Um, and at least in my mind, part of the reason for that is for a data provider to give you access to even a de minimis sample of the data often requires kind of an element of trust of like who you are, why you want the data, why they're provided providing the sample, you're not a competitor, etc. So I'd be interested to know kind of how far do the agents go with say a provider who hasn't made a data sample say downloadable from their website. Like will you go so far as to have an agent emailing uh someone within the vendor company and saying, hi, I'm an agent working on behalf of BrickRoad to um again, we think that your data might be of interest to uh hedge funds for the following reasons. Here's the information we've been able to gather, but we what we would really like at this point is a thousand row sample because that's what our um uh the data buyers are going to need to uh at the very least have some idea of what's in the data set. Is it is it going that far? And uh or is it even going further than that?
SPEAKER_02Yeah, so it at the moment it stops pretty much there. Um the the idea here is to give actual information to our clients. Um and actual information can range from just insight into what alternatives are out there, um and that can just be a list of you know the 300 alternatives to Bloomberg, for example. Um, or it can be going as far as to collect the sample and automating that process. So um it all just to just to back up, the the idea here is to automate the messy middle. And we believe that data strategists are are some of the smartest people in the world, um, and they have this intellect, which is you know this spark of genius, and they have this ability to discern what is really high quality. But that messy middle is going through the process of uh sorting through 4,000 sources or 4,000 pieces of information to quantify whether uh supplier might have a good data set and what the history is on that and what it map how it might map particular, and then for data set and then for you know schema. Um that messy middle, we can do a lot faster with machines. And then, as you see, as you alluded to, we just launched the ability. So all of our agents have their own uh uh email and inbox, um, and we just launched the ability to craft sequences so that you know if you want to contact 3,000 potential suppliers and build that top of funnel pipeline, you can do that through our system uh and craft your sequence, your email sequence that way. We also have um you know, we we call that piece the out the outbox, any anything outside of discovery, we call that our access agents. Um, and the access agents have both the ability to do emails with human approval, always in the loop, um, but also the ability to scan for access points. So some providers um have API access points, some providers have free API access points or paid, depending. Um, and and we share them, and the agents, if if given permission, can go and just tap into those API access points, for example, and get a sample of the data without having to interact with um anybody. That's that's we have to think about it in sense we understand trust is is super important, and that's like really the name of the game. Um, but increasingly people are making their websites and their data sets uh accessible so they don't have to go through that long sales process, at least at the very start.
SPEAKER_00And from the um the uh buy side or sort of the data buyer side, how should we think about your interactions with that with them? Is it the case that, say, a data strategist at some large uh comp shop um would interact with the agents and say, I'm looking for this type of data, and essentially uh your agents go off and they streamline that process and rank potential data vendors so that potentially their time, their human time is better spent looking at the highest likely matches. Or is it an even broader brief than that? Is it simply that you've got you know uh your own internal categorization um uh or data categorization types, and they're almost able to access a catalog that's constantly being filled up by these agents?
SPEAKER_02Yeah, so it is the first. Um we don't maintain uh our own catalog. Um and and as a result, you know, we started because there was um during discovery with with with you know our current batch of clients, um, there was a want to find net new and novel data sets. So what the the the interaction on the buy side is the the buy side will you know type in a query, just like they would email their junior associate. I'm looking for you know US equities data map to take care of XYZ ABC. Um the agent goes out and within a you know between 30 minutes to an hour goes through up to 4,000 sources and ranks on based on the basis of novelty. Um so novelty is uh a multi multi-pronged approach that that we've designed, but effectively it over-indexes for providers that are producing data sets or or valuable data as exhaust from their existing operations. Um so so you know, a company that has an image app, for example, a photography application, um, is a likely it is likely has rights to that photography that's coming through the application. Um, and um as a result is producing those or or has access to those images as a byproduct of their existing application. The same as games, the same as you know, PLS data, uh you can think about the whole gamut.
SPEAKER_01So let's make the assumption then that as we go down this road of um more and more agents in the world and we're we're automating more and more workflows and discovery becomes this way in client. If an AI agent landed on a vendor's website tomorrow, what information would it need to successfully evaluate whether a dataset was interesting or you know, per pertinent to the request of the user?
SPEAKER_02Yes, it's it's an interesting question. I think people conflate traditional STO with agentic actions. Um at the end of the day, uh agents are simply tools to solve a human process. So you need to think about what questions the human behind the agent will be asking. For example, if you're selling to quants, most people listening to this probably know that you need to that that the quants need deep history. Um now with agents, the difference is that agents can infer. Um, and and really that's what makes Brick Road agents special, is uh they carry this diligence intuition uh of a data sourcing team, not just the crawler's ability to read a page. But to answer directly, um the agent needs the same things that a human analyst spends you know multiple calls extracting. So what is in at what is actually in the data, field by field? Um, how far back does it go and is it point in time? What entities does it map to? Um, and how often it updates. Now our agents can infer a lot of that stuff, but um you have to think that when you are being exposed to a deterministic agent, it is going to jump and is going to move on if it doesn't get the questions answered that it wants, that it's told to answer. Um, so making sure that you have some of that information on the website, or that you know, for us, we we we use third-party applications and insights as well, but making sure that that information is available will lend yourself to a greater chance that the agent returns that back to their human.
SPEAKER_01And and does that need to be really obviously clear on websites, or can that be hidden in sort of the code behind the website? Like where does that need to be present and uh and how explicitly?
SPEAKER_02Yes, it's a good question. Um, we run our own tests on like agent crawlers and and and what is what is working and not working. Um we have a variety of properties that we we do these tests on. I will say um that for for all of the um holo hub of you know lm.md texts and and um and markdown and site maps and apis and things like that, the most important thing is your HTML and and so what's on the what's on your page. Um but it doesn't have to be explicit. So it would be great if everybody could say we have uh 10 years of uh of history. But for us at least, for our agents, we're we're really not looking for those established data sellers. So, you know, if you're if you're somebody that wants to fly down under the radar, like one of our agents picked up recently, like a classic one that our agents pick up is how long has this uh app been on the app store? Has it been there for 10 years? Is there an increase in downloads in the last five years? Okay, so it has 10 years of history and it's really scaled up five years ago. Same thing if we know that the company has been incorporated 25 years ago and the website was built 10 years ago. We can make those inferences. It doesn't have to be explicit, but I will say some of the things you may well do well being explicit about like you know, schema map and and whether you map to ticker and stuff like that.
SPEAKER_00How close do you think we are to an uh agent just handling the whole process, like everything from um strategist makes a query all the way to uh the purchase of the data set? And do you think that there are any particular types of data sets where that will happen much sooner than others?
SPEAKER_02Yeah, so it's a complicated um question because exactly how you said what you framed in the last part of it, which is there are certain industries that are going to adopt this a lot faster. I think that the financial services sector is going to be the last adopter of the full cycle, and for good reason. You know, there's a lot of money made on making the right data decision. Um, and and right now we're sitting at like what six months to a year for acquisition of a good data set. Like it takes a long time because there's a lot of bad testing, etc. Um, we are that's why we built, and that's why you know what's public now and it's publicly available now is that like first wedge into data discovery, because people are already searching on the internet, people are already using LOMs for data discovery. We thought you know we can build build a better tool and then just layer things on top as people get more comfortable. Um but the thing, but to answer your the last question, things that are part of the workflow, um, as agents become more integrated into our daily operations, whether it is uh Claude, book me an Uber or uh you know um uh something else, uh book you know, make a flight, the agents are going to need the context around that to make the right decisions. Our thesis in that regard is that we will have we will have no patience for an agent that makes mistakes. So an agent is going to over-index on being accurate. For example, uh if I book an Uber, if I book a flight to JFK, I'm in New York, so if I book a flight to JFK um and I ask my agent to uh book me an Uber and my flight is at 6 p.m., my agent needs to know the traffic patterns of of New York to know that I need to leave at 3 p.m. before before traffic. Um if my agent messes up and books me uh an Uber at 4:30, I'm never going to use that agent again. So we think that for like the basic non-economic um operations that that are intrinsic to our life that we are deciding we are beginning to use agents for. Agents will buy data um through API access or other things, and then we'll do the full cycle. And in that instance, they don't need you know the backtesting and the diligence, etc.
SPEAKER_01And do you think the onus is on the foundational model to pick providers across different data set types that is trustworthy, you know, of good data quality to be able to make those types of decisions, or or will that fall on someone else's plate, i.e. the use uh a third party to to distinguish that that credibility?
SPEAKER_02Success will be determined by um by probably the intermediary. Um and and so you will have multiple uh multiple agent harnesses that will uh seek to differentiate themselves in in terms of their success metric by adding data. Specific data sets and battling over the proprietary data sets. I don't think that I maybe in the next year or so you'll see an influx of uh user users directly purchasing API feeds. Um, but I think you'll probably get to a middle ground of of agent apps uh um seeking to differentiate with higher quality data. I I don't think the foundational models have a have an incentive to purchase you know real-time data like that.
SPEAKER_01As a consumer, I'm really excited to the moment where an agent can just buy me some new new shower gel when that runs out, some some certain groceries that I I use every day where I don't have to physically go onto Amazon or go into a shop to go and buy them just time at my door as as they uh uh as they're finished. I'm excited for that to be off my plate.
SPEAKER_02Yeah, agents have been being used in Amazon's workflow for for for a long, long time. Even uh the predictions on whether uh on where to put FBA shippings or fulfillment by Amazon, um, where to put you know your shampoo, whether that is in Bristol or London proper at the warehouse in London proper, that's that's a agentic decision that's been made. It's been being done for the last 10 years.
SPEAKER_01We actually have a guest coming on in the future that uh that created the recommendation model for for Amazon. So we'll we'll dig in with uh with him on that as well.
SPEAKER_02The head of uh the the former director of AI at uh you know University College London was the one that made the warehousing model for Amazon.
SPEAKER_01Awesome. Well, as we're talking about AI, let's move on to your experience of working on sort of data licensing agreements at the big AI labs. I think the sort of the big AI license deals that they're incredibly publicized. I'd be interested to know from your experience, is the is this noise or are AI labs actually going to be the sort of next frontier of data buyers? And maybe you could split that question into sort of as a training data sets, as we're more commonly seen now. And then I guess we've touched on quite a bit of sort of a genetic buying, but yeah, how you'd separate those two.
SPEAKER_02Yeah, so it's it's a funny one. I think thanks to the large publicized deals, or thanks or no, thanks to the large publicized deals, um, there are the the the pace of content creation it has really outstripped demand. So there's a plethora of uh venture-backed startups that are having a little bit of a race to the bottom, to be honest, on pricing, um, but they are collecting a huge amount of data. Um I don't think that the frontier is selling to large-scale AI models. I think that the labs that are buying, I think that the real frontier and the real, like it's really exciting space, and the the place where there is a sizable profit margin is in the long tail of you know, one or two person labs, well-funded one or two person labs, buying specialized structured, rights-cleared data for post-training e-vales and grounding. So that's small, that that's like the hundreds of thousands of smaller um labs across the world instead of the and in that those instances, they are valuing the full rights-cleared um data sets, they are valuing the same things that a hedge fund would value, you know, latency, access, um uh up-to-dateness, uh economic indicators uh uh uh of a return on investment. They are not necessarily um the same. They are signing long-term deals instead of the labs who are really just signing one-off deals and making everybody fight for for the ability to add their logo to their deck.
SPEAKER_01Interesting. I saw yesterday, although I think that the job's been published for a bit of time now that Meta were advertising a sort of data sourcing type role. I think they've called it a data partnerships business development role. But um, it feels like wider than just the foundational models businesses or sort of AI labs are starting to think about how they can formalize that that sourcing role.
SPEAKER_02It's it's an interesting one. It's it's relatively new. Um, so when I started, data was purchased through um Financial Strategy. And then like one or two labs got partnerships. But it's only in the last, I would say, few months that we started to see job posting and then rules um at the labs for specific data sourcing or data strategy, um, specific to training. Um it's a funny, it's it's I I I I was thinking I think about this a lot, but the the talent pool that goes into data sourcing, data strategy at the funds or financial firms, um, you know, oftentimes there's a talent pool through Eagle Alpha or New Data or Data Rate that goes and effectively sits at second mids at the funds. Um we didn't see that, we didn't see like a uh that that funnel to the foundational labs until pretty recently. And now we see um Turing and Invisible sort of pushing their people into second mids at uh at the funds. So we'll we'll likely see more of those rules come up. But I think you'll probably you're you're you'll you'll probably touch on this later. But um at the end of the day, your ICP for uh for a data deal at a lab is is not the data sourcing guy. It's actually the researcher at the lab who's who's going to be the first one to find it and dictate whether it whether it's high quality or not.
SPEAKER_00From the investment side, and obviously a lot of our listeners are data vendors to um investment funds, the industry has got to a point now where it's mature enough that generally vendors are going to know that there are quant funds, there are fundamental funds, the the quant funds are going to need kind of breadth of coverage and then uh a reasonably long tail of history in order to properly test the data. The fundamental funds are going to be more interested in data sets that give them very deep insights specifically on the stocks that they care about. You mentioned that in the um on the AI lab side of things, the big frontier labs probably aren't the new market that these vendors should necessarily be thinking about, but that it's these, you know, tens of thousands of um of smaller uh labs that are springing up all over the world to build these speculate models. How should vendors be thinking about one, whether their data might be of interest to them, um and two, how they should target those funds? Because I don't I I think probably most people aren't even aware that that that this is going on. That everybody um who's not working in AI at the moment is um probably familiar with Chat GPT, Claude, maybe deep seat, et cetera. But then they're not aware of all this other stuff that's going on.
SPEAKER_02It's my apologies. Can you repeat the the last part of that question? It's not a linear.
SPEAKER_00Sure. Essentially, how how should they think about whether the data um I'm a let's say I'm a data vendor, how should I think about whether or not my data would be of interest to a specialist um lab? And and then two, how would I go about finding who that might be? Do I just have to be really uh switched on about um like where the technology is going and and and therefore be able to kind of um I don't know if I lead on that front and go and find them?
SPEAKER_02Yeah, yeah, okay. Um yeah, so then there's we'll have a thesis of the value of their information. And um whether that thesis is derived from the value to a quant or a fundamental researcher, or it is to a model, it's generally the same thesis that this will help somebody determine their answer question. And so when you're thinking about it in that sense, let's take for example, um uh a weather data company, for example, they know that selling to a quant is valuable because they are able to produce the weather by 30 percent, and as a result, the the the quant can know you know what same day sales look like at Walmart. That's one thing. Um the in the same way that that's their their thesis on selling to the quant, you can uh you can just transpose that to selling to a model in the sense that if you are doing an economic operation, if you are if if the model is sitting in the economic value chain of determining or helping users determine whether to do things um that day or that month, there is value in knowing the weather. Um so I don't think that it's a massive shift in terms of uh knowing whether your data is valuable. If you think your data is valuable, that there probably is a place for it at a model. It's figuring out which which labs are working on this specific problem and also understanding the value to these labs. So the value to a hedge fund is pretty black and white. It's we are going to make money off of this data. The value to a um a frontier or foundational lab, you know, open AI, typically isn't we are going to make money. It's typically we are going to be able to produce better results, um, uh uh, you know, or closer results to reality we are going to save on compute by having like the direct answer. Um and and accordingly making sure that you are uh you are positioning just like a just like great data vendors that are selling to hedge funds will do their own backtesting to evaluate the value, you should probably do be doing your own um your own testing uh to uh evals to figure out the value to a model as well.
SPEAKER_01And to that point, I think hedge funds have a fairly well-tuned backtesting sort of model process. How does that work at an AI lab? What are they looking to assess? And yeah, how can a provide, I guess how can a provider sort of do that work up front to understand what whether they're valuable or not?
SPEAKER_02Yeah, yeah. So in the same way that also doing too much backtesting probably isn't valuable, uh doing too much uh um doing too much evaluates probably isn't valuable as well, and you and you might want to leave some of like the the long tail task to to the fund, the what you're trying to get to when you are selling to to the the labs is a fast yes or no. Um and as a result, uh figuring out where your data set sits in the stack uh of operations is super important. So if you are selling this you know unstructured data set at like millions, hundreds of millions of images, then that's probably a pre-train training task. So you want to do uh something like you know do an eval on on pre-training and and and compare results there. Um if it's another type of data set that is better for task improvements. You want to do like a small tas there, you'll have evals and retrieval. There's there's a bunch of different like tasks that go into uh that that the different labs specialize in. Um and and um and that that you should be thinking about how your data set uh can can augment that. I will say the the best thing to do is do something. Um what is really bothersome to the labs right now is vendors coming in saying, we have all this data. Uh can you take a look at it? Well, the the the data strategists or the the financial strategy guy at uh um one of the big labs is getting 300, 400 emails like that. Um the best way to to to think about it is is who is my ICP for when you're selling to AI labs is is researchers. And so putting up metrics on something like Hugging Face and putting up sample on a hugging face or another platform um will go a lot longer, go a lot farther in terms of appealing to your ICP than you know doing a cold email and saying I have a hundred million images, for example.
SPEAKER_01And in the hedge fund world, providing a level or all historical data is kind of the gold standard, uh, purely because, especially on a quant side, that they they can't really make any money until they have a live, up-to-date feed. Is that what what what what should uh a vendor selling to uh an R lab be thinking? Because that historical data seemingly feels like the value.
SPEAKER_02Yeah, historical data is. Um, it's actually like a great differentiator, which is um in the finance world, the expectation is that they're making money off of the feed. And they are not, they are just using the historical data uh for backtesting and testing. Um in the in the AI world, that historical data, especially if it's priced in the right way, is more valuable than anything. Because you have to think about it, it it the labs are not really that incentivized to give up-to-date data. There's usually a cutoff of you know, whether it's a day or two, it's not real time because they're not doing trading. Most of them aren't doing trading. Um, so the historical data is what trains the models and gets and and and remember that you know most of the labs are are LLMs and they are doing statistical queries, and um, as a result, the historical data is really the most valuable. I will say that back when we were doing a lot of sales to to AI labs, a lot of it was can you find people um again that are producing data as as uh um as uh you know as a byproduct of their operations, but a lot of it came down to how much historical data do you have? We really don't need a live feed. We just want the entire back catalog and sell it to us at uh at you know pennies on the dollar because we know that you're just sitting on it, we'll pay the we'll pay the transfer cost, we'll pay you know the amount it is to for your storage cost, but we're not gonna pay that much more. That's that's also part of it.
SPEAKER_01Interesting. Okay. So we think about we've gone through the testing, you're moving on to contracting, uh, as you sort of just alluded to, you're negotiating price. Where should providers be really vigilant in terms of licensing terms to AI labs? Is there anything that's yeah, uh as a as a business selling that they they should be really considerate of Yeah, so the first question you need to ask yourself is why is this lab buying my dear?
SPEAKER_02Especially if it is public. Um, that needs to be like a come to Jesus moment for you of why is this lab talking to me? Why are they wasting not wasting their time, but like why are they putting the time and effort to to contract with me, especially if it's gonna be you know uh a few hundred thousand dollars, especially if I haven't done the testing required, um especially I don't know the value of it. What we see, what we have seen over and over and over again, is there is an attempt to legitimize what they their their previous scraping habits um or their current scraping habits by uh entering into small and medium-sized contracts with providers and including broad, broad-based uh release language or identification language um related to related to the data. So we've seen things as egregious as you are going to release the lab, its investors, its customers, and all future users from any lawsuit related to uh this data and any data that you ever produce. Um and if you're not paying attention to that and your motivations are just you know to get to Series A or something like that, you can end up really putting yourself in a box because as soon as they purchase it at one time, then they're they're just gonna continue to scrape you for forever. So we've seen that sort of predatory behavior over and over and over again. Um and I think that's that's by far the most important thing to keep in mind is just because a massive lab approaches you to be nice to you doesn't mean that they're doing it sincerely or genuinely.
SPEAKER_00So what what should what would you um recommend a vendor come back with then if the uh you know some large lab comes and says exactly that, we'll give you half a million dollars today, we'll pay really fast, but we want these really, really broad um terms. Is there what what's the model that's beneficial to the vendor? Is it a a yearly subscription and that their right to utilize the data continue so long as they continue to pay, that sort of thing?
SPEAKER_02It really depends on the business motivations. So we talk to a lot of YC companies that get approached by the major labs to license um with this pretty predatory thing. And the reason why that they say yes to it is they just want to prove to their investors that they have a named brand on their deck that they're selling data to. They they're thinking like short term. Um we also talk to a lot of companies that will enter into MSAs, broad-based MSAs, um, for data contracts, you know, up to $50 million. Uh, these are the terms. And they'll go to, again, their investors and say, look, we have an MSA with you know TikTok. I'm just using a random name, but we have an MSA for with TikTok for $50 million. But the reality is that at the end of the day, they have to fight for every deal. And we have, in fact, seen issues where uh people, uh uh startups especially or technology companies especially will sign a $50 million uh MSA and really only extract $5,000 out of that. Uh, the best thing to do is be prepared to walk away and not be in a position where you have to accept a deal uh from a main brand like that.
SPEAKER_01Not to make it feel like it's the AI labs being predatory, but I saw um I can't remember which meet the US media company it was, but they announced a hundred plus million dollar deal with OpenAI. And reading the sort of news article that came out about it, it seemed like this media company was was suggesting that you guys are scraping us. We will get a large lawsuit lawsuit involved if if you don't sort of license this properly. Uh and it just felt like a buyer or or will spend time in court burning dollars, and it's it seemed like a sort of handshake deal was agreed.
SPEAKER_02Yeah, most people aren't in the position to threaten a hundred million dollar lawsuit. But if you are, for example, a company that produces, or you are a publisher especially, that has a copyright to all of this, whether it's news sources, or if you're getting images and uh and you've attached a value to each photo, then it can run up into the billions of dollars and and you have that leverage uh very certainly awesome.
SPEAKER_01Moving on to the last bit, I think a lot of conversation that we've had on the podcast recently has been around consumption models uh in terms of commercial um licensing of uh or paying of uh of data. It seems like where you sit on the market, you're you're probably pro consumption-based models.
SPEAKER_02Um yes. Agents meter naturally. Um so when a human licensed a data set, the unit of consumption uh like a C the unit is consumption. Uh a C a firm or a year really isn't like a good model for agent buyers. Um and I think like in a lot of ways we are moving to the consumption model across a lot of a lot of the things that we we do, um, whether that is Is API pricing for tokens, or you know, the most basic, which is you know, uh uh uh you know, our cell phone bill is sometimes our consumption model. Um, but I would push back on computer consumption, especially for your clients that are building businesses around data set of sales. Um enterprise buyers, which are usually the most valuable and the most stable, um, they really don't like consumption models. And um, and and they're they are unbounded by spend. Um, what they really care about is stability. And so so I wouldn't say it's an all or nothing thing. I think there are some companies, I believe Data Lupa does it, but they they rope people in through consumption. And and I think that that's a good like starting point. If you are able to say and put on your website that buying, getting you they that the agent can access this data for X cents on the dollar um up until a certain amount, that will get you in the front door uh a lot of these labs and and you know some of these traders that are at the small shops, um and and then you can go after them and proof your work that way.
SPEAKER_01That's interesting. So you kind of see it as as both models still existing, but splitting it out to as the sort of longer tail of uh of prospects sort of come to market to buy more and more data, that their entry point is consumption up until a point where it probably becomes more cost effective to have a standard license.
SPEAKER_02Yeah, going back to like how how we started, the ICP for who's buying data is changing. Um, even like even if I were to take agents out of the equation, um, most people, most millennials don't really have the want or desire to jump on a call with somebody. Um, and a consumption model will get them over the bull over the door for them to test. And and the reality is like we are in a place in in technology where testing is is the entry point, and you want to be at that entry point. Um, once you get to a place where you come to rely on this tool as the basis for your operations, um, then then an enterprise arrangement is is the best model.
SPEAKER_01The the thing that scares me as a sales rep around these conversations is where a sales rep fits in. Uh and and maybe this is where data sales changes, but where do you think a sales rep fits in on in a consumption model, if at all?
SPEAKER_02I think that sales reps will be vaulted up the stack. Um, I think that you have to think about the consumption model as the um uh the the top of funnel move. Um so if you can if you can get on the pall with somebody that has already purchased your data, has already played an experience with that data and paid you something, then you're then then your job is that much easier, right? Um so I don't think that we'll see, you know, look, my my company is is effectively six months old. I do all the sales, I'm not great at it. Um I see value in salespeople in a big way. Um, but I also think that uh um they'll be higher up the stack and and probably responsible for the larger deals and the larger enterprise relationship.
SPEAKER_01Do you think that makes the barrier entry into being a sales rep that much higher if it becomes more of an enterprise motion rather than a SB or a mid-market? I think more of like a traditional SaaS sales journey.
SPEAKER_02I think this I think this space will benefit from having more qualified and higher better higher uh more qualified and better salespeople up the stack um uh than you know the quotas that that some people that some firms put on them. So yes, I I I think it I think it will create a barrier of entry, but maybe for the best.
SPEAKER_01Awesome. Well, I appreciate you coming on, Freeman. I think this was a really interesting conversation. I I think the the the way the industry going is going is is somewhat scary, but you guys are clearly at the front of it and doing some really interesting stuff.
SPEAKER_02I appreciate your time and uh uh thanks for having me.
SPEAKER_00Yeah, thanks very much, Freeman. It's great to have you on.