Trends from the Trenches
Trends from the Trenches
Episode: 44 - Knowledge Graphs Turning Life Science Data into Answers
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
AI is moving fast in life sciences, but a confident answer is not the same thing as a correct, reproducible, auditable answer. Knowledge graphs, ontologies, and FAIR data practices are the keys to bridging this gap. Knowledge3's Tom Plasterer, CEO and co-founder, and Eric Little, chief data officer, join host Allison Proffitt to discuss what it takes to turn semantics into something practical: a knowledge product mindset, modular delivery, and “semantic ops” that can better manage data systems. Their conversation challenges the hype cycle around context graphs, explains why life sciences are a natural starting point for regulated, high-stakes AI, and shows how layered models can capture the right level of context without building a giant monolith that never takes off.
Links from this episode:
Bio-IT World
BioTeam
Bio-IT World Europe
Knowledge3
Bio-IT World’s Trends from the Trenches podcast delivers your insider’s look at the science, technology, and executive trends driving the life sciences through conversations with industry leaders.
Welcome And What’s At Stake
Allison ProffittWelcome to Bio-IT World's Trends from the Trenches Podcast, your insider's look at the science, technology, and executive trends driving the life sciences. I'm Allison Proffitt, editor of Bio-IT World, and today I'm talking with Tom Plasterer and Eric Little, both with Knowledge 3, a company Tom co-founded late last year. Welcome, both of you.
Tom PlastererThanks so much, Alli. Great to be with you. Yeah, thanks a lot.
Allison ProffittTom, I'll start with you. You have been beating the FAIR Data Knowledge Graph Drum at Bio-IT World as long as I've been around through AstraZeneca, Exponential, and now Knowledge 3. Tell us about the new venture and what made this the right time to launch the
Why Knowledge 3 Launches Now
Allison Proffittcompany.
Tom PlastererYeah, so thank you. So, really, what we're trying to accomplish and what we're we're building at Knowledge 3 is a knowledge product and services company focusing on life sciences. And what we're really trying to do is make it seamless to get real value out of your knowledge in a very much of a plug-and-play environment. Yeah, but this is something that I've been heading in this direction for many, many years and trying to really close the gap between scientific intent, so what your subject matter experts, what your strategists, what they want to accomplish, and actually realizing that answer in real ways that transform the business of science. I think what we found is that sometimes technology can get in the way. Frequently, technology does get in the way, and we have to come up with ways of making that easier, making it modular, more reliable, more repeatable. And that's where fair and knowledge graph really comes into play. So that's that's you know, broadly speaking, what we're trying to accomplish as the knowledge product company.
Tom PlastererIn terms of the second part of your question, so why now? Why is this the the right time to do this? If you'll indulge me a little bit on the personal side first, yeah, this was a direction I think I've probably been heading in for, I don't know, I'd say probably the last 12 to 15 years. And then toward the the end of my AstraZeneca tenure, I met my other co-founder, Ivan Snakeov, and he's got a very strong background in data engineering and has been just a superstar consultant. And he really introduced a couple of concepts that we now use all the time within Knowledge 3 around problem decomposition, breaking the scientific business challenges into as small as pieces possible, and then building them in an enterprise way so you can just redo it and make it very strong dev ops, which we're now making into semantic ops or semops as a way of kind of bringing that together. So Vivana and I joined forces late in in my tenure at AstraZeneca, built some of these ideas before we both went over into exponential data to show that these are actually commercially viable and make a lot of sense. And then you know, the time was right when that's when exponential data was acquired by Genpact to try this on on our own. And so that's kind of the genesis personally, and we'll talk a little bit about I'll let Eric talk a little bit about when he joined us.
Tom PlastererBut in terms of the right moment for the technology in the industry, I think the biggest thing is really that we're at an interesting inflection point around artificial intelligence, large language models, and agents. We've all seen the promise, and we've all seen the power, and we've all seen kind of, at least in in our industry, the real challenges around hallucinations, the real challenges around reproducibility, the real challenges around audit trail, and all of those pieces. So there's such an appetite now for knowledge grounding that having a knowledge graph and a knowledge product company that really kind of focuses that area that's able to work synergistically with agents, with LLMs, it's like the timing is perfect. And then given the fact that you have such a need for strict provenance, strict audit trail, really deeply understanding domains and life sciences, this is now kind of a perfect fit for us. And so we're we're taking the the many years that we've come together with this technology, the fact that we've proven it out in multiple different you know, pharmaceuticaltechs, and that's kind of where we're getting started with a couple of big clients. So we can we can unpack any of that. So I gave you a giant mouthful.
Eric Little On Semantics And CDO
Allison ProffittWell, Eric, you have a very varied background, maybe more varied than I even realized. But most recently at OSIS, and I guess most recently at Accenture, but now as of just a couple weeks ago, you are the chief data officer at Knowledge 3. So why the move? What pulled you to knowledge three? And what does chief data officer look like for you there?
Eric LittleYeah, like Tom said, thanks for having us. I think there's, you know, I'm gonna echo a lot of what Tom said, and he he said it pretty well. I've been I've been involved in semantics for pretty much all of my career. So my one of my PhDs is actually in philosophy. So I go back to the dawn of you know, where you know, this all started with Aristotle's original 12 categories through Husserl and and everything like that. So for me, you know, why I'm here is I'm a I'm ultimately I'm a product person at heart. I was co-founder and CEO of a of a of a similar company some years back that actually won some some serious awards. We made Fast Company's number three most innovative data science company in the world with this idea of virtualized semantics, which at the time everybody thought was crazy and couldn't be done. And so for me, you know, this was just a really great opportunity to get back into that mindset and to be able to pick up on this. And as Tom said, you know, with Yvonne being our CTO, I was super impressed with the technology. Every once in a while, you just meet people that I mean, it's just this kind of perfect storm of everyone shares the same ideas, the same drive, the same ambitions, and it all aligns, and we're all thinking the same way in the same directions. And, you know, we have some pretty lofty goals to attain.
Eric LittleFor me, spending the last five plus years at Accenture, you know, I learned a lot. I got to work with a lot of great people there, and I got to work with a lot of great companies. But when this AI thing took off some years ago, I was by myself in the firm waving my hands around to anybody who would listen, saying, This stuff isn't magic. And I've seen this before. And so having done my postdoc in in industrial engineering and multi-sensor data fusion, for example, you know, a lot of this AI stuff looked very similar. It's it's all based on statistics and things like this. And I kept explaining to people that you can't solve all of your problems with mathematics and statistics. And I understand that this this thing, you know, looks great, right? Like the as the old as the old saying goes, you know, , there's a point where you know technology is indistinguishable from magic. And I think I think we're a little bit in that space yet. I think everyone is just so enamored with the fact that you can ask this thing a question and it gives you these verbose answers and it can do this quick internet search. It's it's like it's like Google with a with a brain and it or a mouth or something attached, you know. Right. But it's not really thinking. What these things do is they're they're like a parrot that you would teach a 10 billion word vocabulary to. It's gonna say really profound things to you, but not necessarily know anything about what it's talking about. Whereas semantics is all about meaning. And so for me, what CDO looks like at this company is I bring to the table the ability to think about this from the data, the models, the carefulness of the models, logics, reasoning, and how all that stuff applies into engineering practices and and technology and such.
Eric LittleAnd so Tom and Yvonne approached me, and it was just it was just an offer I couldn't refuse. So I decided to come over and you know, if they would have me and put up with me and you know deal with my idiosyncrasies, then I thought we could probably build something really, really useful here. And I think what the world needs right now is they need something that generates more value and answers, a lot less promises and hand waving. And oh, don't worry, it's gonna be great in six months. Most of the clients I've been working with, you know. I mean, you've seen the statistics, 95 plus percent of all these projects can't get out of POC land. And so what people need is they they need everyone's got this data engine, and then they've got this, this, you know, all their legacy engines, right? Their data lakes and their cloud and their all their fancy stuff, and now they've got their LLMs. But there's they're missing this cog in the middle. They're missing this thing that can that can put those together and form a flywheel so that both of those engines can work in combination. And to my mind, that's what knowledge three has. So it's a combination of having really top-notch technology and also having some of the smartest people I've had the pleasure of, you know, being around. And so I think that we form a a pretty formidable team. And so that was really exciting to me to get back into something small and and I think focused on exactly the point in the industry that everybody needs help with.
Allison ProffittSo
Why Start With Life Sciences
Allison Proffittneither one of you are describing something that is life sciences specific. I mean, I'm not even sure you're describing something that is science specific. These are problems that are across all industries and verticals. Why focus on life sciences, or am I wrong? Are you not focused on life sciences?
Tom PlastererYeah, that's that's that's an easy one. That that's we're focused on life sciences because that's the the area we know best.
Allison ProffittOkay.
Tom PlastererBut it doesn't necessarily mean it's the area that we're always going to be in. You know, even kind of thinking about where I've spent a decent amount of my career, if we were that narrow, it would be biomarkers and translational medicine. But you know, you quickly kind of go, oh, well, what's the next part of the pharmaceutical development chain that you need to interact with? And so then like how we were doing knowledge crafts, knowledge engineering at Ashes, and I could start with clinical operations, then to biomarkers, then to translational medicine. And then eventually we landed on oncology strategy. So you kind of can expand that as you move with the business. So we've proven that it works, you know, within the the pharma environment, all the way up to sort of real-world evidence and regulatory. But I agree with you, there's nothing about the approach that's necessarily specific just to pharma, just to life sciences. And certainly Eric's been in other domains, you know, Yvonne, our other co-founder, has done knowledge graphs in finance and energy. So we could we could head in that direction at some point, but we want to, you know, make sure that we can call walk before we run.
Allison ProffittOkay. You didn't meet at Bio-IT World, Eric said earlier. So so I appreciate you coming to serve our area first.
Eric LittleYeah, I think I think the I think another thing to point out is is for me the the areas where AI needs the most help are are heavily regulated industries. So where AI, you know, we get asked sometimes, like, you know, well, well, can't I just do this just with AI or AI isn't, you know, we're not we're not saying AI is bad, we're just saying that it's limited. There are areas where where I think AI can be more successful on its own. Like I've seen it do a lot of good work in marketing because the answer doesn't have to be very exact, right? If I say, hey, I want you to come up with a with an ad campaign or some slogans or draw me some some imagery, you know, that shows you know passengers on an airline having fun or you know being comfortable, it can do that really, really well because it can mimic a lot of the things it can find on the internet. But when you get to, you know, reading scientific literature and the thing trying to parse out the facts from the conjectures, like when it says, Well, we know that XYZ is a biomarker for this disease, and now I'm looking at whether this biomarker is also , you know, this thing is also a biomarker for a different disease, they get that wrong a lot.
Eric LittleLike it will say, Oh, that thing's a biomarker for both diseases. It doesn't understand the the conjecture in there that one thing it's saying, well, this part, this one it's a fact, whereas this one is one that we're hypothesizing about. And so I I've seen a lot of areas where you know where this is needed. And but you're right. I mean, you know, banking and finance is also a regulated environment that has a lot of trouble. And same with oil and gas, same with supply chain. But life sciences is is is kind of unique in that the the problems are so vast that even the even what you may call the facts in the thing are still only scientific facts, meaning they're only facts insofar as we don't disprove them tomorrow, right? So that makes it extra hard and extra complex. And I mean, if you know, if you look at it, for example, there's what there's one real finance and banking ontology, FIBO. There's something on the last count I made around 1,300 biomedical ontologies. So that tells you something about the space. Biomedicine, healthcare, and such has by far the most semantics, the most invested in semantics and has done the most in this space of any other industry by orders of magnitude. And so, you know, it's a complex space. A lot of money gets bet on these drugs, you know, and and frankly, being able to change, you know, these patterns to be able to get people to use data better and such, I mean, you're really impacting people's lives. I mean, you're getting medicines to market faster. You're you're helping people to understand how to run clinical trials and and and you know, understand side effects better and these kinds of things. So it has real impacts on people's lives. And I think that that piece is also really important.
Eric LittleSo, you know, you've got a complex environment, it's regulated, it's it's you know, the the data it relies on a lot of incomplete science, and you have these impacts that you know are around people's health and well-being, which is is just interesting, you know. So why not start there? I think that's that's a great place to be. And, you know, we we want to be, we also want to be focused, you know, like you don't want to be all things to all people. We're a small company, so we want to be focused and and and solve these problems now. And you know, we'll we'll think about moving on later, but no, for right now, it's it's a no-brainer. We've been in this a long time, all of us, and and we all understand it. And we have a lot of connections, and so it makes a lot of sense to be here.
Allison ProffittYeah.
Context Graphs Versus Real Semantics
Allison ProffittSo you both kind of referred to the timing of of AI and where the industry is and and maybe the scope of the problem in life sciences. But earlier this year, Tom, at the Bio-IT World Knowledge Graph panel, which you've hosted, I think, or chaired maybe for a few years now, you were expressing some skepticism that context graphs is a genuinely new idea. So I'm I'm wondering where is this a timing issue that we're kind of finally ready to have conversations you've been trying to have for a long time, and where is their new semantic thinking that we can do here that we really couldn't do before?
Tom PlastererYeah, so I think for those of us that have been in the in the field for a long time, there's a little bit of a double-edged sword when a term like constant context graph appears. And this is you know part of it is it's it's not a new idea. And so those of us that have seen it look at that and say, oh, well, what do you mean previous graphs were not capturing context? Wasn't that not part of the reason that we wanted to think about focusing on the relationships between classes or concepts? So so I think in to some extent you kind of have that, you know, that immediate knee-jerk reaction. So I should say that anything that brings attention to understanding the context of a decision of understanding, you know, how you want to qualify relationships between things, that's a good thing. Now, the fact that the knowledge graph community has been doing this for a while really just means that this is a new buzzword, a new branding on top of things that we've already had in place. So the way that I think about this is, you know, a relationship between any two things is usually sort of at the simplest level to treat it as a fact. And so what the context graph part of it does is it allows you to think about that relationship a little bit more nuance.
Tom PlastererIs it true for this period of time? Is it true under these conditions? Is it true with this probability? And so those are the sort of relationships that you can represent in graphs pretty easily. I especially see that there's a convergence in technology here between the labeled property graph community and the RDF graph community that makes it even easier to go back and forth between these different representations. So then we don't get stuck in holy wars, we don't get stuck in you know different syntaxes that really you know make it so that you can't move past the particular problem space. So where I was being skeptical wasn't so much around the value of the things that context graphs are supposed to capture. So nuance around that relationship, temporality, for example. It's more around, you know, this is something we've been doing for a while. Let's not think of this as inventing something new. Let's think of this as a way of getting deeper meaning, deeper semantics around the problems you're trying to solve.
Allison ProffittSo, do you do your customers need to know that? Do you are or are they your customers because they understand that?
Tom PlastererThis that's a really interesting question. And I'm gonna let Eric jump on this one too, because he's also done some very recent work with one of our customers in this area. We have to be able to operate at multiple different levels. Yeah, sometimes we have customers that you know are from our fans and friends and family network, if you will, and they already have a similar level of understanding and they'll look at what we're doing, be like, oh great, these guys are gonna come in and be able to take apart our problem, you know, reconstruct it, make it scale, do those things that we know that is is part of their approach. And in other cases, we have customers that are just getting into the graph space, that are just sort of hitting the wall of, you know, what can you do with relational, what can you do with table? And so we need to be able to communicate at both of those levels. And that's not always easy. And you know, me personally, you know, there's like a level that I immediately go to and like, oh, wait, wait, wait, we're gonna have to back down on this a little bit because we're we're not speaking at the right level that makes this idea make sense. Right.
Tom PlastererYou know, if we're talking to scientists, they usually have this idea of there's an intent in my question. There's something that I want to understand. Frequently, there's a lot of data and a lot of analysis that I have to bring together to do that. And the infrastructure piece usually doesn't resonate with them, how you're gonna actually keep track of it as you move toward an answer. So they're passing that off to somebody else. And so that that other group, usually an IT, is gonna have to be the one that's gonna manifest that for them. And then you're gonna come back and show them that answer. They're going to make their judgment on that and hopefully make it a virtuous cycle where they can keep building on top of that. So we need to be able to talk to the SME level and shorten that distance, and that's a big piece of what we try to do. But we also need to make it easy for IT groups to implement, understand, and be comfortable with this approach. So again, we're talking multiple different levels there. Eric Bill, I'll I'll turn this one back to you because you just had to do this for us pretty recently. I think you're in day two with this company.
Layered Ontologies That Actually Scale
Eric LittleYeah, that's so I mean for me, yeah, the the the word context graph is is a is a bit, you know, redundant because you know that's the whole point of these things is to brought is to provide some kind of context to your data, like Tom said. That's that's why we cared so much always about formalisms and relationships and models and things like this. Nowadays, what we're seeing though is is there's a lot of people borrowing a lot of the terms from semantics. I'm seeing a lot of companies that use the word like ontology, we ontology. And then you go and look at what they have, and it's just a data model, right? Or it's just a catalog or something like this. And so recently I kind of I kind of got a buzz on on LinkedIn a little bit because I I made a comment to on somebody's stuff, and I said, you know, there's a lot of people out there cosplaying as ontologists these days, it seems, you know, and so for me, when I see, you know, even the word knowledge graph now is I'm always I'm always forced to walk into any client and and and try to level set and say, so what are we talking about?
Eric LittleLike, what do you guys know about this, right? Where are you? Is this something you recently heard about at some conferences? And so so you've co-opted the words and now now you're you're trying to do this, or is this something you have a long history in doing? Have you built stuff already? Do you guys have any experts with any formal training in any of this stuff? What kind of what kind of graphs have you worked with? Not all graphs are the same. And when it comes to to building these things, I'm very much a layered and leveled and and and fragmented ontologist in terms of I want to build the right size models for people. So long ago in in the in the way back when a lot of people used to have these ideas to build these great big monolithic graphs. And those things didn't work because they don't scale. You know, if you're asking something about proteins, I don't need my query to be traversing all kinds of stuff around, you know, specific side effects of a clinical trial or about a disease, or there's a lot of information in there that I don't need the thing to be scanning the graph about. So you need to make these things fit the right spot in the in the in the layering. So if you think about your data levels, right? So your non semantic data, your relational systems, your data lakes, your CSV file systems, your Excel. Your SharePoints, whatever you've got your data in, right? Whatever kind of data it is, structured, semi-structured, unstructured, whatever. Maybe it's image data, whatever. You want to have a layer of ontologies on top of that that are data source-like ontologies.
Eric LittleSo the first step you do is you say, okay, well, I'm going to make some graph-like structures. I'm going to, I'm going to define some of these entities that are in the data, but I'm going to really be faithful to the data. So I'm going to kind of call these things by the same labels they are in the data. I'm going to put them in a graph structure so that I have nodes and arcs and I can connect them that way. But I really want those models kind of to reflect exactly the data so that I have this one-to-one traceability. But I don't want to just keep the the semantics at that level because you haven't really gone in expressivity much beyond what the data sources themselves could do, right? I mean, relational tables, you can link them via primary and foreign key relationships, right? But it's very rigid and it doesn't say a lot. If you've ever looked at a database structure, entity relationship diagrams only tell you very loose connections and very loose models and mappings of how that is. It doesn't really explain why it's built the way it is. You don't really understand any of the deep stuff. There's not a lot of definitions in there. So when I go from this data source level ontology and I want it to look like the data immediately, I want to bring those models up another layer into domains and subdomains. So this is where I want to have things broken into, like the domain of cells or the domain of diseases, or the domain of patients, or the domain of, you know, hospitals or sites or something like this or clinical trials.
Eric LittleYou can imagine breaking the value chain, right? From research into development, into manufacturing, into production, into regulatory, into commercialization, whatever, right? You can break it into a lot of these different domains, however you want to set them up. And then, of course, because you've fragmented this now, you've you've you've you've taken these models apart. You basically decompose things into these substructures so that they're they're they're you know localized and focused on specific areas you care about. Well, now you have a problem that I've decomposed the world so much. How do I ask a question about patients in a study that have a disease that are reporting a side effect? Because I've got side effects in one model, I've got diseases in a model, I've got patients in a model, I've got sites and hospitals and things in a different model, I've I've broken everything apart. How do I run? How do I put them back together? Well, I have to recompose by by every layer of the ontology that you move up in, you move up in abstraction and you move up basically in like an import.
Eric LittleSo everything falls under this uppermost layer of that's a physical thing, that's that's an information thing, that's a temporal process, right? Like those very, very high-level abstract things that you can break into, okay. Well, it's not only maybe it's not only a physical thing, maybe it's an aggregate of physical things, right? So a stone versus a pile of stones, things like this. So you can you can do that at the very top level, then get into domains, and then from domains into subdomains and subdomains into these these data structures, and then down to the data. Now, when you stitch that all together, I've done two things. I can route my queries and I can route my information through those different graphs, only using the parts of the graph that I really care about or that matter for the the question I'm asking or the reasoning I'm running. But because I built them in layers and and I have this ability to nest them, you know, like like little nesting dolls inside of each other, I can go from the most abstract thing right down to the specific row in a in a piece of data and it's all linked. Right.
Allison ProffittYeah.
Eric LittleSo you you decompose, but then once they're decomposed, you reconnect and you recompose so that it's semantically, ontologically, and metaphysically correct. And and this is this is where we start to apply not only the engineering principles that you'll hear about from you know the RDF and OWL world and those guys, but more what you hear when you talk to some of the formal ontologists out there that that really care about the logics and getting the descriptions of the world correct and such. So there's a there's a balance you can strike in there, and you can make these these highly scalable engineering, you know, you know, tractable in an engineering process, but also very metaphysically complete, very expressive, and they can support a lot of different logics and reasoning and a lot of different kinds of questions and queries and all those kinds of things. So if you if you do this in the right way, and this is kind of what Tom said before when he talked about semantic ops, like semops, is about thinking about your modeling and your metadata, the way that we think about delivering, you know, developing and operationalizing our software.
Eric LittleThink about the end state in mind, right? Think about this thing scaling. Don't build toy things that you get to a certain point, realize it will never scale. You have to rip it apart and redo it to make it, you know, to work at at sort of you know, the big level. We don't want to do that. And I think a lot of people did do that for a while. And so, in this sense, you know, by having these layers and these levels, we can capture context at a lot of different areas, we can capture different perspectives. So the the chemist and the structural biologist can both look at receptor data, but they can look at it in very different ways because you can build a model and then label it like, hey, this is the chemistry view of this, and this is the structural biology view. The chemist cares, did my molecule bind and did it do what I wanted it to do, and did it metabolize the way I wanted it to? And then my job's done. The structural person or the person that cares about pathways, they're sitting there saying, okay, well, yeah, but now you just you just did something to this receptor. Over time, it's gonna upgrade or downgrade. That's gonna send signals through the chain, that's gonna affect other receptors on other cells, they're gonna upgrade and downgrade. I'm gonna get sort of macro level reverberations in the system biologically, based on what you did by by you know injecting this new chemical component in there.
Eric LittleSo, same thing on receptors, same data, but very different views. So, we want to capture and break apart the things in the world so we can describe them, but also, you know, we use fancy words like multi-perspectivalism, meaning we can capture a lot of different perspectives. We talk about granular partitioning, which means you can divide the world up into macro level, really big things, mesoscopic levels like medium things or microscopic levels, very, very small things. Things look different at the microscopic than at the macroscopic level, right? I mean, you know, so if you can have all of these perspectives, all of these different items, you've captured context now in a variety of ways. And so, you know, what we're saying here is if we say context graph, we really are using a mouthful of stuff. There's a lot of concepts associated with the context graph for us. It doesn't just mean, oh, I have some nodes and arcs and I'm calling it a context graph because you know it's it's not a relational table. We're really interested in getting to that deeper stuff.
Allison ProffittAre you enjoying the conversation? We'd love to hear from you. Please subscribe to the podcast and give us a rating. It helps other people find and join the conversation. If you've got speaker or topic ideas, we'd love to hear those too. You can send them in a podcast review.
Modeling Reality Versus Viewpoints
Allison ProffittSo when you create this, when you break down all the work and you you build this, these connections, are you illustrating reality or is the process of creating it defining the reality?
Eric LittleIt's both. So, you know, what you're asking is kind of the classic, pardon me if I get a little philosophical here. It's my training. You're right, you're asking me a little bit the the, you know, and this always comes up when I was at the knowledge graph conference and in the and and one of in the panel I was on, you know, it was kind of funny because we got into this this exact kind of line of thinking and and reasoning. And and I said, look, there's a big difference between ontology, which is the study of existence, and epistemology, which is the study of knowledge. So you care about both, but they don't say the same things, right? So ontologies try, if you're a realist ontologist like I am, and like like many other people in the world are, you care about building faithful models that are really good representations of the world. I mean, think of an ontology. , to quote my my mentor Barry Smith, an ontology should be like your spectacles, it should be like your glasses. It just helps you to focus on the world. It doesn't change the world, it doesn't alter it.
Eric LittleYou're not, you shouldn't, you shouldn't use warped lenses that you know make colors different colors or alter the shapes or sizes of things. You're you're trying instead to focus and be very clear as to what those things are and and accurately and faithfully describe them. However, you always know that you're doing that from some perspective. So you automatically have some kind of bias, cultural bias, or some type of, you know, you're you're at some point, we don't share the exact same view of the world, but there is ultimately, you know, one world that we share. We're trying to get to it faithfully and describe it, but at the same time, you may want to describe some of the perspectives on this, right? So you may want to capture some of those epistemological items as well. So it really comes down to just being clear about describing the things themselves, as well as being able to describe how you're processing, thinking about, and capturing those things themselves. That's why I was saying, you know, looking at receptor data, getting the data or getting the model around receptors themselves is important, but then understanding that there are different people with different tasks, different jobs, and different viewpoints of those receptors, that's also really important to capture that.
Eric LittleBut what you don't want to do is blend those things together and then try to say, well, there's this is where everyone always used to ask the question how do I get the one model or the one semantics that everyone agrees with? Well, you won't, because agreement is epistemological. What you need are you need the right models that that capture the right things we objectively agree on as facts or as close to facts as they can be. But then let's be clear in the context about this model is from a perspective, and that model is from a different perspective. And using those different models allows you to look at the same information in reality, the same piece of of the world, but from a different viewpoint, right? And that can be really useful. It's it's really about clarity and it's really about expressing what it is you're actually doing. And the problem that you see in most data structures is none of that work gets done. We we push data into into whatever format we push it into, we lock it into some storage device, and then we try to query it later on. And then, you know, a lot of the perspective gets baked into the facts, and a lot of the facts are not really facts, they're data type entities, so they're representations of facts, and none of that is spelled out clearly. So you wind up with the the the you know, you can get a lot of errors or you can get a lot of problems using that data. Semantics should just be about clarity, definition, and making sure we're getting those things very right and expressive.
Allison ProffittSo when you get it very right, I'm assuming it's rather large. You've built something that is not necessarily. So tell me
Just Enough Semantics And Competency Questions
Allison Proffittmore.
Tom PlastererThis is this is one of the things that we built to distinguish our approach. Okay. So it's it's very much driven, like I was saying, around how subject matter experts are trying to explore a particular question, a particular hypothesis, a series of hypotheses. And so what we want to start with is really deeply understanding their user story, to some extent, their user journey, and breaking that apart into competency questions that are going to satisfy that. And it's usually just a handful of them, five to ten. And then these competency questions are, you know, a path or two through a graph. They're not that big, but they become the contract, basically, the query contract that will go along with the entire question. So then the goal is to narrowly go through and satisfy that question with just enough semantics so that you build out just enough of that ontological structure to address it, but you build it in such a way that when the next question comes, the next, the next, they all snap together. So this is again modular thinking, problem decomposition, all of those pieces allowing you to not overbuild, but to build it in such a way that you can reliably repeat that same answer.
Tom PlastererAnd in some ways, this is a direct descendant of the fair data idea where you really want to prove and push machine interoperability and then later people interoperability, because you're really trying to set this up so that machines can help you here, you know, including AI, including agents. And so you really don't need to overbuild it. You just need to get a few things right that are following directly from fair principles, reusing vocabularies, especially reusing things like operational metadata. And so this is you know, how are we describing data sets? How are we describing, you know, the provenance of the information that we're pushing through the system? How do we make sure that we have an audit trail that we can go back to? Those sort of things are gonna need no matter what. And those are the sort of things that allow you to really correct AI when it's gonna go off in the wrong direction and your agent's gonna get off in the wrong direction. So I think that's one way that you can make it so that you're not boiling the ocean, you're not having to build this gigantic ontology to solve everything. Yeah. The other piece is, you know, like Eric said, we've got, you know, more than a thousand well-used vocabularies, taxonomies to a lesser extent, ontologies in this space. So you're not building from scratch. You're you're more or less kind of saying, all right, what is the world already built? Is it already fit for what I need to solve this particular problem? Do I have agreement here? And then just adding just enough to solve that problem and then rinse and repeat.
Eric LittleA lot of times what we find too is you go in and just by asking clients what they understand about something, you you start to uncover all of the gaps in their knowledge or the gaps in their data, right? That that they realize that there's a bunch of implicit information in their head that nobody ever wrote down anywhere, or they don't really have certain things defined really well. Or , you know, here's an example of something that that we see a lot is you see, people, even if they get into graphs, they'll overuse really vague connections, like the the one I've loved to rail on is associated with. So you'll you'll see a lot of these models where it will have like Eric is associated with Florida, Eric is associated with his wife, Eric is associated with his dog, you know, Eric is associated with his guitar or something like this. And it's like, no, I'm married to my wife, I own my dog, I'm a resident of Florida and I'm playing my guitar. Let's be clear about what the relationship is, you know.
Allison ProffittRight.
Eric LittleAnd it's not like you have to build some really big fancy model, and it's not about making them huge and verbose. In fact, it's the opposite. We want to break them down and make them as small as possible. We have a we have a concept we use all the time. It's called just enough semantics. So just enough to get the job done around the competency question that you care about. But when you unpack those competency questions, you know, a lot of times you find out there's so much implicit information in there. I mean, recently I can give you an example from another domain. I was working with a bank. And if you look at loan data, you've got a table of people and their attributes. You've got a table of the loan, and then you know, it was taken out on this data as an interest rate. And I've got a repayment table. This is telling me like how often I have to pay it, what my minimum payment is every month, blah, blah, blah. Okay, that's great. There's all my data on loans. What's a loan?
Eric LittleRight. It doesn't say anything in the data, it doesn't say anything anywhere in a bank and what's a loan. So how do they do loans? Well, because everybody who works at a bank automatically knows what a loan is, right? Yeah. But it's not like that's written down in the data that's not in the schema. But you know, a loan is I have a lender, the lender has the initial money, and I have a borrower. The lender gives the money to the borrower. It's a one-direction relationship, and then that direction reverses, and the borrower has to give the money back over a period of time, and they have to give extra money back called interest. And that interest is at a certain rate, and that's what's on that table. And when they have to pay it back is on that repayment table. And who that person is and where they live and how you go hunt them down or call them if they're late, that's on the person table. But there's nothing in there that talks about what the actual structure of loans in general are, right? And and that's that's the kind of stuff that we're talking about is if you get just enough of that semantics in to help to describe in general what we're talking about, suddenly all those data tables and all those fields start to make a lot more sense. And you can start to figure out which ones you specifically need to tackle which problems or answer which questions.
Eric LittleSo the goal should be if you do this right for scalability and for speed and performance, you're actually only using just enough metadata to pull just enough data from the sources and to assemble it just in the right way to answer the kinds of questions you want and move on, or hand it to your agents, or let AI do something with it now or whatever. So, in that sense, yeah, we're not trying to boil the ocean. We're not saying build these big monolithic ontologies. And we and there's a whole nother concept here we haven't really touched on, which is like this kind of top-down, bottom-up approach that people get confused on. So if you go strictly from the data and try to build all your semantics up from the data, you wind up with all your ontologies and semantics. Guess what? Looking just like your data looked like. It doesn't really have a lot of added information in there, right? Like like the loan thing I just talked about. Well, if you go too top-down, I start doing the metaphysics of the world and I start thinking about, you know, all the ways I could describe a disease and all the ways I could describe a patient, and all the ways I could really describe everything. I mean, you could spend your life ontologizing whatever room you're sitting in right now, right? With all the stuff that's in there. So you don't have any data about that, though.
Eric LittleSo why would you bother to build a thousand classes on things that I don't have any instances about? So if you're in the middle and you're thinking, I need enough top-down things to structure these concepts correctly, but I also am going to constrain myself with what I actually have data for. And I want to keep it just small enough to answer those competency questions and move along. That's the sweet spot. And I think there's a lot of people, you know, out there in the world that need some help doing this. And that's kind of, you know, that's why we're doing this. That's why we built this company. And, you know, we're trying to do this as a you know product with some services around it, that basically, again, is that sort of missing piece that you can come in and say, plug this in, tweak it, set it to your data, and it will start to give you value and results, right? And then you can go bit by bit, don't boil the ocean, build it up piece by piece, and and and after a short time, you're able to do a lot in your in your ecosystem.
Serving Scientists IT Regulators And Agents
Allison ProffittSo you're describing different users of this structure. There's researchers who are doing the research, the you know, the chemist that wants to know about bonding affinity. There's also maybe agents that you've built, like the machines want to use this as well to do tasks assigned to them. And then there's also regulators later who want to know how this all came together and want to be able to trace some audit trail back to ensure that you've that it's safe or or whatever. How do you serve, and there may be more, how do you serve those different kinds of users who have different needs?
Eric LittleTom, do you want to start?
Tom PlastererYeah. So I think the the first thing kind of goes back to, you know, what I was describing, really capturing good user stories and really just having that strong sense of where are they going, you know, and keeping track of all of those perspectives that it could be the perspective of your scientists, your strategists, and what they're going to want out of the system when it's done, how it's going to reinforce their research direction. So that's one perspective. And the second perspective I talked about a little bit was your IT professional that needs to support it. And so they're going to have, you know, concerns around the governance of that data. They're going to have concerns about the scalability of that platform, about how it's going to work with existing investments.
Tom PlastererSo that's another audience that you need to satisfy. And then there's potentially partners and other consumers like regulators who are going to want to see the data that comes out of that or the information, the knowledge that comes out of that system and understand where it came from, you know, both from the perspective of what was the original source data set to, you know, what was this built upon, how was it assembled, what software was used to assemble it. So that's the whole, you know, provenance audit trail piece. You can build all those perspectives into your model. And there's a couple of them that we see all the time. So we just start out with those. And there's other ones that are more fit for the particular question in play. So the the approach, the process, and the the product supports all of that. And then it really becomes a matter of you know satisfying those multiple audiences and really having a clear sense of what a good answer is going to look like for those multiple audiences. I think, you know, honestly, that's been part of the challenge is that we tend to just talk to one group or another.
Allison ProffittRight. Yeah.
Tom PlastererAnd yeah, then the yeah, then if it becomes well, it if you're if you're in it, your number one job is to lose anything. And so then it's, you know, give me a platform that I can put it in that I'm never gonna lose it. It's it's not necessarily how am I gonna get value out of this. I mean, the incentives are are are different and the what they're charged with doing is different. So so I think you do need to think about engaging with the multiple audiences at the same time, and that becomes a little bit tricky, but you can you can do it if you break it into this this smaller part with problem decomposition and modularity around your ontology, it's just enough semantics. And then you just have to have strong communication as you go through so that these groups stay aligned.
Eric LittleAnd it's nice because semantics is kind of you know built for that type of stuff, right? I mean, you know, if you think about this, let's take an example. Imagine you're you're the person who is doing medical review, right? Yeah, and so you're you're looking at a study and you're you're trying to go in and you're trying to say, all right, maybe I have a product and I and I want to understand, you know, all the reported side. Effects that were in there on this product, but I also want to look at you know what else does it say in the literature from other people's studies and so on, right? Because I want to be able to understand how to review this literature. Now, what you're seeing is okay, well, to speed that up and to allow me to read a lot more information than human eyes can, I'm going to want to use some kind of a tool and assistance, maybe some AI or something. But what we see with AI is AI has a really hard time replicating things, it has a hard time doing the same task over and over. I if you've ever used any of these, Allison, I'm sure you've seen that, you know, even if you tell it to draw an image, tell it to draw an image, give it the image back and say, don't change anything else in this image, just change this one tiny little piece. Usually it will redraw the whole image and something else gets varied a bit, right?
Allison ProffittYeah.
Eric LittleIt has a hard time just doing the same thing over and over, but but other systems like semantic systems or other kinds of data systems, they don't have that problem at all. It will just run the same. If you make a rules-based system, it just runs the same rule on the same data over and over and over and just gives you a consistent result, right? So you want that determinism, you want that to be built in. The other thing is the mappings. You don't want, like in an in an AI, it's going to always pick some path through the vector database based on the weights it puts on on what's the next step it should take, right? I've got a billion paths, I pick one, then I have a billion more paths and I pick one of those. Semantics, though, you map these things together and the mappings become very consistent. They're very static in a sense, right? You can remap anytime, it's flexible, but the mapping holds. So anytime I say, hey, go back to PubMed and find me this article, it's just like if I say go to a specific website like Google or go to a specific website like LinkedIn, it's not going to just whimsically every fifth time take me to some other website, right? Because maybe it thought I should go somewhere else, right? That's what AI will do. So people want yes, people want that consistency. And semantics are really designed to handle this kind of structure and these mappings.
Eric LittleSo when we want to add this information for these different users and you're doing something like medical review, one of the big you know problems that you see consistently is oh, the things spit out some PubMed articles and you start following links, and the links don't always take you to the article it said it was, right? Sometimes there's errors, sometimes it hallucinates, sometimes it takes you to the wrong paper, sometimes the link is not the paper it said it was, and so on and so on. That that doesn't happen in these other systems. So for us, we're looking at, you know, at designing these entities that these items so that they could be used by human users, they could be used by agents and such, but you want to put that deterministic backbone behind these engines to say, look, AI, I know you have this desire to always rethink the problem every time you're asked, right? You get a prompt, you get an answer, you get a prompt, you get another answer. We want there to be some consistency. I want the same answer every time I give you this exact prompt. And so don't rely on AI for that. Put put a semantic or a knowledge graph kind of backbone in there. And now say, hey, AI, don't worry about it. We're gonna let the, we're gonna let the knowledge graph or the semantic system handle this kind of deterministic connectivity. You do what you're good at, read lots of of of textual data, summarize that data, give me answers and you know, and give me a way to just talk to you so I don't have to write queries and do things like that. You can do those things for me. But but don't worry about maintaining the consistency of the model and the mappings and all of these other things. We're gonna use other tools for that, if if that makes sense to you.
Allison ProffittYeah, absolutely.
The First Step To Start
Allison ProffittIf a listener today has been putting off a knowledge graph investment, what is the single most important step you would tell them to take right away?
Tom PlastererSo I think part of it goes back to the conversation we just had around finding those competency questions, those user stories that are really challenging.
Allison ProffittYeah.
Tom PlastererThere's a couple of places where we found the fit is really good. And so if we're thinking about problems in the business where you have a handoff between information or decision making between different groups, that's really challenging because they have their information stuck in different applications and different data silos. And what they really need is a bridging between the information, the semantics, if you will, that that across and mirrors the business process. Those are perfect. So we we kind of look at sort of cross-functional use cases as a great place to start. I'd say the other thing that's it's really worth pointing out, especially the way that we approach this within our product stack, is that we don't need you to necessarily move your data out of your existing systems. And so you can think about it, you know, where where do I have my points of business friction? And then how can we come together with a set of computational services that will bring this together off of a particular use case or use pattern that will allow me to bridge these sort of you know traditional friction impediments here? So I think those sort of use cases are really, really well designed for a knowledge-centric approach, a knowledge graph approach where you don't even have to start big. You just want to show that value from the very beginning to the very end and then you know rinse and repeat on top of that.
Allison ProffittGreat. Thank you.
Why AI Forces Metadata To Mature
Eric LittleCan I jump in with one one thing I'd like to add on to that? And and maybe if you'll if you'll allow me to be a bit of a provocateur for a second. I think up to now, people have kicked the can down the road for a few decades on on getting their semantics and their metadata straight in a lot of companies. And they've been able to do that because there's always the promise of some technology that'll kind of help them and do it for them. So we saw relational systems move to data warehouses. Well, once I manage my relational systems together and I link them into data warehouses, now I have this big thing of my relational systems and I'm good, right? Because why do I need semantics? I have all my relational systems linked. Well, too brittle, don't handle unstructured data, don't handle image data, a lot of things like that. Okay, never mind. , maybe we need semantics. So the semantics people were saying, you know, no, you still need semantics. Well, hang on, Hadoop showed up and big data showed up, and now I've got NoSQL databases. So now I can put my image and my unstructured and my other stuff in there. So again, maybe I don't need to really worry about my data.
Eric LittleBig data will handle the problem. No, because well, actually, to get the information back out of even those big data tables, you've actually amplified the problem. Now you're relying on the index on that thing, which is still the lookup service. That means you have to know what that stuff is about and you have to make these connections and linkages, you're back to the metadata problem again. And then it was like, okay, so do we really need semantics? Well, hang on, now we have AI. This looks really smart. I can just talk to it. It answers questions, it knows a lot of things. But now that people are working with that, they're realizing, okay, it's not deterministic, it hallucinates, it makes a lot of mistakes, and so on and so on. However, here's where I think things are different. I don't think that people can ignore this metadata problem anymore and getting this straight. And the reason is because they're making the investment now in AI. People have decided we're gonna go with AI. It is this new thing, it's very powerful.
Eric LittleAgain, we're I'm I'm not trying to be overly negative on AI. I'm just trying to be a realist about what it's good at and what it's not good at. And if they're gonna make the investment in AI now, and we're gonna use this AI to you know augment people, replace people, use it for important decision making, write our documents, do all this kind of work, you know, we're gonna deeply integrate these models in, and we're gonna pay for this through tokens, and we're gonna put it in a cloud platform that's also very expensive, and so on. How are you gonna scale and trust this over time? And that's where I think everyone's hitting a wall now, right? We've got this AI stuff, it looks very promising, but I can't get out of POC land with it. Right. And so, so the the answer is you have to put some determinism behind this. You need now to have still this backbone. So, again, the semantics people like Tom and I are sitting here saying, Hey, we're still over here banging this drum saying you still need to describe the things you care about, you still need to describe your domains and entities to some level, you need to be able to get information out of your head that's implicit and make it explicit and connect it up to your data so that your data can be, you know, decomposed, recomposed, and and and be usable in this kind of a framework.
Eric LittleSo I think we're at an inflection point, Allison, where I would argue, you know, go kind of going back a little bit to the other question, what would I say to a C-level executive who's who's thinking about this? I would say you you just can't ignore this problem anymore. Okay. It's it's been a it's been a couple of decades. We've tried all these alternatives to I can fix my inherent problem with this tech or this tech or this tech. And we all come back every time to it's it's not the tech, it's the metaphysics, it's the description, it's the discussion, right? It's it and that comes out in these competency questions and what you know about your domain, what you can say about these things and so on. And if we're gonna, if we're now gonna make this investment in AI, which everybody seems to have decided they're gonna do, then you better make that AI trustworthy. You better make it scalable. And and, you know, I see our role as, you know, when when people ask me kind of like, so what do you do with a philosophy degree and and a neuroscience degree and and you know, and an industrial engineering degree in this space? And my answer is, you know, kind of a cheeky one, but but kind of a serious one. I say, look, I'm trying to take the artificial out of artificial intelligence. I'm trying to make it just intelligence.
Eric LittleAnd to do real intelligence, human intelligence, you need logic systems and you need math systems or statistics systems. AI gives you a good statistics-based system, but it doesn't do logic well. The Apple paper has shown that very, very clearly. Apple wrote a paper called The Illusion of Thinking. And in that paper, they showed that these AI systems will use an immense amount of GPU processing and will ultimately fail at very, very simple tasks after only a little bit of minor ramping up of the complexity of a problem, like the Tower of Hanoi problem, moving the disks from one pole to another, right? You got to move it across these three poles. You can't put a bigger disk on top of a smaller one. Okay, they can only handle like only so many disks and those engines give up and fail. Expert systems that were built 20 some, 30 some years ago can handle thousands of those disks with no problem. Does that make them smarter than AI? No, it just means that they can solve problems, different kinds of problems in different ways, because they're very rule-based and very logic-oriented. So if you can put the logics of semantics and ontologies together with the statistics, you've really got something.
Eric LittleSo if you're gonna invest all this money in the statistic side of the engine, you better invest some money in the deterministic side of the engine that's gonna help ground all of this, or you're gonna be stuck in this problem everyone's in now, where you know, you can't get out of POC land and turn this into something scalable that you can trust and make your $2 billion bet on your drug with. You know, I that's where I'm kind of seeing it. So I think this is critical now that people get something like this. And I don't think with the push to AI that people have the opportunity now to wait, you know, for two or three years to build all of this manually by themselves and figure it out and such. I think they need help. And I think that they need, you know, a a you know, products and things that can come in and provide that 80-20 rule. 80% of it's done, 20% of it, you know, you you build out, make it yours, tweak it, you know, do the configurations and such. But people have to go fast. They've made the decision in AI. Let's make it good now. Let's take the artificial out of it and actually make it something useful. So that would be my answer to the question, you know, of like why now and why is this important?
Allison ProffittYeah, awesome.
Final Thanks And Goodbye
Allison ProffittWell, Tom and Eric, thank you both so much for your time. And thank you for joining us for Bio-IT World's Trends from the Trenches podcast.
Tom PlastererReal pleasure. Thank you, Alli.
Eric LittleYeah, thanks so much.