Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov
Hi, we’re Tim Berglund, Adi Polak, and Viktor Gamov and we’re excited to bring you the Confluent Developer podcast (formerly “Streaming Audio.”) Our hand-crafted weekly episodes feature in-depth interviews with our community of software developers (actual human beings - not AI) talking about some of the most interesting challenges they’ve faced in their careers. We aim to explore the conditions that gave rise to each person’s technical hurdles, as well as how their experiences transformed their understanding and approach to building systems.
Whether you’re a seasoned open source data streaming engineer, or just someone who’s interested in learning more about Apache Kafka®, Apache Flink® and real-time data, we hope you’ll appreciate the stories, the discussion, and our effort to bring you a high-quality show worth your time.
Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov
Apache Kafka and Apache Druid – The Perfect Pair ft. Rachel Pedreschi
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
As the head of global field engineering and community at Imply, Rachel Pedreschi is passionate about engaging both externally with customers and internally with departments all across the board, from sales to engineering. Rachel’s involvement in the open source community focuses primarily on Apache Druid, a real-time, high-performance datastore that provides fast, sub-second analytics and complements another powerful open source project as well: Apache Kafka®. Together, Kafka and Druid provide real-time event streaming and high-performance streaming analytics with powerful visualizations.
EPISODE LINKS
- How To Use Kafka and Druid to Tame Your Router Data
- ETL and Event Streaming Explained ft. Stewart Bryson
- Who is Abraham Wald?
- How Not to Be Wrong: The Power of Mathematical Thinking
- Join the Confluent Community Slack
- Fully managed Apache Kafka as a service! Try free.
SEASON 2
Hosted by Tim Berglund, Adi Polak and Viktor Gamov
Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
Music by Coastal Kites
Artwork by Phil Vo
- 🎧 Subscribe to Confluent Developer wherever you listen to podcasts.
- ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
- 👍 If you enjoyed this, please leave us a rating.
- 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
What's Apache Druid and what does it have to do with Apache Kafka? Rachel Padresky, a longtime analytics expert, technical pre-sales leader, and community advocate, is going to tell us on today's episode of Streaming Audio, a podcast about Kafka, Confluent, and the cloud. Welcome back to another episode of Streaming Audio. I am your host, Tim Berglund, and I'm joined in the virtual studio today by a very good friend of mine, first-time guest on this program, Rachel Padresky. Rachel, welcome to Streaming Audio.
SPEAKER_01Thank you very, very much, Tim. It is an honor to be here.
SPEAKER_00It is great to have you. Why don't you tell us who you are, why you are on the show, and why you would you would say things that are of interest to our audience.
SPEAKER_01I wouldn't that's an interesting way of putting it.
SPEAKER_00I just want to really, I think I want to put you on the spot. You know, like you have to demonstrate there, they're listening now. You know, this is gonna be a good 45 minutes long. Why, why do they want to commit to this? You gotta you gotta tell them.
SPEAKER_01Well, I think they should commit to it because you and I have some really epic banter that at least amuses the two of us. It might be some other human.
SPEAKER_00My priority, yeah.
SPEAKER_01Right. So I think that's why I'm on the show. Whether or not I have anything relevant to say, well, I guess we'll find out at the end of the 45 minutes.
SPEAKER_00I guarantee you do. Where do you work and what is your role there? How's that?
SPEAKER_01Okay, I work for a Confluent partner called Imply Data. We are the company behind the Apache Druid project. And what do I do there? Well, it's a startup, so I do a lot of things, but officially my job is to run the worldwide field engineering organization, which it other names is also sales engineering, solutions engineering, uh pre-sales consulting. We've got so many names.
SPEAKER_00Yeah, and um you are to paraphrase uh Sheriff Behan from Tombstone, uh a woman of many parts, but I I think sales engineering is one of the like the defining things of your career. And everybody that sales engineering is what pre-sales technical people call themselves. Um you know, they're they're the the partner to the salesperson proper who actually helps them helps make the technology work and answer technology questions. We've we've talked to SEs, sometimes we call them system engine systems engineers. There have been SEs uh on this program before. I want to say most recently Danny Traphagen. Um that's an S that I might know.
SPEAKER_01I might know that.
SPEAKER_00That is actually happens to be a mutual acquaintance of the two of us. Um so anyway, yeah, you lead the organization of the technical people that partner with the salespeople to help people when they're help customers when they're trying to buy stuff uh from uh your company, right?
SPEAKER_01That is very true. Um, I think uh SC job is the best job in software because we get all the best part portions of a of a customer. Like we get them when they're happy and they're excited about the technology. And that's our job is to get them excited about it and help them figure out if it makes sense in their infrastructure. And we don't get to be we don't have to be like that salesperson who you know asks them if they need to get into newer used software today. Like it's a very cool job to have. And I've been doing it for the most part of 20 years. Yes.
SPEAKER_00Some number.
SPEAKER_01Some number of years.
SPEAKER_00We can probably estimate by by judicious Googling, but we won't. Um yeah, it's some some amount of time, and you're pretty good at it at this point. And yeah, it is an interesting point about kind of being on the front end of you know, timeline-wise, of the relationship where you're figuring out what can be and what can be done, as opposed to say uh post-sales support or or consulting or delivery of those things when you're more finding where bodies are buried and and where things are difficult and all that sort of business.
SPEAKER_01Yeah, I call that working for a living. And yeah, pre-sales is not working for a living.
SPEAKER_00I think it's a very hard job. You have to have a uh unique skill set of of uh really having some capability as a salesperson and all of the you know, sort of political savvy and communication skill that comes with that. Um, but you have to be deeply technical. So um you remind me of my years ago in my independent consulting uh that that period of my professional life. I was always on that same side of things. The hey, let me help you get this planned out, let's uh sketch out some architecture, spike a prototype, get your team trained, and peace out, I'm gone. And then you know everybody else writes the code and figures out where the uh difficult things are. So there you go. Now there's more to what you do. I want to get into that later. There's a lot of we the the kind of community aspects of what you do, I think, which is where you've also been a leader in your career, and I want to talk about that. But you mentioned some technologies and this is a show about Kafka, Kafka, Confluent, and the cloud, which I said went during the opener.
SPEAKER_01Is the cloud a thing? I I I wasn't aware.
SPEAKER_00100%. Oh wow. I think uh actually, if you like some free industry analysis for me, everybody, I think the cloud is going to be big. I heard it here first. So the cloud's a thing. Kafka's a thing, definitely. Uh Confluent Cloud's a thing. We talk about the confluent platform, all these all these kind of streaming technologies. Um Druid is properly speaking not one of those. Um and you're you said that uh you're on the show because you're a confluent partner. You're you're not just on the show because you're a confluent partner. This is not uh everybody behind the scenes. This is not our mutual business development teams pulling the strings and and you know, trading favors to get Rachel on the show. I actually think it's an interesting topic, and it's a topic that you kind of go deep in from a technology standpoint, but I'd like you to tell us what Druid is. I kind of want to just talk about that for a little while.
SPEAKER_01Oh, what druid is. Well, let's start with Druid is a database. And that means a lot of things to a lot of people. So it's a very kind of almost controversial thing to say, because we can say, what is a database? You know, that but at the end of the day, Druid is a place where you put data and then you take data out of it to do something with it.
SPEAKER_00Is that, by the way, um I actually I have a talk that I've been giving for the last year or so that I start by asking the question, what is a database? Um it's kind of it's kind of like the you know, the the undergraduate philosophy, what is a chair question, and you're like, Well, I know what a chair is, and then you actually try to answer the question and you realize you're you don't know anything. So is that your definition of a data database, that thing you just said? Place where you put data and then get it back later?
SPEAKER_01Where you store data for a period of time and then you take some of that data out at some point to do something with it, but it pretty much stays there. Yeah, I don't know if I have a personal definition of a database. I mean, I that might be a flaw in my brand. I'm I might need to have done that in my life to define it.
SPEAKER_00I'm I'm just here to help you grow.
SPEAKER_01Thank you. Thank you. I appreciate that.
SPEAKER_00So I I interrupted you. Druid is a database.
SPEAKER_01Yeah, and it's a so we'll go back, you know, in time a little bit. And I started my career with a also a database company called Red Brick. And Red Brick was a database that was specifically designed for data warehousing. And it was started by a gentleman named Ralph Kimball, who, if you are a data warehousing person, you might have heard of.
SPEAKER_00You might have heard of that name.
SPEAKER_01Yes. And um I remember my interview for that job. So I was fresh out of college, didn't really know anything about databases, and I was looking in the um in the appendix of a database book on what an index was in order to answer a question, that question during my interview. So I really didn't know anything about databases. My mathematics degree from University of California, Santa Cruz, did not prepare me for what is a future index question. No.
SPEAKER_00It just didn't.
SPEAKER_01No. Um, but I started my career with thinking about databases in the terms of analytics and not in the terms of what I learned later was third normal form or OLTP systems. Um and Red Brick was a was really cool. Um it ended up being purchased by another database company called Informix. And they Informix eventually got purchased by a company you might have heard of called IBM. Yeah. So um, and it's all it's interesting. We we look forward to what Drew it is, and it's almost like it's the next generation, or the third generation, or ninth generation, or something. It is the um, it is this the great, great grandchild of technologies like Red Brick. And um so it is a it's a database that's for analytics versus a database that is primary for transactions. Except that those worlds have sort of merged together. And um, if you saw my Kafka Summit slides, and I'll put in a plug for my Kafka Summit San Francisco 2019 slides, I have this this picture of there's OLAP technologies or online analytical processing technologies, and there's OLTP technologies, and they've all sort of merged now into what I just call PE processing. Um because that whole line between what's OLTP and what's OLAP is really blurred.
SPEAKER_00Um could you, real quick, for uh the new folks in the back row define those two things, OLTP and OLAP.
SPEAKER_01Sure. OLTP, online transaction processing. It is if you think of the grocery store analogy, when you go through a grocery store in the United States, you um you have your your things and they get scanned. And every time it gets scanned, it is updating some database somewhere with inventory information. It's a very quick transaction. Oh, Rachel bought chicken nuggets, right? Okay, so let's let's update that. Actually, not Rachel brought it, chicken nuggets, more importantly. Chicken nuggets were uh were purchased. Let's take one of those off of our inventory. And then it also creates like a register of the things that I bought as well, so I can get a total at the end. And I consider that OLTP. Those are very short little transactions that keep track of what's going on.
SPEAKER_00And they reflect things that happened in the world. Uh they they correspond to events in the world.
SPEAKER_01Yeah, and there's, you know, we can we could even talk about acid here, and we could talk about all these interesting technology technological ins and outs, but for the most part, there's short request. And I think that is a term that I sold from DBMS2, um, short request. And OLAP is the idea that in taking, you gather all that information and you can ask it questions. So now you can say, all right, what are mothers of two young children buying at grocery stores on Friday nights? Ah, chicken nuggets. I'm sure that was an analysis somebody did at some point. Um and uh it's about more long request processing. So asking questions over larger pieces of data to get insights. And I think if if you if you are old like some of us are, and remember when they started offering loyalty cards at grocery stores. Yeah. Yeah, I think that's when people try to figure out oh yeah, we've got some good data here. We should like do something with it. And let's try to track who the person is with what they're buying so we can market to them better and we can market to the group better. And that's um, at least in my grocery store analogy, how I define the difference between OLTP and OLAP.
SPEAKER_00Very nice. Because those are uh probably everybody listening, well, I'm gonna I'm gonna go with everyone listening to this knows what a database is. Uh everyone listening to this is probably broadly familiar with the idea of analytics, but there has been historically a sharp distinction between those two kinds of databases, and that's for you know highly implementation-specific technology reasons. They're very different kinds of uh ways of accessing data. Uh you know, one prioritizes writing and a certain kind of retrieval that is of small amounts of data and has to move fast, and the other one uh you know, analytics prioritizes kind of looking at all the things. Uh you don't look at all the things when you're doing OLTP, when you're doing transactions. When you're doing analytics, you you kind of do. This big bucket of stuff and you look at all of them. And uh it's it's you know, for for good, uh reasonable uh implementation reasons, it's just hard to build one thing that does them both. And so that's those terms. And and it's funny that you say that they're blending, because it just seems like I don't hear those words, those those acronyms as much as I used to. And I don't know if they're falling out of favor or what's happening. They're still there's still totally valid categories, and I wanted to make sure everybody knew I think that's part of where you work.
SPEAKER_01I mean, I would say that Kafka is one of the technologies that is blurring the lines between the two. And and there's a chicken and egg analogy here that we can say, right? Is it because the industry is moving towards no longer that distinction that Kafka emerged from the brains of smart people? Or is it the other way around? Is because technologies like Kafka have been enabling democratization of this ETL process and the democratization and um the ability to do polygot polyglot persistence in a meaningful and real way. That, you know, there's there's a question for you.
SPEAKER_00Yeah, it's kind of the uh the the great man versus the zeitgeist uh theory of of software or great great project. You know, is it is it sort of uh a trend that's happening, and if there's a very influential technology that emerges, it's a part of that trend, or is the technology emerges and causes the trend? And that's a topic for another podcast. But I see where you're going with that. I definitely uh Kafka broadly is certainly trying to be a part of that that that blending.
SPEAKER_01Well, and that's exactly the reason why Druid exists. Because the the the type of applications that a Kafka allows for, this kind of idea of data that is streaming from wherever all the time, the the applications that you build on it are no longer just um set to be those transaction processing, but you now can do more of these long requests, these you know, looking at all the data type of um applications in real time and look at historical data. And that didn't used to be available because ETL made things long. Um I met a C-level executive at Disney last year, and I was trying to explain what we did, and I mentioned ETL at some point, and and he's like, Oh, oh, I know what ETL is. Like his eyes got all bright. He goes, Those are the things that make my report slow. I'm like well, yeah. I mean, it it at one point Gartner was saying that ETL took 18 to 24 months to create an enterprise data warehouse. And when I built one for Vail Resorts, it took me like 10 months just to do it for a very small data set. So it that that world of batch and the world of hey, or even micro batch for that, is no longer relevant. People want to have data-driven applications at their fingertips and they and they want to have conversations with the data, but they need to see the data as it's happening, not what it happened five minutes ago. I mean, Druid came from the world of digital marketing, where presenting the most appropriate advertisement is going to get you more money at the end of the day. So it's like the ultimate of taking OLTP and OLAP and smashing them together and say, here, go figure that out. And that's where Druid came from.
SPEAKER_00And by the way, we'll link in the show show notes to a recent interview with Stuart Bryson, who talked what all that stuff Rachel just said. Uh he and I deep dived, deep dove, dove deeply. Uh that question in particular, like what is driving the move from traditional ETL to streaming? Uh, you know, we talked about, took apart the idea of what is traditional ETL and how does streaming differ. And some of that he actually had some surprising uh conclusions that had come from his practice. So check that episode out if you haven't listened already. That's a good way to get more into that. But it sounds like uh Druid is a database, Druid is an analytics database, Druid is uh an analytics database that expects to be answering questions in real time. I'm gonna use that highly inflammatory word and let you back off of it. Am I putting words in your mouth? Is this good?
SPEAKER_01No, that's I'm I'm glad that I all my words got so succinctly summarized.
SPEAKER_00It is my one job on this program.
SPEAKER_01Excellent. Thank you.
SPEAKER_00So um tell us a little bit more. It's it's a real-time analytics database. Uh what makes an analytics database an analytics database? Tell us about the tech a little bit.
SPEAKER_01Yeah, it's um actually extremely cool. Uh and it takes ideas that are not like that far out. They're actually things that other technologies have done over the years. They've just not quite put them together this way and you know, added a little sprinkling of magic druid dust on top of it to make it make it even more like unicorn magical. But um, for the most part, we're we're looking at ideas like uh real-time ingest. So uh search systems are really good at that, uh, enabling technologies like Kafka and other streaming technologies. Those are provide this real-time ingest. So that was a number one um requirement uh when they were building out Druid. Um so data needed to be fresh. Uh it also needed to be schema flexible. Uh this was uh data changes slowly. Uh I used to do talks on slowly changing dimensions, you know, type one, type two, type three.
SPEAKER_00Uh good traditional data warehouse stuff.
SPEAKER_01Yeah, and but data changes and the types of dimensions that you're going to want to query on isn't not always exactly what you decided at the beginning. So this idea that the scheme on right that traditional data warehouses implemented needed to be loosened up a little bit. But maybe not as loose as it happens in a data lake with Hadoops, um, because then you end up with, I don't know, maybe a mess, maybe a swamp, maybe not really knowing what data belongs where. So uh Drew design has a concept of a schema flexible design. Um it also uh aggregates data upon ingest, kind of as you want it to be aggregated. So there's a lot of data that's out there at the millisecond level, and honestly, most humans don't really care about data at the millisecond level. We might care about it at the second level or the minute level. Um so you can set Druid to ingest and aggregate data appropriately at that time. Uh also it stores it in a columnar uh sense. So uh you can add new columns. It's super easy, it's no longer the days of alter table, add column, wait 15 years, your database goes down, people yell at you. Uh just add a new column. And those columns are by default indexed and compressed. Um, so we typically say that your raw data to data in Druid is about um 10 to 1. So you'll get about um 10 to 1 compression within Druid. Yeah. So it means you can say, you know, save more data for longer. Um it's also a post-cloud database, because we already established earlier today that you know the cloud is a thing. Uh Druid was developed after the cloud became a thing, which was not just a few minutes ago, but probably a while ago. And um has been designed to minimize the um your your compute costs so you can move different types of processing threads to different types of machines to you know get a better cost performance ratio there.
SPEAKER_00Got it, got it. Okay. So just designed to be cost efficient with compute resources. In the cloud.
SPEAKER_01Totally. And then also tear off data too. So you can put you know five-year-old data on maybe sort of S3 style storage, while you might put last month's data more in RAM.
SPEAKER_00So that tier data add ingest is a tiered storage is baked in.
SPEAKER_01Yep. Excellent.
SPEAKER_00Excellent. That is a uh I appreciate that you see post cloud and not cloud native. Um I mean, I'm not above saying cloud native. It does mean something. It is uh it is a meaningful term with uh specific semantic content associated with it, but um you do kind of have to back off and say what you mean uh because it it does get abused. And uh post cloud, I think, is is more specific and I think drives your point there. Uh you said columner in storage. For anyone who does not know what a columnar database is, what does that even mean?
SPEAKER_01The data is stored in columns versus rows. So at the end of the day, everything's always just gonna be files. It's just real talk, you know? It's all files. Excuse me. Um, the file, the each column is stored as a set of files. And therefore, when you query your data, you are only going to be opening up the files of the columns that you need, and therefore more efficient, blah, blah, blah, blah, blah. So um columnar databases are not new. Um, they're they've been around for a while. Uh, I worked for a company called Vertica that is a columner database, and there's others out there. So that was one of the OLAppy features that the authors of Druid uh picked up was columnar databases make for much faster table scans, to put in lack of a better term.
SPEAKER_00Let me um let me throw a scenario at you, uh, which I believe to be true. Again, if if this idea of a columnar database is a new thing to you, this is for your benefit. But let me try to summarize this and you tell me if I'm right. So the typical uh OLTP query is to on write, uh, create a new row or modify one row somewhere, usually, right? So you have a file or a set of files, and you have some index that tells you where the where's the offset in the file of the block of data that has that row, and you go read that offset and you change the thing and you write it back, or you go put a new one in a free block and you write it in there. Something like that, right? The OLTP is all about inserting a row, changing a row. Um, or when you read, you just read probably a row. Um analytics generally, uh, like I said before, you're reading all the things. You wanna you wanna read all the things and do some sort of reducing operation to those things, like say add them or average them or find the biggest. Or of course there are more sophisticated things, but let's just stick with those. Um and maybe I I want to go through some time range of rows. Uh and it's a lot less efficient to have to do that on a row by row basis with an OLTP style storage format, because you got all this I.O. you have to do, where if you're really just aggregating one column or a small number of columns, this file with columns in it, and now I can index into that and seek to the place in the file where my time range or whatever starts, and I can now do this big sequential read of the day or the year or whatever it is, and it's all there in this file sequentially, and I do my aggregation. Is that right?
SPEAKER_01Sounds good to me. Um I always love hearing you describe these things because you do them so well, and I go, oh, wow, I need to steal all of that for my box.
SPEAKER_00By definition being recorded, so we can get you a transcript.
unknownUm excellent.
SPEAKER_00Yeah, so uh and that's another again, if you're just if databases are not your world, um then you know you can hear things like OLTP, OLAP, Columner, and they just mean nothing. Or if you're if you're relatively new uh to the the business, this just might be things you don't know yet. So Druid is one of those. Um, and obviously there's lots and lots of ways to be a columnar database, and it is kind of optimized for this life in the cloud. Um so do you I I have another question. I I feel like I'm I'm moving you on and you had something else to say. Are you?
SPEAKER_01No, I am much more happy to answer questions.
SPEAKER_00All right, good, good. Just making sure. Um what does this have to do with Kafka? So Druid's cool, and I kind of wanted people to just be able to learn a little bit about it. It's here's another Apache open source project. It's kind of a uh brother or sister to to Kafka.
SPEAKER_01Um but uh why do why why are these things discovered in the wild together because Druid is really, really, really good at taking data from Kafka and ingesting it. And in my world of pre-sales and helping customers get up and running with Druid and with the Imply Stack, I so prefer it if they're using Kafka because boy, it makes my life easy. Like the integrations just works extremely well. So we can get a, you know, you can get queries coming off of your Kafka data, you know, in a matter of minutes. Um, and that's a real game changer from the world that I used to work in, where when you had a new data set, it may take weeks, or if not months, in order to be able to get to the point where you can query it. Um so one of the reasons that you see these things in the wild is that uh druid works really well as Kafka as an ingest, um, to the point where there is a mutual, I don't know, customer or user of both Kafka and Druid. And um, they're a pretty you know prominent tech company. And they actually told me they brought in Kafka because they were using Druid, uh, because they wanted to be able to stream gate in real time, and Druid's integration with Kafka was so good that they started using Kafka, and then it just went viral within their organization. Nice as it does.
SPEAKER_00Very nice. And we'd say who that customer is, but two things. Number one, you know if they're referenceable for you, because you're in pre-sales and you pre-sales people always know that. Um I am in developer relations. I never know, Rachel, off the top of my head, whether a given customer is referenceable or not. We didn't, folks, we didn't coordinate this before the show, so we just can't say the name. Maybe if we figure out later we can. We'll put it in the show notes, and you can know, hey, here's this customer uh who's the super compelling user of both. Um you can look over the show.
SPEAKER_01Well, they're not they're not actually, I don't think they're a customer of ours.
SPEAKER_00They're actually open source users.
SPEAKER_01Yeah.
SPEAKER_00Got it. Okay. Um, so uh and I want to come back to that, actually that in a minute, but uh that makes sense. It is uh real time, and I'll say I always feel like a little piece of my piece of me dies when I say real time in this sense. But this real time analytics, cloud, post-cloud database. And by real time, we mean you're you're not gonna be waiting for long, right? The idea is you ingest data, it is rapidly indexed, and will be available for queries uh sooner than the technology it's replacing is is kind of what real time tends to mean. It it it has a uh it is an imprecise term, but it gets the idea across. And that's druid. And so people who are dealing with invented data, um, well, you know, they may make a commitment to Druid and say, hey, we need Kafka for ingest, or people who are using Kafka for integration uh, you know, between microservices or ingest of events or just kind of building event-driven systems in general, if they've embraced that architectural paradigm. Uh you want to answer questions about your data. And uh obviously Kafka and Confluent, you know, we have a story. We've got Kafka streams and uh Confluent has KSQL, and these are all these real-time analytics things, or at least tools with which you can build uh software that does uh event evented uh analysis, you know, real-time, real-time analytics computations. Um but uh those are uh we'll say highly purpose-built analytics queries. Uh a KSQL job that's running is something that has been operationalized, right? You just you know what that query is and you just want that query running all the time. If you want to have a bunch of data in Kafka topics and then go do exploration or exploratory analytics on it, um that's not gonna work real well. Um that's just not uh not a real Kafka use case just yet. And so when you have this Kafka Druid combination, um that vastly broadens the horizon of the kinds of analytics work you can do. You can do, it seems like in Druid, you could do um I mean people do operational stuff, right? People build a dashboard that somebody is expecting to use and they put that in front of Druid, right?
SPEAKER_01Yeah, totally. And unloading capabilities, et cetera.
SPEAKER_00Totally, totally. Um and then if you want because the the the the capability I'm always talking about how to enable is the exploratory data science thing. You know, how do I how do I give something to someone to play with and you know, business stakeholders ask crazy questions about, hey, you know, is it more likely uh if for this person you know accounts uh the accounts convert uh at a higher uh probability in Q3? If I know the internal champion has a turtle for a pet or something, you know, whatever, that'll you know, some business analyst will come to a data science team with a question like that, and you probably didn't have a KSQ query that was answering that. You know, you need data sitting somewhere to and to be able to ask it.
SPEAKER_01Well, there's very few universal truths in this world, and I do think that one of the universal truths I have is that no matter what, how no matter how much talking to your users you do about what reports they need, they're always going to lie to you. They're always gonna want something more. So as soon as you give them something and you say, Yes, this is everything you wanted, they go, Yes, this is great. And three minutes later, they're like, but what about this? Um, and that is exactly the type of use cases that you're gonna want a database like Druid for. You're just like, here, have fun. Go, go have fun with the data. I'm just gonna step out of this situation now. Exactly.
SPEAKER_00That makes a lot of sense. Now, um, turning the page, uh, I want to talk about something else. You and I have worked together in the past. Um, and I I think we said that before, right? That we're well, I don't know if we did. Rachel and I have been coworkers in the past at the same company. It's been delightful. Uh I will say once and future co-workers. Um and uh we have worked together on community efforts. Of course, that's what I do. I run developer relations here at uh Confluent. Uh you're in sales engineering, but you uh you've got some thoughts about community that honestly have been helpful to me as I form my sort of uh developer relations worldview and and ideas about strategy. Uh and you've got some thoughts on how it integrates with sales engineering, since that's you know, I think the the better part of your heart is is being an SE leader, but there's also this this sort of inner community leader Rachel that seems like she's struggling to break out sometimes. So just tell us your thoughts.
SPEAKER_01I mean, honestly, the two roles are really, really similar now. Yeah, especially in an open source company. Yeah. Um it's something I spend a lot of time thinking about because there's a I I've spent my career working for both companies that are open core. So they have an open source project that they provide support services and other uh software around. And I've also worked for completely closed source um companies. And it's and in pre-sales, you get night, your your whole life is pretty much helping people through a POC or trying to avoid POCs where possible, because POCs are work, and as I mentioned earlier, you know, SEs, you know, we we leave the hard stuff like down the line, right? So these proof of concepts is what I mean by a POC. And when you work for an open source company, or when you do um when you have um an open core product, a lot of the POCs are done without you knowing about it. Um and can I tell a little story, Tim?
SPEAKER_00Yeah, you can.
SPEAKER_01Has to do with airplanes and you know mathematicians.
SPEAKER_00Rachel, will you tell will you tell me a story?
SPEAKER_01Okay, great. Yeah, so uh there is this great book uh that I read um a few years back about um how to make decisions um using mathematical logic. And I've you know it's in the show notes. I added that to the document earlier. And in the prologue of all places, they tell this great story about an Austrian-born mathematician named Abraham Wald, um, who helped the US military during World War II. And one of the questions that the military had asked um Mr. Wald was hey, we're we're getting all these planes back from the European theater, and there are bullet holes all over the fuselage. And we're always wondering, you know, uh armoring the planes is expensive and it makes it heavy. So can you look at all where all the bullet holes are and figure out where we put where we should put the armor? And in in the wise mind of Austrian-born mathematicians, he said, well, probably we should look at the planes that are at the bottom of the Atlantic Ocean rather than the ones that came back. And what they found was that the holes that weren't were in the engines because those planes didn't make it back. And I think about open source proof of concepts like that. And I even to the point where internally I always talk about the planes at the bottom of the Atlantic Ocean. Where are the POCs that we're losing that we don't know about? Where do people struggle with Druid in open source and then abandon it? Or where does Druid not show up as a possible solution to a problem when it could be because the right blog posts aren't out there or the right.
SPEAKER_00So instead of planes come back with holes in them and we add armor there to the holes and and you know, actually planes perform worse with that armor, you know, you got to look for the the planes where did they get hit that prevented them from coming back. And so the analogy is all of your you know, closed lost analysis on POCs is uh not exactly what you don't want, but one of the things you don't want, that's that's potentially deceptive.
SPEAKER_01Exactly. And when you're looking at your incoming prospects, and if they have had a successful experience in the community with their proof of concept, the SD has a lot less to do, right? They're gonna have to build the value around the support the services, and in my case, that Imply provides, versus spending time getting Druid up and running. Uh so it lowers the total cost of sale. So investing in community is actually lowers the total cost of sale for a software company. And so this is where I've been looking at for the last, oh God, five, six, seven years now, because it it's so important that you have a strong, vibrant community that are bringing you customers that have been successful. And I like to talk, and so we we hear a lot about you know community as the the top of the funnel of the marketing funnel. And yeah, it may be if you're looking at it from the perspective of a software company. Um, but if you're looking at its perspective kind of at the at the higher level, community is really your flywheel. The the more energy you have in your community to move that wheel, the more people are going to move it for you. And the more people are going to be successful with your project, and that will cause more people to be successful with your project. And therefore, you're it's going to trickle down into the experience that they have with the company that's behind the project as well. So, my my idea around community is that it has to delight the user from every moment because they're going to be community is what the user is going to see throughout their entire life with not just the project, but also with your company, hopefully.
SPEAKER_00I like it. So, how how can community delight those users? What what needs to be present in your view?
SPEAKER_01Oh, and that's a big question, right?
SPEAKER_00There's a whole other I ask you at 40 minutes into the podcast, you know.
SPEAKER_01Right. Well, that I I think a lot of that is not there's no playbook for it yet. It's almost like when you do demand gen marketing for software, there's a fairly defined playbook. Like you do X, Y, Z, you're going to get these types of results, you know, money in, money out. And I think in community, um, we don't have a playbook because every community is going to be different. And you're going to, and every community is in a different um stage of their flywheel development. It's almost like you have to continuously adapt your community program in order to reflect the way your community looks at that particular moment. You know, I'm a big proponent of um doing surveys. And so if you do see a survey from the Druid community coming out, it's 100% because we want to make the community better for the users and attract more users. Uh so I think asking a lot of questions and trying to figure out, at least for me, who is the ideal community person? Like who is going to be the user of Druid at the end of the day? Who's going to be making those decisions and who's going to be influential in there? And they're not easy questions to answer.
SPEAKER_00No, no, indeed. And uh you said flywheel a couple of times, and I would like to ask you to expand on it because this is uh just one of the most insightful things about community that I've ever heard you say. Uh, and for the benefit of the listening audience, I have heard Rachel say this before, very influential concept. What do you mean by flywheel and flywheel in contrast to what? Talk us through that.
SPEAKER_01Yeah, um, so flywheel and to me, in contrast to that top of the funnel that I mentioned earlier. And so a flywheel is overcomes inertia. And I think there's a tendency for communities as anything to fall into inertia if they're not um you know, if if they're not pushed in some way. And a community as a flywheel, you know, helps a commun uh helps a group overcome inertia. And then the more power that you put behind it, the more effective that community gets. And that's and that's what I'm saying is if we think about our community organizations in terms of the power that fuels the organization versus a place that we get leads, we're going to we're gonna end up creating a more delightful experience throughout the customer's entire journey with the project and with the company.
SPEAKER_00There you go. And where I have taken this, by the way, which again, just amazingly insightful analogy, where I've taken it is to uh measurement. Um with a funnel, uh you know, I I have not spent a career as a marketing professional. I've just really in the last few years been close enough to um, you know, some really, I would say highly scientific marketing operations to appreciate just how dang quantitative it is and how m mechanical it is. Uh most of my fellow engineers, when they think of marketing, they think of, you know, this big party with a two-drink minimum and a bunch of people who don't work, or maybe they make big logos or something like that. You know, it's it's they think of I'll say with logos are important.
SPEAKER_01I like logos.
SPEAKER_00They they are, but uh with respect to my colleagues in corporate marketing, they they think of that two drink minimums are good too. They think of like the the worst stereotype of corporate marketing is is just you know, fluff and pump and blah, blah, blah. You know, but marketing is this incredible mechanistic thing where You are you are identifying people and you are uh communicating with them and measuring when they go away or what it takes to optimize the probability with which they will convert to being somebody who pays you money. And it's all like science. It's amazing the machines these people build. So I think there's a lot of my fellow engineers don't appreciate how much uh quant what a quantitative game that is.
SPEAKER_01And dude, marketing is hard.
SPEAKER_00Right. And if you do that with if you if you take that game over to community, okay, we've got people coming to meetups, let's sign them up, let's get them on a drip campaign, send them an email uh with some helpful content and uh you know click here for to download the white paper and all that kind of stuff, uh they'll never come to your meetup again and they'll tweet about you, right? Uh that doesn't work. So you can't apply that funnel discipline to community because I think we, and by when I say we, I mean technical people, expect when we access community resources, uh like a meetup, like a conference talk, like a uh tutorial video or a blog post, you don't want to sign up for that. You don't, you don't want a salesperson to call you. You don't have budget anyway. Like, don't bother me. I'm just trying to learn here. And so the question becomes like everybody knows how to measure a funnel. And that's just that's this whole discipline, and that's uh, you know, that that side of marketing and growth, and it's a it's a science and it's cool. But you go over to community and Rachel's funnel analogy, or pardon me, uh flywheel analogy, you don't you don't collect names, you don't uh measure when they drop off. What you do is you measure how much energy is in the flywheel, right? How fast is it spinning, how much mass is in it, and how fast is it going? Uh and and more specifically, what that means is you can count activities in the community. How many uniques does the blog post get? Does the video tutorial get? How many people are at your meetups, how many are RSVPing, how many meetups do you have, uh, et cetera, et cetera, you know, whatever it is that you're doing to nurture your particular community, uh measuring the energy in the flywheel is sufficient. Um because, well, it's not just sufficient, it's that you can't measure it like a funnel because it'll go away. If you if you if you think of it like a flywheel and measure it like a flywheel, then you're still able to nurture it if you have some skilled community leadership behind the behind the rudder there.
SPEAKER_01Word. That was awesome.
SPEAKER_00Okay, well that's uh and I can only say those words because this is like I don't know, two years ago, three years ago, we had some conversation on the phone. I don't remember exactly what the context was, but you're like, think of it like a flywheel, and there was the mind-blown gif sort of appearing over my head, and it was it was just very helpful. So that's uh Rachel Pedresky. Words of Tim Berglund, ideas of Rachel Padresky. It's a funnel.
SPEAKER_01Sometimes they're better than the cloud has become a thing.
unknownRight.
SPEAKER_00Not untrue. Um what uh uh any other thoughts on community? I kind of I kind of took over that for you there for a minute, but I just get excited about this idea.
SPEAKER_01Oh, I have so many thoughts on community, but I uh you know probably need to have it have save it for a conversation for another time. Um but uh if I leave anything, it's that the you know treat your sales engineers as evangelists because they're gonna be the best ones out there. And there is no better uh outcome from a first call than to have the person that you're going to be meeting with say, Oh, that's right, I saw that person's video on that particular topic. Um, talk about creating technical credibility right away. So there's lots and lots of good reasons that sales engineering and community should be friends, should uh we should rethink maybe how we talk about those two organizations.
SPEAKER_00My guest today has been Rachel Padreski. Rachel, thanks for being a part of Streaming Audio.
SPEAKER_01And thank you, Tim. I hope I enjoyed my time, and let's go out there and community some.
SPEAKER_00And there you have it. Before I go, I want to tell you that we have a pretty cool new offer to help you get started with Confluent Cloud without you having to pay for anything. If you're a new user and you go through the regular signup process and start using Confluent Cloud, your first $50 of usage per month are free. This will last for the first three months after you sign up. So that's $50 per month of serverless Kafka for three months at no cost to you. So go to the sign-up link in the show notes. I don't want to read you the URL, and sign up now. I think the only thing I could really do more is write your code for you. And I think we can both agree that's too much to ask. So check it out and hey, let us know how you like it. Anyway, as always, I hope this podcast was helpful to you. If you want to discuss it or ask a question, you can reach out to us on Twitter at Confluent Inc. or reach out to me at TL Burgland. That's T-L-B-E-R-G-L-U-N-D. Or you can hit us up in Community Slack. There's a sign-up link for that in the show notes as well. And while you're at it, please subscribe to our YouTube channel and to this podcast wherever fine podcasts are sold. And if you subscribe through iTunes, be sure to leave us a review there. That helps other people discover the podcast, which is a good thing. Thanks a lot for your support, and we'll see you next time.