Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov

Kafka in Action with Dylan Scott

Confluent, original creators of Apache Kafka® Season 1 Episode 42

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 38:15

Author Dylan Scott tells all about his upcoming Manning title Kafka in Action, which shares how Apache Kafka® can be used by beginners who are just starting out their own projects and dispels common Hadoop-related myths, as Kafka has grown to become a powerful event streaming platform beyond big data ecosystems alone.

To get 40% off Manning products, use the following code: podcon19

EPISODE LINKS

SEASON 2
Hosted by Tim Berglund, Adi Polak and Viktor Gamov
Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
Music by Coastal Kites 
Artwork by Phil Vo 

  •  🎧 Subscribe to Confluent Developer wherever you listen to podcasts. 
  • ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
  • 👍 If you enjoyed this, please leave us a rating. 
  • 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
SPEAKER_01

As Apache Kafka continues to grow in sophistication and also in adoption, helping to get new developers up to speed with it is a key concern for the community. I mean, hey, that's a big part of what my own work is here at Confluent. Dylan Scott is a longtime software developer who wants to help solve this same problem. So he's writing Kafka in Action, a book that's an introduction to the basics of Kafka. I got to talk to him today about the book, the process of writing it, and some technology trends in general. It's all on today's episode of Streaming Audio, a podcast about Kafka, Confluent, and the cloud. Dylan, welcome to the studio.

SPEAKER_00

Yeah, thanks, um, Tim. Excited to be here. Um, you know, have to you know plug it. I'm a fan of this podcast and what you and Gwen do, um, even on the YouTube channel. So super excited to be here.

SPEAKER_01

Oh, thanks. I appreciate that. It's always nice to know it's hitting the mark. Yep. Um, so you, sir, are uh in the, I should say, in the process of writing this book. It's uh in MEAP right now, right? The Manning Early Access Program. Correct, yep. Okay. Uh and it's Kafka in action. Now um tell us a little bit about you. What uh what do you do?

SPEAKER_00

Yeah, I'm a software developer.

SPEAKER_01

You're not bookwriting, I should say.

SPEAKER_00

Yeah, yeah. So full-time um software developer, um, kind of done a bunch throughout my career. Um, you know, started an infrastructure automation, doing some Perl scripting. So um always like to point that out. Perl still alive and kicking, um, even though it doesn't get the press. Um but kind of transitioned throughout my career to more app dev, um, and then recently um really focused on DevOps enablement. So kind of the gamut, including some Hadoop in there, um, which is when I first got introduced to Kafka, um, and then been able to kind of leverage that knowledge throughout some of my my other skills and and work through uh in the in the insurance industry. So um yeah.

SPEAKER_01

Cool. Cool. Um so insurance, so you're you're dealing with occasionally large sets of data and kind of yeah, uh insurance is one of those businesses that really only works at scale, you know. That's kind of the that's kind of the idea. Yep. Uh so you can get big data sets. Yeah, and you know, oh sorry, go ahead. I just I realized I said big data. I didn't mean to say big data. Let's come back to that later. I meant just data sets which are larger in size than in some other lines of work.

SPEAKER_00

Sure, sure. And you know, it's not always the most um, you know, I wouldn't say the flashiest uh thing, but yeah, there's just so much data out there um and you know, over the years and and some of the sources can get, like you said, humongous, so just the amount and the speed of data.

SPEAKER_01

So yeah. It's funny. Um you say it's not it's not the flashiest business. And um I've done I've I've spoken at conferences, sort of uh regional conferences local to the Midwestern United States, which I mean, if you're not from the US, even if you are from the US, you might not know financial services are a dominant industry. There's like a lot of insurance in the middle part of the comp the country. Sure. Um and I have found that uh you know, like the the New York fintech, high-speed trading things, there's and you know, Chicago exchanges and all this, you get some pretty exotic technologies to get things done uh at scale and with extreme performance. But insurance in the Midwest has a reputation for adopting technology more slowly uh than on the coasts yet. Uh here you are uh writing a book on Kafka, which is which is still a fairly forward-leading thing. Sure.

SPEAKER_00

Yeah, it's amazing to see though, it's explosion. Um, you know, even from I think I started with Kafka 0.8.2 and um you know, past five years, whatever, it's it's already exploded, you know, version two officially. Um not sure what the latest latest is, but um, yeah, I mean it's it's just grown so much. Um kind of cool to see.

SPEAKER_01

Yeah, it is. So that we're recording this on, I don't know if I should even say this, but uh June 25th. This this may may be published uh weeks later, you know, we've got a backlog of things. And so this will come out sometime after June 25th, okay, 2019. But as of this very morning, it's uh Apache Kafka 2.3 was released. Cool. Uh so yeah, but what uh the question that is really top of mind for me, and I say this as a person who has written uh a couple of technical books, what possessed you to write a book on Kafka? And I'm speaking of it as if it's done, I know it's in meep and you're still working on it, but why did you do this?

SPEAKER_00

Yeah, so definitely um just one of those things is I kind of learn, um, you know, different people have different styles. I'm one of those people who still love um kind of learning from manuscripts and books um better than even lectures. So um it's kind of been one of those things that's always how I learned. And um, you know, Kafka, you know, it definitely is very simple. And uh, I mean, let me rephrase that. It can look very simple when you start looking at it, but then there's some complexities. And I think that coming from a beginner standpoint, which is what I'm trying to do with this work, is you know, you say, oh, cool, I have a producer-consumer, um, I read and write messages, but then there's so much more under there. It's how you handle your committing offsets, um, what consumer groups mean, what you do in failure scenarios. So I think that it's deceptively um simple when you look at it the first time and you're like, oh yeah, you know, it's it's a queuing system. Um, but I think when you start to look at the power, what you can leverage, um, that's where it gets really interested. And so try to take this book and say, you know, maybe you've lived in a in a situation where you've you've always had nightly batch jobs, um, and that's kind of just how it's always been. And um, maybe we can look at Kafka and see different ways you can leverage um, you know, just the speed of that data, um, or or how often do you get updates? Does it have to be nightly, or can you present that information a little bit sooner? So that's kind of the driver, just trying to work through some of those complexities and make them hopefully easier for beginners um as they start their Kafka journey.

SPEAKER_01

I love it. You you made some good points there. The uh one of the things I love about being a part of this community is that the the technology at its center, which is Apache Kafka, um you said, oh, it's a queuing system. You know, that's the the beginner approaches it, like, oh come on, uh, message queues. You know, we've done this, we know how to do this, it's fine. Um but there are a few things, a few uh I think fundamental architectural decisions that Kafka has made that make it not an enterprise message queue and cause all of these other uh like they they they have implications that work themselves out in the kinds of architectures that grow up around Kafka. Um and so it's this deceptively simple data structure that anybody who could partially program their way out of a wet paper bag can understand. Um yet then uh it it really like has a claim on the way you think about architecture in in ways that are surprising.

SPEAKER_00

Sure, sure. Yeah, and you know, just the fact like when I always initially thought of a queue, you know, I thought a queue, it's a temporary between applications that actually do work and just kind of again, it really takes the power of your data and kind of puts it the center of your infrastructure. So um, yeah, I mean, uh we can go anywhere we want to go, dig it into that stuff, but yeah, that's kind of one of the most powerful aspects. Um, you know, it's it doesn't have to be a server database uh web app. Um so definitely changes things.

SPEAKER_01

The other thing you said there was that you were trying to make this accessible for the beginner. And uh I really appreciate that. So talk talk to that a little bit. Is that is that a goal of yours in the book and why?

SPEAKER_00

Yeah, so we talk a lot, um, developing the book with Manning. Um, you know, what's our our qualified readers? Like, what do you have to bring to the table? So it's really about um someone who's probably familiar a little bit with Linux, uh, you know, how to run commands. Um, you know, also just it would be nice if you have some distributed systems, um, you know, if it's not totally foreign about everything not running on one VM or node, but you know, really that's it. There's not too many more requirements. Um, we try and take you from beginner, not having um used Kafka before and kind of introduce you step by step to okay, I I have it started on my machine. Um, let me do the consumer producer and produce a message to the topic, and then how do I look at it using like the console consumer? So um, and then from there we go into producers, um, using the Java client, just because it's kind of the most standard. Um, I think obviously Java just has that um, you know, language-wise, it has a pretty big um market that a lot of people are familiar with. Um so yeah, just trying to do that step by step uh and get people up to speed. Um, if they feel like, oh, you know, it it's it's too complicated. That's what I'm trying to kind of um you know battle with this book is is making people feel confident as they approach it.

SPEAKER_01

Yeah, no, I I really appreciate that because that's uh that's just a personal passion for me, too, is how do I help the new person get onboarded into this thing? Uh and so I really appreciate that. It might be so the the uncharitable interpretation of that is that I'm I've always been kind of a generalist and never really super deep dived into anything. So, like the uh you know, how to optimize uh uh buffering on the producer to to and and axe to uh maximize throughput, minimize latency, and all the like I'm not your guy. I I know I know who is, I know how to go do that, but I'm I've never been that person. Um and so you know, the uncharitable interpretation of of me is that, well, you know, he doesn't know the details, so he has to stay with the beginner stuff. But the more charitable interpretation is I'm the same way. And it just matters to me that that there are a lot of people, you know, that this is a technology on the rise.

unknown

Yeah.

SPEAKER_01

Which means there's a lot of people trying to come into it.

SPEAKER_00

I agree. Um, and you just think of you know, the other thing is we have Kafka streams in in action, we have um the Kafka the definitive guide, and you're like, well, do we need another Kafka book? Um and then you start to think about other technologies that are foundational. Um, you think of how many Java books are there, how many Pythons. So um, you know, I do think there's room for it in the market, and I think that certain people, you know, I I hope they can grab it and go um kind of develop. So um yeah, definitely I I agree. I think it's a it's exploding. Um so hopefully help people on that journey.

SPEAKER_01

Yeah, because there's a lot of people who are trying to get in. And uh you mentioned Kafka Streams in Action. Um, that's uh uh the author, Bill Bedric, has been a guest on this podcast a few months ago uh talking about that book. And that's an example of like you need to know what's in Dylan's book before you can appreciate what's in Bill's book. You're you're gonna be super confused if you're not at Dylan level before you go to Bill level. Um you know, Kafka Streams is about how to how to do stream processing and build applications on top of Kafka, but you gotta know Kafka first.

SPEAKER_00

Yeah, agreed. You know, and to be fair, you know, he does give a overview. I think a chapter or maybe even chapter four in his book is about core Kafka, but um, you know, even listening to that podcast um with him, you know, the references to your consumers and consumer groups, um, yeah, this is, you know, it's kind of hard to build on that um knowledge and get to streams without having those core foundational pieces really nailed. So that is part of, you know, why why I think this is kind of that expanding upon that that overview um because his book definitely goes deeper into the streams and the um not sure if it touches on KSQL, but um, you know, where those where those technologies lie past those foundational concepts.

SPEAKER_01

So Right, right. Um and again, just the you know, when when you look at a technology that's on the rise and is suddenly important in the the kinds of systems that people build in their jobs, uh you know, the certain certain of us in the community who like to teach, and if you write this book, you're one of those people, um you know, you have a responsibility to make it easy and pleasant to come on board with the new technology, right? Because there's a lot of people who it's it's frankly a little scary sometimes to be a developer because everything is new every five years, and like what you were doing if you've been at it for a long time, what you were doing 10 years ago, it almost doesn't matter. Uh so you're always learning new things and just helping those those developers who are uh you know trying to make an honest living doing stuff. Uh, here's this brand new thing that upends everything I knew about how I thought I was going to build applications. Let me make it easy. Yep.

SPEAKER_00

And you know, cool. And that's part of it is just trying that's really the details, kind of that bridge where um, you know, maybe your traditional data um database admin and you're familiar with right-ahead logs and kind of trying to type that yeah, or tie that into um, you know, that concept of the log and what's happening and those events and you know, even how that can be done with with Connect or something like that. I think that's kind of where you're trying to grasp for those straws and try and bring people in with some of that stuff they already are familiar with.

SPEAKER_01

So yeah, that's definitely a huge part of trying uh to connect with people when uh when people are learning Kafka, uh and let's assume they have some experience as software developers and they're they're new to Kafka, uh what are some gotchas that you've seen people trip up on? And uh whether whether you address these in your book or not, and if you do, you know, tell us. But what are what are gotchas that you see?

SPEAKER_00

Um yeah, a lot of it's um kind of thinking about the usually producing, I feel like's pretty straightforward for most of our the folks I've worked with, but the consuming end, um really it's about um I have a consumer group, um, and I have these five consumers, and I don't understand why they're not all busy working. Well, it's really about how many partitions did you have for that topic. And um, so I think that's part of it is kind of not realizing how consumer groups work. Um, you know, and along with that is really how you commit messages. So I've handled this message, but I failed. But some reason it looks like in Kafka I've already seen it before. So really taking the discipline and looking at um how do you how do you commit those offsets as you've seen them? Is it asynchronously with a callback um to handle those failures, or is it um, you know, kind of that some people are like, oh well, if I put the put the you know auto-commit on, it will be it'll be okay. But um, so I think really that's one of those points is people are used to consuming a message um and then it's gone. They don't think about, well, what happens if I reread it? Um did I commit it properly? What happens if I do see that message again? I think that's what really trips people up, after especially those that have used other queuing systems when they're starting out, if that if that makes sense.

SPEAKER_01

Oh, it does. It does. Consumer uh consumer commits and just basically understanding uh consumer groups. So yeah, and you also have a background in ops. You said you've kind of moved more towards app dev recently, but are there any operational gotchas that you see?

SPEAKER_00

You know, basically make sure um you you have enough disk space for what you're planning. But um really live by, I guess. Yeah.

SPEAKER_01

Um that's a things just don't like nothing works, right? Operating systems make this assumption about the world that that there is disk space. It's like, you know, the strong force and the weak force are gonna be a thing. And if suddenly they're not a thing, I don't know. Nothing's gonna work.

SPEAKER_00

Yeah. Well, yeah.

SPEAKER_01

Okay, so disk space.

SPEAKER_00

And that really kind of hints towards operational-wise. A lot of people, I think, um, and it's hard, is you just need to kind of know the shape of your data. Um, so you know, what what are you looking at? Are you looking at I'm gonna put a gig a day of events on this disk, or is it gonna be a terabyte? I mean, are your messages gonna be one megabyte, which you know, hopefully not, but um, that's really what it kind of looks like, is I think not every topic looks the same. Um, you know, are topics keyed? Um so really the hardest part, I think, even operational wise, is kind of getting close to those developers knowing what the data is gonna be shaped like. Um is it compacted or not? Again, those all play into um it's it's not a system that you just send a message through and it's gone. It's kind of one of those you do care about um how your how your data is retained. Um so that would probably be the biggest challenge is right.

SPEAKER_01

You um and this is um I'm I'm promoting you to consultant from author here for a second, but you work in financial services, and it's typical in financial services for uh deployment to be on-prem and less in the cloud. Um that's not universal, but uh it's a trend. Sure. Right. Is all your operational, is the bulk of your operational experience with on-prem, Kafka? Um correct.

SPEAKER_00

Yeah, so definitely everything involves risk and and right now, um, you know, definitely looking out the future though, playing with some of the stuff in Kubernetes now, with obviously step one was getting some of that persistent storage worked out in the state full set um type of logic. But yeah, definitely still looking ahead though, even though most has been on-prem. Um excited for to see what the the Confluent operator looks like when it's finished. Um, but yeah, definitely keeping an eye on it and kind of doing some POC stuff. But yeah, it it is a different world, um, operational.

SPEAKER_01

Being being on-prem, you mean correct. Yeah, yeah. Um and hey, shameless plug for uh confluent networking. I appreciate that. Yeah. Uh it is, I mean, that that's in the world of Kubernetes operators is is uh a thousand flowers are definitely blooming right now. Cool. Uh because you know uh Kubernetes is is uh not everybody's doing it, but it's one of it's it's become one of those technologies where you're doing it, you're planning to do it, or you're embarrassed that you can't or something. You know, it's it's it's assumed. Uh and for a system like Kafka, you know, you need you need an operator to make that reasonable. Either you're gonna build one or you're gonna adopt one or you're gonna buy one.

SPEAKER_00

Yeah, and you know, that's part of the you know, it's it's not addressed as much the Kubernetes just um because yeah, the the most of my stuff is um on-prem, but um, you know, just those fundamentals of you know that you need storage and then how do you handle it, and you know you need a quorum. So when your pod goes through a um a disruption event in Kubernetes, right, if the administrator is rolling, how do you make sure you get those? Um you you want to make sure you have a quorum. So hopefully you'll know what that means when you get to different environments. Um again, it's laying that foundation, hopefully that crosses no matter where you deploy. Um so that's that's what it's you know, hopefully helping, even though um we don't have like an operator in the book itself or you know, right, right.

SPEAKER_01

That's not that doesn't need to be in fact, uh, you know, just thinking of the title Kafka in action and being this we would like to introduce beginners to Kafka, uh getting a chapter on how to deploy Kafka in Kubernetes uh doesn't necessarily seem like you would even be on the hook to write that for this book.

SPEAKER_00

Yeah, you know, and what I've always addressed it, which again, it's you know, Kafka itself is can be a complex animal. Um, and then Kubernetes itself can be very complex. So I'm gonna try and teach those in this book. It it is kind of one of those things that's just such a big scope. Um, I didn't think it would be fair. Um so yeah, to kind of address that point, it's something that's definitely crossed our crossed our minds, uh, my editors and myself. So yeah.

SPEAKER_01

Well, if when you're done, if you have not completely sworn off the written word, um when you're done with Kafka in action, you could always do like a deploying Kafka or Kafka on Kubernetes or not nice. Yeah, nice. You've you've got options.

SPEAKER_00

Yeah, there you go.

SPEAKER_01

Cool. Um so uh uh yeah, good good thoughts there. What um so far, what has been your favorite thing to write in the book?

SPEAKER_00

Uh you know, definitely it's it's kind of interesting. Um yeah, you get you got me. How about how about least favorite, I'm sure. Yeah, yeah, yeah. So security hasn't been published out there, it's coming chapter 10. Um that's probably been the the hardest because um honestly, I I know you talk about Kubernetes and some of those principal stuff. So um trying to make that something you can play with locally, I think, you know, again, it's kind of pushing that. I don't want this to be a book about um Kerberos security, you know, but um so that's probably been the most challenging, trying to make sure we keep that uh accessible. Again, that's that'll be chapter 10. Um, you know, but really the I love digging in the other chapters, you know, it's kind of hard um because in daily work you do what you need to get done. You go to the API docs, you you try and knock it out. So um that's kind of been the joy of the book is getting a dig in. Um some of those other pieces about uh you know even the how the consumer group works. Um, you know, that's that's kind of been the most interesting, I think, from from the po you know, the positive side.

SPEAKER_01

So sure. Uh well let's uh let's dig into that a little bit. Um you mentioned consumer groups as a beginner gotcha. Sure. But that's a that's a fairly elementary thing. Like you know you you can't you can't run more consumers in a group than you have partitions and and scale that way. And you know you have to think about consumer offset committing and all that stuff.

SPEAKER_00

But is this like the rebalance protocol and and that yeah yeah and I I think it's it's definitely interesting to think of you know and part of that is even that you know you s you seem a basic but then you're talking about well what if you do have a keyed partition that suddenly s switches consumers. So um that's where you know digging into well what about these rebalanced listeners? How can we hook in and kind of notice that is happening even, you know, nobody's pushing a button or no manual change, but being able to rebalance it and get those events kind of because yeah it it always you know keyless are um those without keys and those that are keyed seems to kind of throw some kinks in the plans. And initially you might not think of that um you know oh great you know somebody took over the work but um is that what you really wanted to happen um if you grabbed another key partition. So yeah that's kind of the those gotchas that we're trying to um you know you get one level and um just trying to see how far we can go and and and get that more clear um kind of take that mystery out.

SPEAKER_01

Are you referring there to uh a partition with no key and then elastically scaling the consumer group like at at runtime adding a a node and having the work be rebalanced?

SPEAKER_00

Correct. And even you know if a consumer drops out of the group and suddenly um you have a consumer that that's already doing one keyed partition take over another one um you know just the just that sort of um you know those changes happen and and is your code robust enough to kind of to handle that situation. So yeah I I think it applies both ways. I'm not trying to think.

SPEAKER_01

And I guess in general if you're if your consumer is stateful then that's potentially a pain in the butt.

SPEAKER_00

Correct. Yep. Yep state you know it's funny state gets a bad name but it's everywhere. So that's the thing.

SPEAKER_01

I I I mean you know because uh when I explain again I'm I'm usually doing this at the beginner level and I'm trying to get people to understand what a consumer group is cool uh and and what what the consumer library API gives you. Like what what is the contract there? Um I always say it works great and it's elastically horizontally scalable for free. You don't have to worry about a thing as long as your consumer is stateless. Yet you know think about what your con what any code you're ever going to write is going to do. How much filtering is there really pretty much everything's stateless. So states of pain but I mean everything is stateful and states of pain but you have no choice.

SPEAKER_00

Yeah it's um yeah it it's interesting but but yeah it's it would be nice. I mean ideal yeah it would be be awesome to not have to worry about state but yeah um yeah it's it's one of those things that it does complicate say things some things and um but yeah for business value you you have to you know it's there.

SPEAKER_01

So and how far because this and this whole question sort of is coming up the stack from Kafka into applications on Kafka. Sure. And how far do you go in the book? Is that legitimate territory for you to explore?

SPEAKER_00

You know we it's really small scenarios so really the case is like well this one um I think I have an alerting example where what happens if you you can't miss this alert? So we talk about making sure we have enough acknowledgments um you know do you wait for the response um once you come you know you produce the message um so yeah it's more uh small based on on small app examples um really trying to say you know as we explore the patterns you know maybe this use case for this little example um can we use it most once or um do we need to do many um we don't really cover exactly once two depth you know because heck that's a that seems to be still a controversial topic but um you know and and that has some other bounds of being inside the Kafka um ecosystem right so um we we really try and touch on little examples not not full blown enterprise apps but kind of those situations where hey I I know I need this message so what's my best shot um what kind of config do I tweak um to make sure I get this message across and hey maybe I see it twice right um but then your code knows how to handle that so that's really um hopefully that answers what you what you're asking.

SPEAKER_01

Yeah totally totally does it's um I think Kafka Kafka in production is a future book and software architecture on Kafka is really where that if you if you go down that path of the question that I asked too far then you're talking about how to build event driven systems. And you're not really talking about Kafka anymore.

SPEAKER_00

Yeah you know there's some great um works out there that do cover kind of that theory um but yeah this is it's supposed to be practical I hope it is uh you know especially um I want to write this I want to make sure I get my message hopefully I have an example that's suitable for for readers to see that this is how I would do a producer um yeah that's kind of what it's geared towards versus the um you know I'm trying to think of the you know it's not as much of the eventing philosophy as hopefully the practice so right right um on that point um is Kafka a big data tool yeah so one of the things I I try and talk about um you know when I originally started working with Kafka it was brought in as part of uh Hortonworks or Cloudera is kind of bundled with those tools um and when you got Hadoop correct sorry yep yeah so clearly the answer is yes right well it's very interesting because a lot of the the things of being built for failure um I think Hadoop and big data some of those labels got applied to some of this um just you know it's big data they had to just distribute they had partitions some of this terminology um kind of got out there and I think that when I started and especially where my coworkers it felt like it was it was bundled um in that space but you know looking further um you know looking at it so why it does have some of those same philosophies as yeah things are going to fail let's build in um replicas and hey we're gonna have so many um consumers let's kind of balance out the workload with with the partitions um I know I'm not saying that elegantly but um but taking that and making it its own system so again even in the book I try and address that when I started some people thought that Kafka was bundled and that it actually wrote to um you know HDFS which it has you know it doesn't use that file system but um I can see where users who are used to that environment only might get that that feeling um so yeah just trying to kind of lift it out of that space it doesn't have to be tied to big data um I know there's terms like fast data out there all that stuff but um yeah definitely I I do want to put that that myth to bed that it's it's strictly a big data tool. So um I I hope I did a fair job and say even though some of the terminologies are are similar that um yeah the this this tool is definitely powerful enough to to be its own infrastructure it doesn't have to be tied.

SPEAKER_01

Nice I like it. So yeah it's it's funny what's happened with that term anyway it's a little bit embarrassing even to say these days um it's it's it was an era and it gave birth to a number of important technologies and I think genetically you know Kafka sort of co-evolved with Hadoop and it was initially always paired with Hadoop it was this way of getting things into Hadoop and people used it for their queuing purposes. So I you can understand why people make the association.

SPEAKER_00

Yeah and you know let's talk uh Apache Flume right I mean it was I I mean I don't know but to me that was one of the first usage of an app that used that as an underlying um technology for ingestion right so yeah it was definitely um you know when I saw it it was this tool that was a part of a suite of other tools so that's kind of how I was introduced um so yeah I can I can see where that started from but yeah it's it's interesting to see how it's kind of taken on a life of its own and um really kind of I think you know jumped out of that suite to stand on its own if it ever was part of the suite.

SPEAKER_01

But yeah it it it probably was you know in practice in the minds of some but it it's pretty clear now that having a distributed log with the properties that Kafka has, you know, has its own uh it is its own reason for being and it it's not it's not tied to some distributed computation or distributed storage framework. It's it's like it's uh its own thing.

SPEAKER_00

Yeah yeah I like the way you phrased it that's um definitely a good good take um more elegant than than I said so cool.

SPEAKER_01

Yeah and and where I I guess in summary um and uh this this so I'm not setting you up to say Kafka belongs everywhere by the way but where does Kafka belong like clearly the the days of it being uh a little helper to your Hadoop cluster are over but what what's the right place for it?

SPEAKER_00

Yeah so you know I really think that if you have um an organization that really um uses its data um needs that data available um I think that if if you have that data source that really powers your organization it's really a fair place to say Kafka might be um where it's at so just because like I grew up in an environment where um to get access to data it seemed like you always had to go through a service or an application that kind of did its own view of the data. If you wanted to get something out of that data that wasn't there, you had to go to the app owner or the service owner and kind of say hey can you make this data available? So I really think that for for me where I see it fitting in is where you can have whether it's marketing or operations or analytics, they can go to that data. They don't have to worry about it being munged for that service's purposes. They can go to that true source of data um kind of use it without having to ask permission, ask for formatting. So that's really where I see it shines in what I've been looking at personally I you know and definitely welcome to your thoughts if you know if if that makes sense or that's too out there. But um that's kind of where I've seen it shine is just that availability of data is kind of um really really what's powered some of our our stuff moving forward.

SPEAKER_01

Yeah I'm not I'm not sure I could have said that better. That's a that's a very good and general kind of set of guidelines for where it makes sense. Cool. There's you know some people tell like a real-time ETL story about events happening and I I need to make I need to do computations on them now and make results now. That's totally legit.

SPEAKER_00

Yeah and you know one of the things I can think of is even like a proprietary database that we we deal with from a vendor we don't have access um but we can take that data out of um the database whether using connect or or um you know there's there's different connectors but um you know just the fact that we have access to that data and even if it's not made available we can make it available because just the operational that proprietary nature makes it so we can't really touch much of it. So that I think hopefully that's another example of where it's really helped not only extract the data but then made it more available where we're not going against a proprietary vendor and that we're not slowing down the operational workloads that traditional we'd we'd have to if we were trying to mine that data from from their UI or API.

SPEAKER_01

And that would be um a uh species of the genus uh mainframe offload where in this case maybe instead of mainframe it's database I don't like offload. Yeah it's it's fair yep yep yeah and um the other the other story is the microservices story that people talk about which is kind of just a different form of sure the use case you described you know I want all my data in one central accessible place except you know there's this idea of the services that hang off of it being reactive.

SPEAKER_00

Yeah.

SPEAKER_01

Is that a thing that you see in your definitely um I'd say I've seen it um not not huge at where I'm at right now to be honest but yeah it's definitely one of those things again it's just um you know we're all lo those those core um you know software tenants you want to be loosely coupled um you don't want to wait for a um a totally separate app to respond to you that whole request response I think that apps kind of fell into for years um yeah it's it's definitely opening up um those options so pretty pretty cool to see um you know honestly it just hasn't taken off here yet but I think there's so many use cases um as you know Tim that um it it might be right around the corner for us so nice nice yeah request response is uh easy paradigm because it looks like method calls and you can even write wrappers that make it method calls as far as you know sure sure uh it doesn't scale right it's probably not a good thing for building the little distributed systems that we're building now yep cool uh best guess how many books are you gonna write in your lifetime um you know I'm hoping to finish this one hopefully people um find value and again it's just um you know such a such a cool technology that I hope um can do it justice um hopefully get people interested so yeah no no commitments though that's probably wise yep yeah cool my guest today has been Dylan Scott Dylan thanks for being a part of streaming audio yeah thanks a lot Tim I appreciate the time and thanks for having me. Hey you know what you get for listening to the end a Kafka Summit discount code. Kafka Summit is coming up on September 30th and October 1st in downtown San Francisco and you can get 30% off if you go to Kafka-summit.org and use the discount code Audio19 during checkout. Just enter Audio19 while registering at Kafka-summit.org and that 30% off is all yours. I'd love to see you there. But hey I hope this podcast was helpful to you. If you want to discuss it or ask a question you can always reach out to me at TLBergland on Twitter. That's T-L-B-E-R-G-L-U-N-D or you can leave a comment on a YouTube video or reach out in our community Slack. There's a Slack signup link in the show notes if you want to register there. And while you're at it please subscribe to our YouTube channel and to this podcast wherever fine podcasts are sold. And if you subscribe through iTunes be sure to leave us a review there. That helps other people discover the podcast which is a good thing. Thanks for your support and we'll see you next time