Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov
Hi, we’re Tim Berglund, Adi Polak, and Viktor Gamov and we’re excited to bring you the Confluent Developer podcast (formerly “Streaming Audio.”) Our hand-crafted weekly episodes feature in-depth interviews with our community of software developers (actual human beings - not AI) talking about some of the most interesting challenges they’ve faced in their careers. We aim to explore the conditions that gave rise to each person’s technical hurdles, as well as how their experiences transformed their understanding and approach to building systems.
Whether you’re a seasoned open source data streaming engineer, or just someone who’s interested in learning more about Apache Kafka®, Apache Flink® and real-time data, we hope you’ll appreciate the stories, the discussion, and our effort to bring you a high-quality show worth your time.
Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov
Ask Confluent #5: Kafka, KSQL and Viktor Gamov
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
Gwen is joined in studio by co-host Tim Berglund and special guest, Viktor Gamov, a new member of Confluent’s Developer Experience Team specializing in Kafka, KSQL and Kubernetes. In this episode, we’ll find out: Does Viktor know what he’s talking about?
EPISODE LINKS
SEASON 2
Hosted by Tim Berglund, Adi Polak and Viktor Gamov
Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
Music by Coastal Kites
Artwork by Phil Vo
- 🎧 Subscribe to Confluent Developer wherever you listen to podcasts.
- ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
- 👍 If you enjoyed this, please leave us a rating.
- 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
On today's episode of Streaming Audio, we take questions from the community about KSQL, Kafka streams, and some core Kafka stuff. Plus, we have a special guest. Welcome back, everyone. I'm your host, Tim Berglund. Let's get to it.
SPEAKER_01Hi, everyone. Welcome back to As Confluent, the show where we answer Kafka questions from the internet. And with me today, we have my beloved co-host uh Tim Berglund.
SPEAKER_00Hi, Gwen.
SPEAKER_01And we have a special guest.
SPEAKER_00We do. Hello, internet. Um, it's Victor, longtime listener, first time caller.
SPEAKER_01So Victor is the newest member of Team Steam.
SPEAKER_04That is correct. Uh developer advocate here at Confluent. Uh delighted to have Victor on the team.
SPEAKER_00Delighted to be finally on this podcast, guys. Sorry, but I have to plug it because I love the show. He's pretty excited.
unknownYes.
SPEAKER_01And you're going to see him all over the world. If you did not yet see him all over the world, you are going to because we're going to basically be sending him to speak all over the place.
SPEAKER_04We keep this guy doing it.
SPEAKER_01He'll talk about Kafka and KSQL and Kubernetes and clouds and all the cool stuff. So let's ask him some questions. So, question from our friend of Confluent, Alex Ott, who's a very friendly one of our tweets.
SPEAKER_04That would be fair to say. Alex, thank you, by the way.
SPEAKER_01Is it he asks, is it possible to extract information about window start time gener window start generated by aggregation with KSQL if I consume the topic from a normal client, presumably a consumer, and not KSQL?
SPEAKER_04Yes. May I? Yes. So you've got um just to frame that, if you don't understand the question, um when you've got a KSQL query and it's a persistent query, it's going to produce its results back to a Kafka topic. So you'll say something like create table monkey as select, blah, blah, blah, and there's an aggregation happening there. I think you said aggregation in the question, right? Yeah. There's an aggregation. So you're making a table, and that table is the result of an aggregation on a stream. And in case, you think, well, it's a table, I'll deal with it as a table, but that table has an underlying Kafka topic. Those messages are being produced to that topic all the time. So if you want to access them from something which is not Kaseqal, you can do that. That's one of the cool things about KSEQ, is it's just a Kafka topic. You can have a conventional consumer then go consume from that, or you can Kafka connect it, or you can do whatever you want is just a topic, right? So and the question is, can I get at the aggregation window start time in those aggregated results? Because there's a window, obviously, if there's an aggregation. Anyway, all that to say, yes. Next question. No, okay, let's try a little bit more difficult than that. The answer is yes, you can get at those things. And uh our friend and occasional podcast guest and colleague, uh Hohat Jafarpur, answered this actually on Twitter as well. There is a CERTI called the windowed key CERTI. And if you use that, um then you it's it's basically you're like looking at the results as a windowed thing, and that that window time is a part of the key, and you can get that back. So use the windowed key CERTI and um when you're consuming from that topic and you get what you want. So it's a lengthy answer, but to a very good and specific question.
SPEAKER_02What do you say, Victor? He got it right?
SPEAKER_00Uh yeah.
SPEAKER_02So You're not just saying it because he's your boss?
SPEAKER_00No. One of the things that uh probably some people can notice if they start to consume in this like a KeySQL produced topics, that's some uh additional columns. Uh in general, uh they're they're not there, but like if you want to read this from the conventional client, you need to use like a specific uh serializer to you know send what exactly is there because Kafka doesn't really care what's inside, uh it's only your app. So your app needs to understand, and for that matter, we have a special deserializer that will bring the things in, make it more sensible.
SPEAKER_01Okay, next question. Tamar asks Is there such a thing as oversubscribing to Kafka? Where do we draw the line and how? So I think everyone has their own opinion about what how much subscribing is oversubscribing?
SPEAKER_04Oversubscribing like uh consumers?
SPEAKER_01Uh like is that what I saw? Do you have a different interpretation? For me, you have consumers, I subscribe. How else do you oversubscribe to the code?
SPEAKER_04Yeah, I just I thought maybe like can I give it too big of a place in my heart? There's probably not.
SPEAKER_01Obviously, no. Obviously, not that.
SPEAKER_04But yeah, okay, so consumer.
SPEAKER_01That's that was my guess. Like too because you know, something like kinesis limits your fan out quite a lot.
SPEAKER_00Uh profoundly. Yes. Yes. So um idea about Kafka again is that we have a lot of readers, much more more, much, much, much, much more readers than actual producers. So this is why essentially Kafka is designed to be like oversubscribed. Like um, I I can assume that this question comes from uh many people who has a traditional messaging background where they think, okay, so if my broker will do the push to me, um do the actually destroys uh by multiple like hundreds of consumers, so they need to do like a push for hundreds of the uh subscribers. Kafka is works a little bit different. So I don't think it's kind of like oversubscribing. In other thought about oversubscribing is like putting a lot of use cases around. Yes.
SPEAKER_01And that's a good thing. You want to have a lot of use cases around it.
SPEAKER_04Trevor Burrus, Jr. And so I'm I'm always tempted to say no too, right? Because again, architecturally brokers themselves don't know who's listening. They don't need to know who's consuming. There are just requests that come in for messages in a partition at an offset, and you know, consumers are managing that. And as long as from a raw pub-sub standpoint, they can get the work done and everybody's CPU and I. The answer is no. You know, the fan out is infinite, and that's that's always the answer you want to give. But then you dig into a question like this, and we have not dug into Tamar's question, but then they're like, oh yeah, because I wanted one per user and I've got three quarters of a million, three-quarters of a million users. Is that okay? And I'm sending files to Kafka.
unknownRight.
SPEAKER_04Right, right.
SPEAKER_01So there's always this thing that even though you can, should you do it? Is that a good idea? Trevor Burrus, Jr.: This is perfect.
SPEAKER_04We need to say this in the Jeff Goldblum voice. You know, we didn't we didn't stop to think whether we whether we should uh Jurassic Park. So um So where does it fall down, right? There's like metadata, there's the consumer offsets topic. I mean what if you just crank the exponent in ridiculous ways, where do you fall down with lots and lots of consumers? Where would you guys predict?
SPEAKER_00Well in my my favorite answer to this, um if you cannot measure, you cannot control it. So if you know what exactly your numbers you're targeting here, like how many consumers and you can project at some point because Kafka has this ability, you can project that uh growth, and this is what we usually suggest to people to do is maybe even use Confluent Control Center to monitor your bottleneck because we can guess autal bottleneck.
SPEAKER_01I mean, my guess was probably CPU, but possibly network connections. But you know, who knows? You need to have a tool for measuring it. So exactly.
SPEAKER_04Watch it and turn the crank and see. Because I'm I'm guessing there is something behind the question where somebody's wondering.
SPEAKER_01A lot of times I you go to people like on Twitter, you ask how many consumers do you have? And people say, Oh, I have 30,000 consumers, I have 50,000 consumers. And they say, interesting, how big is your cluster? And they say, Well, I have 100 workers. So like you have to remember that when you at the end of the day, hardware is hardware, a machine is a machine. If you want to go with a lot of consumers, you may need to have a larger cluster too.
SPEAKER_03Got it.
SPEAKER_01Okay. Next question from David Newberg, who is uh Newberger. Very uh good friend of my Twitter and uh shows up at random Kafka conferences. Like we know him, we like him. Uh he asked, generally speaking, is it more costly to upconvert or downconvert message formats for messages sent to Kafka? And I this was a response from me tweeting hey, we have a new guy in the podcast, let's ask him really hard questions. What's the hardest question you have?
SPEAKER_04And before you answer the question, um what does that mean?
SPEAKER_01Make sure everybody we're totally going to explain because this is a question that Victor and I had to go to the experts. So thank you, Jason and Ismail, uh Kafka committers, Kafka PMC members, for explaining everything about upconversion and down conversion because it turned out to be a lot trickier. My first guess was that it's identical, it's going to be exactly the same thing. What's the difference? Upconvert, downconvert, same thing, same thing. And it turned out that I was dead wrong.
SPEAKER_04Okay, so Victor, is it better to give or to receive?
SPEAKER_00Um let's um let's dissect this question. So, first of all, um a couple things, uh, what I like to do before I'm you know jumping to the conclusion, jumping to the answers, is to uh figure out what exactly uh this gentleman means in terms of cost, uh cost for whom. And we did this assumption that uh we're talking about Kafka brokers, and let's talk about this. And uh never go to uh the Ask Confluence show unprepared. So I prepared a little bit of props. So let's talk about this, let's dissect it. So uh first important uh things that this is how you draw Kafka in your uh diagrams. You're not using tubes, bricks, and stuff like that. You this is how this is Kafka broker. So and um we have, let's assume we have a Kafka broker with version 1.1. So um we have a producer with older version, which is uh 0.1. And we assume that uh when we publish this message to Kafka broker, this is what we meant by upconverting. We need to bring the format of the message that we have uh from producer to the internal message format that broker supports. So in this case, we're gonna be having upconvert. Now, on the other side because of a version mismatch in producer and consumer? So we will talk about this, but just right now I'm just trying to establish like baseline and uh common vocabulary so we can you know talk the things through. Okay. On the other side, we have a consumer that also has a version of the uh message different from version of the broker. So and we say that when a broker has a different version from uh from the consumer, we need to do down consumption. Why is that? Why is that why we need to do this? So, first of all, important thing that uh Gwen uh taught me is that first thing that every client will do during handshake, they receive the version. So in this particular case, um producer will uh receive version of the broker and it will establish version of protocol because producer version is higher.
SPEAKER_01Can I nitpick? 0.9 is a bit too old to know the version of the broker. So the broker in this case has to do all the work. I think we only added the version protocol in zero, is there zero ten or zero eleven? I think zero ten.
SPEAKER_03Okay.
SPEAKER_01So if you have a zero nine client, it will just send the old uh version regardless. Everything else in the diagram actually uses the version check.
SPEAKER_00Yeah. And in this particular case, it's gonna be up converting. So broker will bring the message format to the message format that broker has in configuration. Yep. By default, if you're not specifying anything, it will use whatever is version of uh Kafka broker. Now, in the down convert, uh, we will use the message format that uh consumer supports. Now, what is more costly? In this case, it's also I like the segue from the previous question where we're talking about oversubscribing. And one of the things that we usually tell about Kafka is that uh we produce less that we consume, we read more. The thing that we already mentioned here is a consumer fan out when we have more consumers that we uh usually you know produce.
SPEAKER_01So in this case remember, Kafka is an author. He wrote his book once, millions of people read it all over the world.
SPEAKER_00Exactly. This is this is this is pretty pretty good metaphor uh for I learned it from a very smart new developer advocate that we got. We can oh yeah, we can talk about this, but essentially um another thing that will affect uh performance of the broker and add the little bit of cost here is that message needs to be uh stored in memory. So when we're talking about uh like a typical producer and consumer, they have a same version. In this case, uh the message will be streamed from the file to the socket. This is what we call zero copy. Now, in this case, we're losing zero copy because message needs to be read into uh heap of the broker, needs to be converted uh to the specific format, and after that will be sent. So that'll be a copy and a copy. Yes, and this will be uh will grow exponentially, depends on the number of consumers, uh, and depends on the version of this consumer. So in this case, it will be doing this for every uh consumer um that will connect it here.
SPEAKER_01And before 2.0, we were really bad at doing it in some ways. So you could actually cause garbage collection pressure, run out of memory.
SPEAKER_04It used to be pretty painful by having client and broker version mismatches. Yes, especially if you remember those consumers. From the release notes of two now, talking about this. Yes. Exactly.
SPEAKER_00And I remember this from a very awesome talk that I learned from uh Kafka Summit London when the Gwen and Xavier were talking about monitoring, and one of the use cases is uh was explained in this talk. You just go and find this uh how we do this, like find the link below. Um uh in this talk, you will you will learn what kind of uh uh things to monitor and what kinds to see if you do like this upgrading of versions and things like that. Now let's talk a little bit about like new stuff. Like uh if we're running uh producer and consumer version higher than version of the broker, will this something like this happen? It turns out that um Kafka, producer and consumer, Kafka clients, they're very smart. And the one of the things that I already mentioned and the Gwen mentioned that they communicate in the the negotiated version. So in this case, producer with higher version know more about uh broker with older version. So in this case, uh producer will just operate on the same protocol level as a broker. So in this case, there would be no heat in the both sides. In this case, they will be just talking to the version that the protocol supports. So this is what I like about uh it's kind of like a shows the metaphor of um developer experience team. So we lift people up, we don't let them know. So we we're providing it. Yeah. So this is uh this is a picture that uh we draw before the show. I think I hope it is explains uh what we're trying to say here. I think that was great. But uh last but not least, always uh be measuring. So if you some cannot some uh if you cannot measure something, you cannot control it. So that's why um you probably need to do your own test because it really depends. Um in theory, it works.
SPEAKER_01And now we have new metrics to tell you exactly how many clients of which version you have connecting to Kafka. So this is all good.
SPEAKER_04Console control center.
SPEAKER_01Yes? Yes? No, not in Kafka, no. Actually, no.
SPEAKER_04Not as well. Okay, but new GMX, right? Yes. Okay, also in GMX.
SPEAKER_01Yeah. Uh but um yeah, the nice thing is that this really informs how you do upgrades. Because now that you know that the worst things that can possibly happen to Kafka is old consumers, and all producers are only the second worst thing to happen to Kafka, you know that you have to hunt down and upgrade all the consumers first, which is a problem because there is a lot more of them.
SPEAKER_04But upgrades are I already advanced for you. Oh, yeah, so I'm just I'm jumping in the gun.
SPEAKER_01I said you're the best co-host there. Uh so next question. Can you please give some recommendations on using Kafka headers? When should I put event metadata in the event itself? And when should I use headers for that? So uh people have been asking for headers in Kafka since at least version 0.8 when I joined.
SPEAKER_03That's when yeah.
SPEAKER_01And we only added them at version 0.11. So kind of makes you wonder what happened in between to make people change their minds.
SPEAKER_04A lot of no.
SPEAKER_01A lot of no happened in between. And it used to be that like the biggest Kafka player at the time was LinkedIn. Uh they owned most of the Kafka developer, not all of it, but a lot. And for LinkedIn, this they basically did what Marek Gerhardt um suggests. Why don't we just put all the information, all the metadata in the event itself? And LinkedIn was using Avro at the time, so they just had kind of a standardized uh Avro extras that they included everywhere, and all their Avro schemas included in it, and it was fine. But LinkedIn grew a lot. Time passed. People left the company, started new companies, who knows? And it turned out that new teams didn't want to use Avro. And between us, who can blame them? And they wanted to use things like JSON and Protobuff, but there were centralized teams that needed the metadata from every single message going to Kafka. For example, the team that is responsible for how many events are produced by this application versus the other application. They didn't want to learn, oh, you're doing Avron, you're doing JSON, you're doing protobuf, and let me figure out all of that. They needed something very standardized. So headers give you the ability to have something that is standardized across the various different message formats uh that go around in your company. And this is a this happened to be a huge deal. But since you're asking not just when to use um whether to use put metadata in the event or in the header, I think the good question is what do you do people do in the headers? What do you want to use them for? And I heard that Victor has a lot of experience with customers and actually know what people use it for.
SPEAKER_00I always have opinion about something, so this is uh what we call the opinionated developer at the kitchen. Um so one of the things that uh the Gwen mentioned that usually putting some metadata for perspective of audit, for perspective of tracing, for perspective of lineage, this is something that people uh putting in the special like format. For example, like Spring Framework use uh this concept of headers, which is not called headers and they were not available in uh in the Kafka. Um I'm talking right now about uh Spring uh Kafka uh library, shameless plug. Come to see me talking about Spring Kafka at the Spring platform uh next month. Next month in September, if you will be in the Washington DC. Um another uh use case uh that you can use uh they use it in the Spring Cloud stream to do routing, basically. So based on this header, you can specifically always cross-cutting things, though.
SPEAKER_04Exactly.
SPEAKER_00And one of the things that uh Gwen mentioned is for many um uh for for few organizations, um apparently there is a legal implications where you cannot read the content of the message. So in this case, it should be somewhere outside. And the header is the first perfect place like in terms of like all these regulations with the data you know, dealing with data and things like that. Um another thing that's uh I recently learned that the frameworks like OpenTracing that has a Kafka client, that they use headers to put this like a trace ID. So in this case, uh you get a lot of different things.
SPEAKER_01That makes a lot of sense, yeah, because OpenTrace doesn't want to care about your message format and whether you use JSON or other.
SPEAKER_00Yeah, precisely.
SPEAKER_01Um it's also used for lineage. I think that's pretty common. It's kind of like OpenTracing, but slightly different. A lot of time you need to know who produced the message again, sometimes for legal reasons, sometimes for GDPI reasons. Message that was produced in Europe may not be legal to replicate it into the United States. So now replicator, for example, has to be aware of the headers and rules around them.
SPEAKER_00Yeah. So Campbell thinks that uh with like with headers right now, pretty much every framework supports this. Um in in terms of like uh for producers, consumers.
SPEAKER_02Streams, connect, everyone.
SPEAKER_00For streams, uh we recently added support of headers in um in the in the um uh processor API. So we do have you have ability to. This is a good question. This is a good question. This will take some time to research for us. Um we'll get back to you. Yeah, yeah, I don't think I don't think KSQL uses this right now. But this is really, really good um, you know, really good use case for for taking care of uh things not related to data. Like when some other people need to manage. Encryption keys.
SPEAKER_01That's a good place to put your public encryption key.
SPEAKER_00This is an interesting idea. Yeah, it's probably uh worse to explore. Yes.
SPEAKER_04But the the the basic idea is when it's a cross-cutting concern like lineage or provenance, uh stuff like that.
SPEAKER_00If I may add a little bit of interactivity to our uh to our podcast, go write in the comment comments if using headers and why you're using headers and how using headers. So it would be also interesting to know how do people use it.
SPEAKER_04Maybe on a future episode. If it's good. So make it make it good.
SPEAKER_01We all we love stories. Okay, and I think last question for this episode uh this is a question for our previous episode where we talked about an Italian named Scarface. Scarface? Scarface. Isn't Scarface 48? Yeah, it's like in I think there's the Disney thing with Lions.
SPEAKER_04Yeah, that's Scar. Scar Scar. And then there was a film called Scarface. Got it.
SPEAKER_00There was a dark night where there was a there was a character who was asking, Do you want to know how I get scars? Ah yes.
SPEAKER_01We do not know want to know how to get scars. But but yeah, this is a family-friendly show. We do not want to talk about scars here. But Scarface has a question about storing data in Kafka forever. Because we keep saying that, you know, we give a lot of examples saying that it's okay or it's a good idea, and why not really? And then he says, well, it's a good idea, but it's expensive. At work, I have to use RAID disks, which are expensive.
SPEAKER_03That's a way to spend money.
SPEAKER_01And then he has two concerns. First about just storing data forever, it's not cheap. But also, if you want to use something like Kafka Streams and go and process historical data, Kafka Streams itself has its own state, which also uses storage. So there's like somehow compounded concerns. Do you have any thoughts about that?
SPEAKER_04Yes. Yes. There's a um I think there is an implied proposition that's not stated that we need to state. And he says, you know, well, if if it's so expensive to store data in Kafka, then we're going to have to use Spark to do our data processing. The implication there is that the data is going to live when I'm doing Spark. I'm going to put the data somewhere that's cheaper, right? And that's, of course, HDFS, which is always the silent partner of Spark.
SPEAKER_01Silent partner?
SPEAKER_04We never say HDFS. We say Spark because it sounds cool.
SPEAKER_01HDFS is kind of dead. The cool kids are doing S3 these days. Aaron Ross Powell So there you go. Okay.
SPEAKER_04So um somewhere cheaper. Yes. Right?
SPEAKER_01And maybe for Cassandra. Is Cassandra cheaper than Cassandra?
SPEAKER_04It's absolutely not in terms of storage. It's it's byte for byte. So HDFS and S3 are cheap storage. Yeah. Cassandra wouldn't even come up here. Right. So you're doing uh you're doing Spark over just something cheap. And if you're using RAID in your Kafka cluster, well then that's not cheap, right? And that's uh those other things, whether it's it's a legacy HDFS thing or it's S3, those are already getting the replication done for you, right? So you don't think, oh no, this needs to be RAID. It's just a file system that uh replicates. Well, folks, uh so does Kafka, right? So one of the things here is that running Kafka on RAID is probably a more expensive way to run Kafka than you need to.
SPEAKER_01Aaron Ross Powell Yeah. And it used to be that you pretty much needed that because Kafka broker would not survive the loss of a disk, but that's no longer true. Starting as I think 1.1. So definitely all of you move to JBOD and run Kafka on the cheap. Aaron Powell On the cheap, yeah.
SPEAKER_04So that kind of that kind of gets you out of that. Now uh more broadly, is storing data in Kafka over the long term gonna cost more than S3 buckets? Yes, right, because there's there's just more going on in a Kafka broker than uh a file system, a distributed file system. So there's definitely a cost differential there, and long term you're gonna have higher costs, but let's not make them worse, right? And there are interesting things going on in the community, uh thinking about like what happens. Is there a cold storage option? And and uh what what how does this really work when your data is supposed to live in Kafka forever or forever, whatever that means?
SPEAKER_00Here's an interesting uh thought for you guys. Um we can say you not have to have all data stored in Kafka forever. So because every topic has its own retention policy, so some of the data can be you know short-lived. You don't need to store, for example, uh for internity some of the intermediate topics where you do some processing.
SPEAKER_01But so my business does require a lot of long-term analysis. For example, is my marketing campaign this winter better than my marketing campaign five years ago? So do if I don't store it on Kafka, I have to store it somewhere, and now I have to use both Kafka streams and Spark, and I really don't want to use Spark.
SPEAKER_04And you don't. You don't want to do both. Yeah. So I mean if if so be thoughtful about that. The question seems to imply that there is historical data for regulatory reasons.
SPEAKER_01If you don't need historical data, your life is way easier.
SPEAKER_04Trevor Burrus, Jr. Yeah, yeah. And certainly and Gwen, you had mentioned uh Kafka streams and its its storage requirements uh in that that are that are not long-term, but uh you know persisting its uh the state store stuff onto disk.
SPEAKER_01And that can definitely be on cheap disks.
SPEAKER_04Please don't do that on RAID. Yes. If that's if that's on a RAID volume, then um that's just money you don't need to spend. Right. So there are some things that that kind of blunt the force of the objection a little bit, that that can make things a little cheaper, and long-term uh certainly things community is looking at.
SPEAKER_01Aaron Ross Powell And I've even seen people use a different cluster for the streams state stores than they do for the rest of the data, like the actual inputs and outputs.
SPEAKER_00And uh let's not forget uh uh Kafka Connect. So you s you can use the best of the both worlds if you your organization might use and HDFS, let's face it, the main organization use it. Um and you need to use historical data. You not need to use historical data every day. If you do writing reports, you can have have this as a part of your pipeline. You can you know feed data from HDFS using connect uh do your processing, you have a few topics with uh 2016, 2015, 2015.
SPEAKER_01I think that's a fantastic idea.
SPEAKER_00And after that you run this processing uh in the Kafka streams or Spark.
SPEAKER_01And when you're done, you can drop them and exactly.
SPEAKER_00Yeah, and you can also like uh using uh connector to offload this back to HDFS, so it will wait until next year when you need to do this book.
SPEAKER_01Fantastic suggestion. I think oh I thought we were just one one one more.
SPEAKER_04The the Spark versus Kafka thing, like there there is a place in the world for stata analysis of static data, for like experimental analysis of static data. That's that's not real time. No, you don't even think so?
SPEAKER_01I mean in my ideal world, uh streaming can do things across a lot of time horizons. And uh it's good to be like as a developer, it's I think that if I build expertise on one tool, it's kind of good to use it for all its worth, right? Like Spark is not easy, Kafka streams is not 100% trivia, like let's let's do one of them.
SPEAKER_04Yeah.
SPEAKER_01And where I was going with that was capable of, right?
SPEAKER_04Trevor Burrus If it's if it's a real-time system, an event-driven system you're trying to build, um like maybe there's a place for this analysis of static data where data scientists can experiment in Python and things like that. But uh that's never going to be an event-driven system. Like you can't make Spark into that, even in our opinion, with Spark streaming. So it's it's the should I use Kafka or Spark? Um I'm kind of trying starting to see that as a really wrong question.
SPEAKER_01It's Lambda architecture versus Kappa architecture in a way. Yeah. Lambda says you need batch and streaming separate. Yeah. Kappa says streaming is a superset and or includes a batch. And I'm kind of a Kappa architecture counselor.
SPEAKER_04You are. And if in fact streaming is what you need to do, then it's not really a question, right? At that point, it's Kafka and Kafka streams.
SPEAKER_01So I think we're good here. We answered a bunch of questions. We had very interesting debates for sure. That was fun. Great questions from the community, everyone. Thank you. Don't forget to subscribe to the channel and keep an eye out for Victor at the nearest conference. I heard that he would really appreciate it if you come up to him and say, you're so good in that podcast and you're so good in that talk. When can I see you again? And he will do a selfie with you.
SPEAKER_00He will absolutely do a question what I'm known for doing selfie. Um thanks for having me here. Extremely excited. Uh, hopefully, not last time. I really uh hope that you guys will like this episode, and this is why you tell, hey, bring Victor more. He's awesome. Thanks.
SPEAKER_01You are awesome, and I'll gladly have you back any day. Thanks, Gwen.
SPEAKER_04And there you have it. I hope that was helpful to you. If you've got questions, you can ask me at T L Burgland on Twitter, that's T-L-B-E-R-G-L-U-N-D. Or you can ask Gwen at GwenShap, that's G-W-E-N-S-H-A-P, or you can leave a comment on any of our YouTube videos. Your question might be featured on the next episode of Streaming Audio. And feel free to subscribe to our YouTube channel and this podcast wherever fine podcasts are sold. And if you subscribe through iTunes, be sure to leave us a review there. That helps other people discover the podcast and just generally helps us get the word out. We appreciate your support. See you next time.