Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov
Hi, we’re Tim Berglund, Adi Polak, and Viktor Gamov and we’re excited to bring you the Confluent Developer podcast (formerly “Streaming Audio.”) Our hand-crafted weekly episodes feature in-depth interviews with our community of software developers (actual human beings - not AI) talking about some of the most interesting challenges they’ve faced in their careers. We aim to explore the conditions that gave rise to each person’s technical hurdles, as well as how their experiences transformed their understanding and approach to building systems.
Whether you’re a seasoned open source data streaming engineer, or just someone who’s interested in learning more about Apache Kafka®, Apache Flink® and real-time data, we hope you’ll appreciate the stories, the discussion, and our effort to bring you a high-quality show worth your time.
Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov
Ask Confluent #10: Cooperative Rebalances for Kafka Connect ft. Konstantine Karantasis
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
Want to know how Kafka Connect distributes tasks to workers? Always thought Connect rebalances could be improved? In this episode of Ask Confluent, Gwen Shapira speaks with Konstantine Karantasis, software engineer at Confluent, about the latest improvements to Kafka Connect and how to run the Confluent CLI on Windows.
EPISODE LINKS
- Improved rebalancing for Kafka Connect
- Improved rebalancing for Kafka Streams
- The "what would Kafka do?" scenario from Mark Papadakis
- The future of retail at Nordstrom
- Watch the video version of this podcast
SEASON 2
Hosted by Tim Berglund, Adi Polak and Viktor Gamov
Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
Music by Coastal Kites
Artwork by Phil Vo
- 🎧 Subscribe to Confluent Developer wherever you listen to podcasts.
- ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
- 👍 If you enjoyed this, please leave us a rating.
- 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
Welcome back everyone. I'm Tim Burkle. In this episode of Ask Confluent, my co-host Gwen Shapira speaks with Confluent Engineer Constantine Karantov about the latest improvements to Kafka Connect and how to run the Confluent CLI on Windows.
SPEAKER_02Hi everyone, welcome back to Ask Confluent, where we answer questions from the internet. I'm your host, Gwen Shapira, and this is Ask Confluent episode number 10. So yay for 10th anniversary! And with me today special guest Constantin. So Constantin and I have been working together for like ages. And he's an engineer on the Connect Team, and he first joined the team to work on the SS3 connector, which immediately became this huge success. And then he wrote the Confluent CLI, which immediately went on to become a huge success. And then he wrote, I think, the MQTT connector, which is on the proxy, right? Which is on the way to become a huge success. And now he's working on really deep future improvements to the Connect framework that we're going to chat about. So it's really an honor to have you here on the show. I'm super excited about that.
SPEAKER_01Great to be here, Gwen. I would like to be on your show and uh discuss uh interesting questions from the internet and our users, of course.
SPEAKER_02Yes, and this will be a busy episode. We have quite a big uh question backlog. So let's get to it. Okay, so first question was a response. A few episodes back I had Matthias from the streams team here on the podcast, and he uh discussed how streams rebalancing work and how failover work and what happens if I have a streams job and I make a change, I add something, I remove something. And Renato Mendes Figureto Mendes Figuredo said that this is good information. I wonder if the whole rebalance strategy from K-Streams also applies to Cafo Connect. I've been experiencing long delays after deploying a new version of our containers, which causes them to be recycled. Thanks. So is the whole rebalance strategy and streams the same in Connect?
SPEAKER_01That's that's a very interesting question, and yeah, they're very related. Well, um uh we I would say we see the problem from a different uh perspective, a different angle. And uh I believe we will arrive at a common solution uh very soon, combining uh different approaches. So uh what what has been going on in Kafka Consumer and the streams is that um there uh uh we wanted to avoid um rebalancing in most cases. It's uh it's expensive, you have to restore uh state. So in Connect we thought of uh seeing uh the problem differently a little bit, and what we are trying to do is make rebalancing itself much more lightweight and uh less disruptive for the the group.
SPEAKER_02So basically in Connect you don't like every worker doesn't have that much state, right? We just read and write from Kafka. There's some configuration, but not that much state.
SPEAKER_01Yeah, exactly.
SPEAKER_02Uh so what do rebalances really do in that case?
SPEAKER_01Oh yeah, it's very interesting. Um so as you said, uh connect workers are uh locally stateless in a sense. The only state they save it's either in Kafka and the third-party systems, of course. Uh so uh we are we are in a good uh situation there. What happens is that um rebalancing is useful for connect because what you uh distribute and exchange is the number of tasks and connectors that will run in the Connect cluster.
SPEAKER_02So basically when he recycles his containers, he stops the container, and immediately the connect tasks that are running in that container shift over to another container, which causes a rebalance, which right now is pretty disruptive from what I understand.
SPEAKER_01Yeah, and and the reason that is disruptive right now is that we followed a simple strategy where it says that uh when you join, when a new worker joins, or even when a new connector starts, all the other members of the group uh release the resources. Which for connect it means that you stop the connectors and the tasks and you join uh the group uh with a fresh start, making a fresh start. Uh by doing so, you uh you have to stop the tasks and the connectors. So that's where we want to focus uh with this improvement.
SPEAKER_02So we're talking about KIP 415 that I think you created a few weeks back. It's still an ongoing discussion in the community, and this keep basically suggests a whole new strategy for rebalances in Connect.
SPEAKER_01Yeah, exactly. It's an improvement. Uh we are we are still uh uh polishing the details, we are still discussing it, but uh we have we have a good plan. Uh implementation has already started. We try to address the stop the world effect. Yes. Right now the workers won't give up uh the connectors that they don't have to stop, and uh either they will start only the new tasks or they will release only the connectors that need to be reassigned to different tasks.
SPEAKER_02Clarify right now is an engineering right now, which is after the keep will be approved and my patches will be merged and a release will happen. So it's not a real right now. It's a right now for like three months from now, I guess.
SPEAKER_01Uh yeah, I I couldn't guess exactly the time. Uh uh it would be uh tricky to do that, but uh it's it's in advanced development. So uh we have that uh uh actively working and uh yeah, we are confident we have it soon.
SPEAKER_02Yeah, I really like this thought about it, is that we're kind of I I love the name cooperative rebalancing. It's just it sounds more friendly, and you really don't ask every each worker to do more than he absolutely have to. So it's a kind of a more gentle method for uh balancing work. I really like that. I think I think it's a neat idea, yeah. I think the idea is neat, and it's interesting to see the discussion in the community as well. Like people ask very deep questions, and it's uh I personally learn a lot about different ways people approach their distributed systems by just reading how how they what what they ask when they read one of those proposals.
SPEAKER_01Yeah, classic, classic uh distributed system problem. Yeah, it's exciting.
SPEAKER_02Okay, on to the next one. The next uh comment is actually a compliment. So a while back Peter and I here recorded a short presentation about Kafka and the service mesh, which is kind of just me exploring some ideas and it was pretty new stuff at the time, and kind of like how do I approach the service mesh and how it's related to the work we are doing here. And Rob Gruhl said that this is a brilliant talk and I'm ahead of the curve, which is amazing to me because maybe you don't know who Rob Gruhl is, but he's managing uh innovative infrastructure at um Nordstorm, and he published um GitHub repository and a bunch of presentations and blogs about the future of retail architecture, and it has amazing ideas about cloud and serverless and service meshes and events driven and Kafka, and it's all really, really advanced, really, really smart. I really love Rob Gould's work, and I'm still planning to fly over to Seattle to meet him in person. So I'm a huge fangirl, and getting a compliment from him is just a huge moment for me, and I'm super happy. So thank you, you made my day. Next question is a response to our intro to Stream's uh video. And Jeff G says, so Stream's API is only for Java clients. How about .NET, Rust, Python, all those popular languages? What about them? And I think right now our answer is you know, KSQL works in with any client in any language, and it is stream processing API, and SQL is a fairly popular language. So I think that's yeah the main way to go about it. Or you can write your own. It's all open source. That would be welcome too. Yeah, of course. I want to do streams APIs in Rust. That would be a good excuse to go. Uh I know that we a bunch of us are applying to be in a Rust conference uh in Portland. Oh, really? So uh yeah, maybe that that will be our excuse to go.
SPEAKER_01Yeah.
SPEAKER_02Okay, next question. Actually, a comment, maybe a compliment. It's very, very hard to tell. Our ever-popular video, KSQL Introduction, level up your KSQL that team did, I don't know how long ago. I think it's one of our best sellers. There is not a single episode where I don't get some kind of comment or question about that specific video. And SD said Gavin Belson left Juli and now works at Confluent. And everyone in our team responded by thinking it's a super funny comment and it's hilarious. And I don't know what they meant, but I heard that maybe you know what they meant.
SPEAKER_01You're lucky, Gwen, because I'm a fan of the show and uh watched it. So um, yeah, uh I might be able to take a guess here what uh they mean.
SPEAKER_02Uh that's it's from the Silicon Valley show, right?
SPEAKER_01Yeah, it's gl exactly, it's from that show, and uh Gavin Berson is uh this uh Arhedipal uh bad guy in the show. Yeah, he goes after uh small startups, he's uh leading a big corporation that is not innovating. We don't do something, yeah. No, we are definitely exactly he would chase us uh in that sense. Um so replicating their technology, trying to acquire them. So, yeah, but he wasn't always like that. Oh yeah. Um he uh at the very beginning, the early days, he was an innovator himself. Uh he created uh a new technology that led to uh this um success with Juli. So uh what I understand here, and maybe our our users uh uh had in mind is that he returns from Tibet and he feels nostalgic about these early days where uh he he was working in something new and up and coming.
SPEAKER_02And uh yeah, I feel that this common means So maybe he joined Confluent to work on the next big thing. Exactly like we all did.
SPEAKER_01He feels like he has to leave the dark side and uh join uh yeah, something like that.
SPEAKER_02But we'd never really accept him because we don't accept evil people. We have a high value. We have confluent values. Our engineers have to be nice and humble and respectful. We don't do that. So yeah, sorry, Gavin Belsom, you are not accepted at Confluence.
SPEAKER_01He was rejected, yeah.
SPEAKER_02Yes. Next question. A while back, we had a screen cast about introducing the Confluent CLI, which you happen to write. And Jawad Khan had a question: How can I use the CLI on Windows? So you did the unimaginable, you went places where no engineering confluent went before.
SPEAKER_01That's uh yeah, that's that's interesting to do. Uh well, I also uh follow your show here, and uh, that this question has come up before. So I went ahead, I provisioned a Windows VM on AWS with uh Windows Server 2019. Uh and uh yeah, follow follow the steps to install um uh Windows subsystem for Linux. Uh that's the first step if you want to uh have uh a command line shell uh like bus uh at your availability. So that's what you need, the first step. And then you have to install Java and the CallFlow platform, of course. Uh once you do that, you are able to uh run Confluent CLI and start uh the Confluent uh services on Windows.
SPEAKER_02We have proof right here that it actually works, and we have Zookeeper starting and Kafka starting, and I see you didn't push your luck, so we don't see connect here, but we can try.
SPEAKER_01Yeah, I I wanted to leave something for our users to explore and uh discover by themselves.
SPEAKER_02Nice one, yes, and uh amazingly it seems to be working. So thank you so much for resolving this mystery for us. And I have to say that with the subsystem installed, it looks almost like a normal operating system. Almost, not quite. Okay, next question I have from Sachin Deukar, who we already I think responded to one of his questions on the podcast, so it's a repeat uh visitor, and he said the question for the podcast Can you provide us with few popular use cases for headers in Kafka and what it should and what it shouldn't be used for? And I feel like this question comes up once every few episodes. It seems like people keep looking for header use cases. And I heard you had a few.
SPEAKER_01Uh yeah, I have I have a few from uh some products that we've developed, and uh, of course, our users uh use them and uh they found them useful. So uh two use cases that come uh to mind is uh uh how we use headers in MQTT proxy. Uh basically it uh worked well for us because we isolated uh the payload at the Kafka record value, uh the MQTT topic uh at the uh as a key, and then uh metadata such as the quality of service uh level, we put them in headers. And that's the the protocol where uh that is being used by devices like uh a car, a scooter that you can rent, or uh a thermostat at your home uh in order to publish data in customers.
SPEAKER_02So you use headers to basically match what a different protocol would do.
SPEAKER_01Yeah, exactly. Uh uh embed all the metadata of the protocol in headers. Uh it was better encapsulation of the data which I love that.
SPEAKER_02It and that's kind of a very engineer architecture-driven approach. You find good encapsulation and what really belongs in what portion of your payload.
SPEAKER_01Yeah, definitely. We wouldn't like our users to have to unpack uh the value from uh the record uh value and then do some deserialization in order to discover again the for instance the quality of service uh level. So that that uh worked out really well for uh for us with headers. And um another example that comes to mind is um uh with uh confident replicator actually. Uh we use headers there uh in order to uh encode the origin of a specific record and basically do what is uh a classic, a classic application of headers, uh do routing of records. So based on the origin, uh a record might be uh replicated uh in a data center or not. So we avoid uh having cycles, uh infinite cycles of replication between uh data centers, and uh that's how yeah, that's how we use headers in a replicator.
SPEAKER_02Yeah, and I think there's a general pattern here is where you have metadata and you don't want the you have some consumers who only need the metadata. So for example, the replicator it doesn't care about the data itself, it just copies the data from place to place. And the data can be in Avo or JSON or Srift or a format that I invented myself that is hyper-optimized. But the replicator doesn't need to care. But there is bits of metadata that the replicator needs to care about, and you put this metadata in the header so the replicator will be able to access this metadata without having to know anything about the data format that uh I'm using in my messages versus the ones that you're using in your messages, and the replicator has to deal with all of them. And I think that's like the generic pattern that I've been seeing for headers all the time. Metadata that has to be used by consumers that don't always understand all the different serialization formats in the company. So it's things like um timing information, uh lineage information. You if you have an audit process that tracks lineage information, it will have to deal with every data format in your company. It had the you cannot put the information in the payload, you have to put it in the header.
SPEAKER_01Yeah, that's a that's a great another great example that uh it's not the first to come in mind when you build a system, but definitely it's very useful.
SPEAKER_02No, it's not the first for you, but if you worked in a bank, what the auditor would do would be your top of the mind.
SPEAKER_01Oh yeah, for sure.
SPEAKER_02Okay, and another question, that's a new one arriving yesterday on Twitter, but I had to include it because it's uh really good in two different ways. First of all, the person asking the question is Mark Papadakis, who A happened to be a Greek national.
SPEAKER_01Yeah, and you pronounce his name uh really well.
SPEAKER_02He managed to do it, oh my god, yay. And he's also a good friend of mine, and he's CTO of some kind of network uh company, and he's writing a Kafka clone in C. It's called Tank. So he's like a pretty big distributed systems expert. And uh at some point he ran into a fairly complex scenario. We have a producer doing this, a producer doing that, we have five replicas, one of the one of the topics has min ISR or sorry, yeah, I think they have different min ISRs, and then some replicas are online, some replicas are offline. It was a pretty intense scenario. And then he asked, okay, what will happen in that scenario? And it took me a while to figure out, and I uh responded with a big story, and then you know, if something is offline, then what happens if it comes back online and what the timing of the things? So it was a really interesting question to think about. Also, the first time that I got a question via GIST, we already had questions uh via, let's see, Stack Overflow, uh, YouTube, Twitter, uh, Slack. Uh so but GIST from by GitHub is a new one, but I thought it's pretty good.
SPEAKER_01It was a big story, so I think it was always nice to have uh some code along with your questions, right? There was no code there. Okay, it's a good idea.
SPEAKER_02It was a scenario scenario. It was an abstract scenario. Uh but it was a pretty interesting. It's always like it's funny how when you are trying to learn about the distributed system, you keep coming up with those scenarios and trying to reason over what would happen in if we have this network partition and then it heals and then we have another network partition. You know, you come up with these things. But after you really understand the protocol, it's very easy to think through and answer all of them. That's how you kind of know that you understand. So I really enjoy thinking through those types of scenarios. That's I guess my way of geeking out.
SPEAKER_01Yeah, it's super interesting.
SPEAKER_02So yeah, I invite everyone to invent their own their own horrible scenarios and we'll try to figure out what would Kafka do. Okay, and last we'll finish with another compliment, comp comment for my uh last as confluent, uh free music 54, which almost feels like someone has a sock puppet of some sort. Uh they said, nice video, you can learn so much from those videos, amazing. And yes, you can learn a lot from those videos. And if you want to keep learning, you can subscribe to our channel and keep learning a lot from all our videos. Uh, we try to make them fun and educational. So that was a lot of fun. Thanks for joining us, uh Constantine.
SPEAKER_01Great, great pleasure. Yeah, nice to be here. And uh good luck with the rest of it. Uh I bet there's gonna be many more episodes.
SPEAKER_02Yes, and great questions from the community as usual. I learned a ton and thanks for KIP415. I hope it will get voted in soon.
SPEAKER_01Yeah, we're getting there. Thank you.
SPEAKER_00And there you have it. I hope that was helpful to you. If you've got questions, you can ask me at at T L Burkland on Twitter. That's T-L-P-E-R-T-L-U-N-D. Or you can ask Gwen at GwenCap. That's T W E N F. Or you can leave a comment on any of our YouTube videos. And feel free to subscribe to our YouTube channel at the podcast or Five Podcasts. And if you subscribe to our four podcasts,