Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov

Scaling Apache Kafka with Todd Palino

Confluent, original creators of Apache Kafka® Season 1 Episode 57

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 46:03

Todd Palino, a senior SRE at LinkedIn, talks about the start of Apache Kafka® at LinkedIn, what learning to use Kafka was like, how Kafka has changed, and what he and others in the community hope for in the future of Kafka. If you’re curious about life as an SRE, Todd shares the details on that too, and goes into how Kafka is used at LinkedIn, as well as several wins and challenges over the years with the product. 

EPISODE LINKS

SEASON 2
Hosted by Tim Berglund, Adi Polak and Viktor Gamov
Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
Music by Coastal Kites 
Artwork by Phil Vo 

  •  🎧 Subscribe to Confluent Developer wherever you listen to podcasts. 
  • ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
  • 👍 If you enjoyed this, please leave us a rating. 
  • 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
SPEAKER_00

When you want to know where something is going, it's good to talk to somebody who knows where it's been. Todd Polino has been involved with Kafka since 2013, so he definitely qualifies. He's a senior staff site reliability engineer at LinkedIn, which was ground zero for Kafka development at that point in time. I got to talk to him today about what it was like back then, how it's changed, how he thinks about operating Kafka after all these years on the job, and where he thinks the project should go in the future. It's all there on today's episode of Streaming Audio, a podcast about Kafka, Confluent, and the cloud. Maybe he needs a little bit of introduction, but I'm going to introduce him anyway. Todd Polino, Todd, the senior staff SRE at LinkedIn. Welcome to the show.

SPEAKER_01

Thanks very much, Tim. Happy to be here today.

SPEAKER_00

You bet. And I feel like I feel like introducing you by your current job title at LinkedIn does not adequately capture your contribution and your presence in the Kafka community because you you have your guy who's been around for a little bit. And that's always the most interesting thing on my mind when I'm talking to early Kafka contributors and operators and things is, you know, when did you get your start? What, how did you get your start? What was it like then compared to now? Give us some give us some oral history.

SPEAKER_01

Yeah, so I started on Kafka when I started at LinkedIn in December of 2013. Uh so this was pretty early on. Uh we were just making the transition from 07 to 08. We were kind of on the tail end of that transition at that point. The funny thing is, when I started, I had absolutely no idea what Apache Kafka was. I was offered the job as an SRE and given the topic that I'd be working with. And I had to do research to get started. But within uh within six months after that, I was giving talks at ApacheCon with my colleague Clark and talking about how to run Kafka because nobody knew how to run Kafka at the time. We were all just figuring it out. At LinkedIn, of course, we had the benefit of working right there with Jay, Neha, and June at the start, as well as Joel Kashi and Gojong. So we had a lot of dev help to get it going, uh, but really they knew how to write the code and they wrote really good code, but running it was still very difficult.

SPEAKER_00

Okay. Uh what was hard about it back then?

SPEAKER_01

I guess so when you think about the 07 to 08 transition, there was no transition. That was um pick up and move everything from one cluster to another because you couldn't do an upgrade at all.

SPEAKER_00

Those upgrade release notes are easy to write, though.

SPEAKER_01

Yeah, exactly. And you don't really have to worry about backwards compatibility because there's no forwards compatibility either.

SPEAKER_00

They're easy to write, they're not easy to deliver in person, but they're you know, super easy to document. Yes. That was that was and and that was consistent with the kind of Wild West spirit of the time, right?

SPEAKER_01

Exactly, exactly. So I mean we were LinkedIn at the time was running on trunk. Uh we weren't even running release versions because we were actively writing the code as quickly as we could, knocking out the bugs as fast as we could. And so we were always running on the absolute latest code stack. That's it's frightening to think about how much production data we put through Kafka, and we were doing it at the time on completely untested code.

SPEAKER_00

Wow. And it was still uh you know, from the beginning, it was critical operational infrastructure, right? This wasn't like some slide project.

SPEAKER_01

No, no. I mean, from the start, we were running uh tracking data for the website through it. We were running metrics. I mean, when I started, not only was was the metrics for every single application at LinkedIn running through Kafka, but the metrics for Kafka itself were running through Kafka to get to the alerting system.

SPEAKER_00

Oh, I like it.

SPEAKER_01

Yeah. Talk about talk about you know Wild West cowboys.

SPEAKER_00

Uh yeah, yeah. Uh so I and I don't I it's it's hard, I guess it's easy to say, well, you know, five years ago uh life was simpler then. You know, it wasn't. Okay. It was LinkedIn was a major internet company. There was significant traffic, all of those things you described, there are significant parts of the business hinging on them, and people who, in order to do their jobs, need that data to be available and everything. So it wasn't it wasn't like it was 2002, and you know, there were still groups of buddies from high school running ISPs or something. It just it 2013 was the internet economy was sophisticated. So that was a big deal. But um is it still like that? I I don't know to what degree you can talk about operational details at at LinkedIn, but um is uh you though you you're sort of implying that you ran on trunk then and you don't know.

SPEAKER_01

Yeah, the uh the cadence of updates to open source Kafka has really picked up over the years. So it's impossible for anyone to keep up with those at this point and and stay on a trunk version. Uh, there's just so much engagement in the larger community with building Kafka. But even now at LinkedIn, Kafka is central to everything we do, probably more so than even then. When I started, we were running tracking data and metrics over Kafka. Uh, we've now got most of the queuing within uh LinkedIn for any given application running over Kafka. Our deployment systems depend on being able to stream messages through Kafka to other components. We have developed change capture systems for our databases, and some of our larger database systems like Espresso, which is an internal key value store, now does all of its replication via Kafka as well, whereas that wasn't the case. So, whereas back in 2013, you know, we had a couple of sites and we had a little bit of data running through Kafka. It took us to 2016 to hit uh a trillion messages a day. But at this point, we are over five trillion messages a day easily, and you can just see how that that hockey stick has built up just in the last couple of years.

SPEAKER_00

Yeah. Oh man, that's uh that's remarkable. You guys are for for in terms of scale, uh, you're certainly one of the flagship users. But you say it's Kafka itself is changing, you know, master or trunk is just changing too quickly to be able to responsibly engage in that kind of behavior.

SPEAKER_01

Yeah, definitely. Um, whereas whereas it was a lot easier when LinkedIn was the primary contributor to the project, obviously. All the developers sat in the same room. If there was a problem, we could work it out very quickly. So it's not too difficult to run on the latest code base at that point. But now we have hundreds, if not thousands, of contributors from all over the world who are contributing bits and pieces all over the place. You've got not only core Kafka, but you've got the Kafka Connect and Kafka Streams stacks that are all part of the same project. And everything's just shifting at such a constant rate. It's difficult to keep up with the kips coming across the mailing list now for updates, the large updates that are going through.

SPEAKER_00

Tell me about it. Yes, it is. Uh you're you're very right. And you know, sometimes my job is is to just summarize the important ones on a major release, but it's still it it if there are things that you do in your life that are not just keeping up with those kips, yeah, it's hard to do. That's that's very true. What um okay, so it it's changing faster, there's a broader community contributing to it. Um all the core or the the majority of the core committers are not sitting in a room, also being LinkedIn developers. So all of those things that all make sense. Um what else has changed? You know, like what what you said it was hard to operate back then. It's nobody accuses Kafka of being easy to operate now. So has it gotten easier or do you think as the the ecosystem has has kind of broadened its mandate, has gotten harder? What's the deal?

SPEAKER_01

I think it's certainly gotten a lot easier to run Kafka. Um we see a lot more companies who are doing it now and doing it comfortably and running production data over Kafka and not worrying about it. Um, and not just companies who are buying services from a provider such as Confluent or any of the others out there who are selling Kafka or services. Um, because there are tools that are available now. Um, myself, my colleagues at LinkedIn, um, so many other people who I've had the opportunity to engage with at Kafka Summit and elsewhere are writing lots of good tools that help make running Kafka easier. That's that's always been my goal in the community is to try to make running Kafka easier for everyone else, whether that's through documentation tools or just talking a lot about it.

SPEAKER_00

Right. Which is a thing that you do. Uh and speaking of Kafka Summit, we've been uh uh delighted to have you there and sharing your experiences in the past. So that's uh that's a good thing. Thank you for thank you for being one of the people who tries to make it more pleasant. Um my uh the focus of my work in the developer relations team here at Confluent um is often more around application development. Uh you know, deployment and and operations are 100% within our mandate, so that that kind of thing happens. Uh like we've spent a lot of time talking about how to run Kafka on Kubernetes recently, things like that. But if you look at like me personally and the the advocacy that I do, I'm always thinking of uh how how can application developers have a better time with this? And people like you who are thinking primarily of uh you know you being an SRE, like how can my fellow SREs uh love their lives a little bit more? Uh I just really appreciate you being on that job. So I don't know if I've ever had a chance to thank you for that. So thank you, Todd.

SPEAKER_01

Well, thank you very much. I mean, uh I've enjoyed doing it. I have I always enjoy um giving a talk and being able to share a little bit of the experience. And it's been interesting over the years to find out where I'm wrong and spend time correcting myself because so many people have um given me the benefit of listening to me about how to do this. So when I'm when I'm wrong, it uh it then takes me a lot of time to explain that.

SPEAKER_00

Right. Uh can you give me an example?

SPEAKER_01

Absolutely. Uh we could talk about under-replicated partitions because that's my favorite uh favorite topic on which I have been very wrong on. Uh so LinkedIn still monitors uh underreplicated partitions as a primary metric. And I have said for had said for years that that was the best primary metric to look at your Kafka cluster through because it could tell you so many things about what was going on. And what I began to realize as um, especially as we were getting into a space a couple years ago where we had an alert overload in our system, is that it's a very noisy alert. It's often ignored, and it's rarely telling me anything useful that I can't get from another signal. So I started talking about how this was actually a horrible thing to monitor and explained these reasons why, and talked about some different ways of monitoring the system, like SLO-based monitoring. But it's funny because um Gwen Shapira, one of your own, would disagree with me on the benefits of using under-replicated partitions as a primary signal. Until very recently, I gave a talk in May in Poland uh talking once again about under-replicated partitions and other metrics. And afterwards I was talking with Gwen and she said, now that she's responsible for a large deployment of Kafka, she now agrees with me.

SPEAKER_00

Okay, so you uh this is not a time when you had to admit that you were wrong.

SPEAKER_01

Well, I didn't have to admit I was wrong again to Gwen. We we disagreed for a while. I insisted I was right, and she's kind of come around, but it's not entirely um, it's not an entirely bad metric. You still want to monitor it, you still want to use it. All I say is you don't want to alert on it. It's not something that you should should wake you up at night.

SPEAKER_00

Okay. Um well, is there a point? I guess uh this is a detailed question, but I mean there's a point for a particular topic where it can have, you know, you can have enough partitions underreplicated that that's a problem. Like at a at a topic level, that's an alert, right?

SPEAKER_01

Well, not necessarily, because your alerts should generally focus on what your customers care about. Um there's uh there's a lot of talk lately about SLO-based monitoring, and I've given my own variant on on that talk as well.

SPEAKER_00

It's service level objective-based monitoring, right?

SPEAKER_01

Yes, service-level objective. Thank you for expanding that. We can go into a whole tirade about the various service level uh definitions of what they do.

SPEAKER_00

I invite uh a little bit of tirade there, so please bring it.

SPEAKER_01

Um don't confuse your SLOs and your SLAs. SLOs are generally what most of us have. SLAs, uh service level agreements are SLOs with penalties, like having to pay someone money. Um you guys probably have SLAs with your customers, but most of us who run Kafka internally don't. Our database guys aren't going to come and demand a case of beer because Kafka had 98% availability.

SPEAKER_00

Of course. But a serviceable objective is an SLA with uh some sort of consideration attached to it.

SPEAKER_01

Yes. But that's kind of off the uh off the subject a little bit. Um so you know, we all we all give talks about SLO-based monitoring now. And what it comes down to is if you know about problems and you know how to deal with them, you should automate the responses. And therefore, you don't really need alerts. You just need the automation of how to remediate whatever the alert was. What you actually want to alert on is are the things that you don't know about, the problems that you don't know that you have. And the only way to detect those is to monitor your user experience, which is generally a service-level objective that you're going to monitor, something around availability or latency or correctness, um, some combination of those. So um, when you're monitoring SLOs, you can then you always know what your user's doing. So your user doesn't care generally about under-replicated partitions. Now, as you were noting, if it gets to be, if your under-replicated count is so high that you actually have partitions offline, then yes, your user cares about that. But you'll see that through monitoring the availability from your user's point of view.

SPEAKER_00

Your user doesn't directly care that there's a partition offline. Uh there that's a that's a a fairly remote cause, and there's a more proximate cause of the monitor, like, you know, uh, I can't create a new account or status page is unavailable or something like that.

SPEAKER_01

Exactly. So you want to alert on that piece. The under-replicated partitions itself is something that you want to monitor, and you'll definitely want for debugging problems. And you may even want to have some sort of daytime alert on it, but it's gonna be noisy because there are situations where you're naturally going to have under-replicated partitions, such as as if you're rebalancing data within a cluster and moving things around. Your partitions are just going to be under-replicated, that's by definition. So you have to be careful about monitoring those sorts of things.

SPEAKER_00

Okay, that makes a lot of sense. And that's uh a move that you've made personally in your in your thinking as an SRE.

SPEAKER_01

Absolutely. Absolutely. And it would be great if I could convince all of my colleagues to do it, but we're getting there gradually. It takes a while to define SLOs and and get them working properly.

SPEAKER_00

Sure. Um, I mean it it took uh I already said I'm I'm not myself uh any sort of leading ops thinker. Nobody comes to me for my opinion, and they shouldn't. Uh, you know, there are other people who are good at that, but uh you had to convince me. You first said that. I'm like, wait a second, of course you want to know that. Um I guess in my defense, it didn't take long to talk me into it.

SPEAKER_01

It it's a discussion that I've got very practiced at because you get that from ops people too. Anyone who's been monitoring that way will very, very much disagree with you. Uh, and so you get practiced at what the explanation is.

SPEAKER_00

Yeah, so but what what the the mind of the ops professional becomes, though, is this um kind of classifier of all of these noisy signals that come in. And you know, you you take those noisy signals as inputs and determine when you know there's there's really cause for alarm. Uh or you just take action, you just take expensive personal action to look at something that you didn't really need to look at because it didn't it didn't impact the customer, and then you've got a person whose life is less pleasant, right? That SRE has to get woken up in the in the middle of the night more often.

SPEAKER_01

Exactly. I mean, I like to say that SREs are professionally lazy.

SPEAKER_00

Yeah, they should be, yeah, right?

SPEAKER_01

Yeah, it's our job to automate the problems away or make sure they never happen in the first place, because you want the rest of the company to say, oh, why do we have these SREs anyways? The system's always up and running fine.

SPEAKER_00

Right. That's that's the goal. What do we what do we pay these expensive people for? It's uh there's there's no problem. It's uh maybe the opposite of the uh carrying garlic in your pocket and saying that's why there are no vampires. Um maybe you have to think about that. Maybe it's maybe it's the opposite, maybe it's like a legitimate version of that argument. I'm not sure, but I take your point. How about okay, so you you joined um back to kind of history and and future, and I I really I really appreciate being able to talk to you because you're a person with a lot more perspective, uh historical perspective on Kafka than most. Um and so your extrapolations about the future are are interesting on that basis. So you said you joined just before point eight, and point eight was this big disruptive uh upgrade where you you know there you had to build up, stand up a new customer and migrate or cluster and migrate data. And the client interface changed a lot in point eight, right? Am I getting that right?

SPEAKER_01

Yep. Uh the client interface changed. We added replication. Um there were lots of block changes at that point.

SPEAKER_00

Little little things like replication, which is so crazy to look back. That's I said five years before, but we're recording this in 2019. So that was um that was six years ago. To think that there was a production Kafka that didn't replicate. Um that seems strange to you as it does to me.

SPEAKER_01

Oh, it's insane. I mean, the the amount of data we were running through the system and how critical it was at the time, and there was no replication on at the Kafka layer. You lose a you lose lose a broker and you're done. Uh I wouldn't want to run any system like that.

SPEAKER_00

Yeah, no, you and you wouldn't probably. If someone proposed one, you probably would uh would reject it. So point eight brought us replication and the new clients, and then point nine, the big deal was connect, and point ten, the big deal was streams. Am I getting all those right?

SPEAKER_01

Uh I think that was about right. We haven't made much use of connect or streams internally at LinkedIn, so I lose track of exactly where they were.

SPEAKER_00

Gotcha, gotcha. Um, but what is your I mean, I I have a take on those things as being a part of Kafka, and and uh I'll just spill that first. I I I kind of I would love your your you don't use them a lot, but I'm sure you have an opinion. Um but I view those as the things that if if you're a developer, uh team of developers, and you're building interesting things on Kafka, um, you're gonna have to integrate with other systems. And so you're either gonna write a connect or use a connect. It's it's your choice, right? You can you can build one or uh use one that the community has built. Likewise with streams, and I this argument gets a little bit cutesy, right? I mean I realize it's more complicated than this, but um your consumers are the the code in your consumers is code that tends to grow hair. There are things that they do uh that they always do. Like they you you always end up wanting to aggregate and enrich and do these these things, and the the library gives you no help at all. So, you know, you're gonna extract a framework from the first ten consumers you write, and you're then gonna need to care for and feed that that stream processing framework for the rest of your life. So I view Connect and streams as these things that the community solved in uh you know in a in a consensus way because you were gonna do them anyway. But now I would love I love your take on that, particularly in light of the fact that you said you guys don't you don't do those things. So what do you think?

SPEAKER_01

Yeah, so within LinkedIn, we use uh Apache Samsung for stream processing. And then for the Kinect piece, we've used a variety of things over the years. Obviously, before Kinect existed, you grow your own. Uh, but now we have uh Brooklyn, which I believe is on the path to be open source soon as well, kind of started as a change data capture system and grew up into more of a generic connection framework. Um, so like many others, um, we have our own solutions for these problems. I see Kafka Connect and Kafka Streams as both very useful in their own right. Um, but I kind of wish they hadn't been part of the core Kafka uh project. I I wish that they were it was more of an ecosystem project where you had um pieces around Kafka that could update by themselves without having to constantly push the version of Kafka forwards faster and faster.

SPEAKER_00

Got it. And uh what what do you think? What are the good outcomes you're looking for looking at uh you know with that approach? Just the ability to have a more manageable pace of change in core Kafka?

SPEAKER_01

Yeah, it it really is about uh a more manageable uh pace in the the core product, where you have consuming and producing, and then you can layer other things on top of that. Uh right now the the Kafka project moves so quickly that you can't support very many versions. Uh and many of the versions that are coming out have major changes to connect or streams going in, which and they don't impact the core as much. So you end up with a case where all of a sudden you've released 1.0 and then 2.0 almost right on top of it, uh, it seems, as far as the time string goes. Uh and then you're getting to a place where you're rolling off support for the older versions faster than people can keep up with. You know, as we've said, we're talking about uh systems that run a significant amount of production data, and there's definitely a lot of reticence to upgrade those very quickly. Uh, if they're working well, then you don't really want to touch them uh as much as you can avoid.

SPEAKER_00

Right, right. Which is the classical ops answer, right? Upgrades are what break things. Yeah. And once you've got automation and monitoring and stability around a certain version, there isn't there isn't a prima fasci reason to change that, other than access to new features.

SPEAKER_01

Yeah, exactly. And there's there's a lot of change that's in in the core Kafka uh in the broker that's driven by connecting streams. And that's necessary, uh, message format changes and other things like that. But if the projects were separate, then those changes would have to be thought through a lot more than they do right now. And they would have to be slower paced. They would require a little more give and take in terms of timing to get them out, which would lend itself to a little more stability in the broker itself and with the with the uh protocol as well.

SPEAKER_00

Gotcha. I take your point. I take your point. I think uh there is no question that that's true if they were separate projects. Um now you have negotiation between uh more or less separate teams instead of uh, you know, obviously if you look at the community of people who contribute to Kafka, there are people who care more about Connect and more about Streams and more about Core. You know, those are all specialties, but there there is this sense of you know, it's one family and these are bedrooms at best, uh, rather than um, you know, one extended family and different houses on the block, and you kind of have to get up and put pants on and walk down to the other house uh to have a conversation rather than just shouting from the other room.

SPEAKER_01

Exactly, exactly. And you know, again, as you noted earlier, you uh you have a very application uh developer point of view on this, and I have a very operations point of view on this. Uh it's always about uh working together and trying to meet in the middle somewhere.

SPEAKER_00

Yeah, it totally is, it totally is. And the funny thing about that is you say that, and I mean, Todd, your your argument makes all the sense in the world, and I get where you're going. But when you say that and you say, well, we need to do this, so it'll slow down the pace of change. I mean, my immediate reaction is, what are you talking about? Slow down the pace. We can't slow down the pace. We need these features, you know. Uh so it's funny, uh, ladies and gentlemen in the audience, the classical uh developer and ops argument is playing out before your eyes. Uh I want things to change right away. Todd wants uh systems to run. Uh so it's uh and it it's important that those two competing motivations get negotiated. Um I guess what you're saying is separating the projects might make them a healthier negotiation.

SPEAKER_01

Quite possibly. Uh and and um you also there's also other risks with having the projects uh so closely associated with each other, which is um you could risk stifling innovation as well. Uh as you have people who are opinionated about how stream processing should work, you it's very easy to make changes or not make changes in the broker that don't that cause problems for outside systems, outside stream processing systems. Uh so that's something that, especially with the projects being so closely aligned, the community needs to be very cognizant of and make sure is not happening.

SPEAKER_00

True, true. And this part of the discussion now reminds me of the Android iOS uh the the different approaches between the Android and iOS ecosystems, where Android is very much an approach of dang it, let that thing fragment, let a thousand flowers bloom. You know, 993 of them might smell bad or die real soon, but you know, there'll be seven really good ones that we get to discover because of that, you know, that rich, bizarre marketplace of of attempts at innovation, compared to the iOS approach, which is like, well, let's let's kind of get this get together and design this centrally, and we make one very well integrated thing that works well and can provide a better experience than the fragmented thing. Um think that's a fair analogy?

SPEAKER_01

Uh it's not a bad analogy. Um there you can also go down the path of iOS tends to be a lot more stable than Android, which I would say is probably the opposite here. Um but certainly you then have the the problem that you see sometimes, which is oh yeah, I can't quite get an iOS app that does what I want because Apple controls the API too tightly. But I can find a bunch of Android apps that do.

SPEAKER_00

Yep, I uh have seen that very thing. Uh I am a uh hobby photographer, and there are certain kind of fringes of certain types of gear and processes and things like that where you can get programs that will automate things on cameras, and you can get them in the in the Android ecosystem, but you know, it requires like a USB connection from the phone or something that, at least in earlier versions of iOS, you couldn't do, and uh uh for that very reason you can't have the thing. So yeah, that's uh that's a good point. What do you see as uh the future of Kafka? What what do you think is gonna happen? Uh what do you want to happen?

SPEAKER_01

So I think that there's a lot of Kafka's done a great job over the years uh at being very high performance. Uh, but certainly one of the things that I've discovered, and especially recently uh in the last couple of years, as we've been trying to, you know, poke around the edges of Microsoft Azure and how we can use it, is that a lot of the performance we get depends on running on bare metal and running on very specific hardware, uh, which I think is a combination of, well, that's just how it was years ago, um, and the cloud ecosystems weren't as popular. And also that's just how it was at LinkedIn, and so that's how Kafka grew up. Uh, certainly you lose a lot of performance when you move on to virtualized systems, uh, and that's true for anything, uh, but I think it's it's more pronounced for a system like Kafka that just depends on uh disk.io quite so much. Right. Right. So I think there's a lot of changes around those types of things. There's a lot of gains that can be had with performance, a lot of interesting things that you can do with the way you manage disk, for example, or uh manage network. There's a lot of large changes that are necessary for that, though. If you want to start playing around with how you manage disks, if you want to start playing around with uh the placement of things on disk, you need a little bit more of a pluggable architecture when it comes to the underlying storage systems in Kafka. Something that's not present right now and requires a lot of effort and a lot of change to get it into place. And I'm not sure that that can happen with the current state of the project. It's one of the reasons that I'm a fan of de-aggregating uh Connect and streams from the project, if possible, and moving them out so that changes could be made in the broker without having to rev the entire project and keep it in lockstep. Make it more of an API-based approach.

SPEAKER_00

So uh and that's with storage in mind, you said.

SPEAKER_01

Storage is probably one of the big ones. Um again, there are there's a lot of opportunities uh to be had in optimizing the way Kafka works at a low level.

SPEAKER_00

Could you elaborate? You started to mention you know, thinking about where things go on disk as a performance evaluate uh performance um optimization. Could you elaborate on some internals there? That stuff is often fun to talk about.

SPEAKER_01

Well, sure. Just think about the fact that most, you know, so Kafka runs on spinning disk at LinkedIn for the most part. Spinning disk is cheap. You can put a lot of them in a system, um, and you can get you can get decent performance by just adding more spindles, by adding more disks. SSDs are great, they're super fast, they're really expensive. The the good ones are really, really expensive. Um, so you can't run the same type of storage density that you can if you want to run on SSDs. So if you're storing a lot on disk, you need to make use of a spinning disk and SSDs. Conversely, if you don't care about storage on, if you don't care about keeping your data around that long, maybe you want to run purely in memory. Um maybe you want to do something along the lines of infinite retention of data, which is very tricky in a system where you have to allocate partitions onto disk.

SPEAKER_00

Right. Uh so in those cases, uh those, you know, you you describe memory, SSD, and and spinning disk. Basically, for economic trade-offs, performance cost trade-offs, uh suddenly you need to know a little bit more about the storage subsystem than your typical cloud provider block storage service is going to give you.

SPEAKER_01

Absolutely. And and certainly we've seen with other large storage systems, uh, you know, like HDFS, uh, changing how you store data on disk can have significant performance impacts as well. So, for example, if you wanted to use a raw disk to store Kafka data onto, you again have to make some large changes to the storage subsystem. And those might not be for everyone. Uh LinkedIn might like that, and uh Apple might not. Uh LinkedIn might like it, and Microsoft might not. Uh, it may not work in AWS, whereas it you you may have some things that will work in Google Cloud. Uh so there's there's a lot of there's a lot of different things that are going on, even just with regards to storage, that it would be nice if we could take advantage of.

SPEAKER_00

Yeah, and that would be you're really talking about something pluggable, right? So I've I've got the on-prem thing, and here's my storage config uh on my bare metal servers. So that's that's one plugin, and I can optimize such and such a way. Or I'm running in this particular cloud provider, which exposes these APIs that will tell me uh not a whole heck of a lot, but maybe a tiny bit about my storage subsystem. And then I'm running in a second cloud provider that tells me nothing. And so those are sort of my my different scenarios. Is that my reading you right?

SPEAKER_01

Yeah, and pluggable is hard. Um when you're when you start talking about pluggable architectures, you have to have well-defined APIs that stay fairly stable. Um And I know that this is this has been a point of contention in the Kafka community in the past. Even just getting through something like uh message headers was very difficult. Um and interceptors was also extremely difficult because of this idea that you can have an API interface that people can write their own modules for. And oh my gosh, if they write really bad code and insert a bad, uh insert a bad interceptor into the stream and the performance tanks, they're gonna come back to the open source community and say how crappy Kafka is and that it's all our fault. At some point, you have to you have to let go. You have to let people make their own mistakes and have a clear deep delineation of where they made their mistake and how the system works normally.

SPEAKER_00

Yeah, which by the way is always my answer to that question, and I think that's why it's good that I'm not solely responsible for making any of those decisions. Um because even even though uh you know, I'm just kind of thinking back on your disaggregate um versus not question. And again, my gut response there is, you know, man, we need to be it, this the the transaction costs between uh teams are not such that uh or are you know are such that we should make this one team uh that that does this instead of separate things. There's a uh there's an ec there's an economist, uh, early 20th century economist, Ronald Coase, who uh developed a theory of the firm. You know, why do we have companies and not just people trading in the marketplace? And that had to do with at a certain point, uh, it's costly enough to access the market that you you you actually want top-down decision making within you know the open source community around the project or a company or whatever he was thinking about companies. So, you know, my gut take on the aggregate-disaggregate thing is that we're on kind of the the firm side of the Kosian equation where uh it's probably smarter and we'll have a better product if we're all in one place. Um yet, um, and I I may have I may have managed to lose my train of thought here, but the uh the situation you were just framing with plugins and APIs and all these things seems to provide a counterexample, right? Seems to be to say that that you know this would make it a good thing to have a little bit more organizational diversity.

SPEAKER_01

Absolutely. And you can also think about the client that talks to Kafka as well. For a very long time, the only client that worked well was the Java client. And everything else had to go through something that would go through the Java client. So you have, you know, we had a REST interface, and everything went through the REST interface. Now we have a little bit more proliferation of clients now uh that use the protocol more directly, but the protocol is very hard to work with and it changes very quickly, which means it's difficult for anyone external to uh to keep up with the rate of change and have a client that is not owned by the Kafka project itself. And I think um we definitely saw when Confluent got started that the rate of change of the Kafka project changed significantly. I was a fan of LinkedIn owning the project entirely because we could always drive it where we wanted to. However, I was also a fan of the community growing and taking it over because that means it's no longer dependent on a single company. And I think we're in a dangerous spot right now where the community could become too dependent on Confluence. So we were too dependent on LinkedIn for a long time. And now we're very dependent on Confluent and Confluence knowledge of how this works. I fear that we're in a place where we may not have good clients that are working or the ability to work outside of that ecosystem. Not just Confluence specifically. I'm really just using the company as a uh like a central point for the developers at that point. Uh, but I want to see more proliferation outside of the Core Kafka project, of lots of interesting things going on. Clients, plugins, all kinds of interesting tools. And I think that's only possible if you don't put them all together in one basket.

SPEAKER_00

Yeah, you need you need diversity in the community for it to develop uh enough different ideas. Um well, it's just like an ecosystem, right? You need enough genetic diversity in an ecosystem for when uh you know like a pathogen threat or a a new predator or something uh introduces is introduced to the ecosystem, uh, you have enough different ways of responding that you don't get whole species wiped out. So same thing for an open source project, right? Um you always get when it's a successful project, a company forms around it. You get this open core company, right? You and I could rattle off half a dozen of them apart from Confluent. Um, and that's always been an engine of investment in open source, right? People's salaries get paid to be able to develop this free and open software that everybody can use, and that's like uh economically a really good and amazing thing. Uh, but we don't want a complete overconcentration there, because then then the the um we go back to the cathedral and the bizarre sorts of ideas, you know, the the benefits of kind of the market economy of ideas, you start to lose those if too much gets concentrated in one place.

SPEAKER_01

Yeah, absolutely. I mean, Confluent is doing amazing things for the community and for Kafka in general. And you're you're absolutely right. We see this pattern all over the place, and it's necessary. I mean, Kafka wouldn't be what it is today if LinkedIn hadn't put a significant amount of money into it. Kafka wouldn't be where it was is today now if Confluent wasn't putting significant money into it and hiring developers and everything. But at the same time, you need that healthy open source community, and we need to make sure that we're enabling the community to do good things. And when you're talking about a community doing good things in software, you're talking about good, strong APIs, stability, and in you know being able to have all these other things going on. There's plenty of room in the pool for everyone to play.

SPEAKER_00

Yes, yes, there is. And uh to that end, just uh I don't know, this seems like a uh a tooting my team's horn thing, but let me for a minute, just so you can know my money is where my mouth is. Uh we've actually kicked off recently a program to recognize contributors that are that are outside of Confluent, you know, but contributors to Kafka um and some other initiatives around that just to make the process of contributing to Kafka and to Kafka ecosystem projects that Confluent is able to influence. Um we want that to be like a welcoming thing. Like we've kicked off a a way of looking at our README's and uh contributing uh.md files and you know how discussions go on and everything and and and just auditing. Are these thinking from a very GitHub-centric view of the world, right? Are these repos nice repos to hang out in? And when people approach them with contributions, uh, you know, you don't know that your PR is going to get merged. You have no right to demand such, but uh you ought to be treated well. You ought to be treated like a guest who's come with something valuable. And and um it's a thing I've been pushing on, and you know, we collectively on the Confluent side of things have been pushing on uh to try to make sure that it doesn't suck to be a contributor because everybody wants there to be more, and that's good for everybody when that part of the ecosystem is diverse.

SPEAKER_01

Absolutely. I'm glad to hear that uh Confluent is putting effort into that. I have always said that the community is what makes the software, and fortunately, I have always found the Kafka community to be very open, very welcoming. You don't get uh a lot of the down talk on newbies that you see in a lot of uh software communities out on the uh internet. It that's never been present in the Kafka community. Everyone, you know, there there are no stupid questions. There are just people who are not experienced yet, and they've always been helped along. That's one of the best things about this community. So we definitely need to continue putting effort into that and money into that to make sure that it stays that way.

SPEAKER_00

My guest today has been Todd Polino. Todd, thanks for being a part of Streaming Audio.

SPEAKER_01

Well, thank you so much for the time, Tim. As I said, I love talking about Kafka, so I'm glad to have the opportunity.

SPEAKER_00

Hey, you know what you get for listening to the end? A Kafka Summit discount code. Kafka Summit is coming up on September 30th and October 1st in downtown San Francisco, and you can get 30% off if you go to Kafka-summit.org and use the discount code AUDIONET during checkout. Just enter Audio 19 while registering at Kafka-summit.org, and that 30% off is all yours. I'd love to see you there. But hey, I hope this podcast was helpful to you. If you want to discuss it or ask a question, you can always reach out to me at T L Burgland on Twitter. That's T-L-B-E-R-G-L-U-N-D. Or you can leave a comment on a YouTube video or reach out in our community Slack. There's a Slack signup link in the show notes if you want to register there. And while you're at it, please subscribe to our YouTube channel and to this podcast wherever fine podcasts are sold. And if you subscribe through iTunes, be sure to leave us a review there. That helps other people discover the podcast, which is a good thing. Thanks for your support, and we'll see you next time.