Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov

Inside OpenAI’s Streaming Backbone with Aravind Suresh | Ep. 24

Confluent Season 2 Episode 24

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 30:51

Adi Polak talks to Aravind Suresh (OpenAI) about his career in distributed systems and real-time streaming. Arvind’s first job: coding at school. His challenge: turning OpenAI’s fragile Kafka setup into a reliable, multi-region streaming backbone.

SEASON 2
Hosted by Tim Berglund, Adi Polak and Viktor Gamov
Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
Music by Coastal Kites 
Artwork by Phil Vo 

  •  🎧 Subscribe to Confluent Developer wherever you listen to podcasts. 
  • ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
  • 👍 If you enjoyed this, please leave us a rating. 
  • 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
SPEAKER_02

As long as we are able to build a system, it could be uh abstraction that makes it easy for our users, and as long as it solves a problem, that's good. I've noticed time and again that simple systems scale better, any kind of like clever, complex code. Ultimately, in the long run, although if the overall architecture is simple to explain to others, will it stand the test of time?

SPEAKER_01

Hello, I am Adi Polak, and you're listening to Confluent Developer, where we explore the fascinating journeys of software developer tackling complex problems. In this episode, I'm interviewing Arveen Suresh from his early days exploring distributed systems to becoming a key figure behind open AI, real-time infrastructure. Arvine shares a story of scale, simplicity, and building the right abstractions for AI. His early passion for coding, starting back in school, eventually led him into the world of data streaming, large-scale systems, and enterprise software. So, what did OpenAI do differently to scale their data streaming platform? Have a listen. Hi, Arvind. I'm so excited to have you on our show.

SPEAKER_02

Hi, Andy. Uh executive here. Nice to meet you.

SPEAKER_01

Yeah, well, thank you. First of all, thank you so much for being part of Corinth Nola Program Committee. Your insights and comments were fantastic and really helped us build a good program. And also, thank you for being a speaker and sharing your expertise. I know, you know, behind the scenes, there's a lot of work that you put into actually taking everything that you know and turning it into a presentation that people can learn from. So thank you for doing all these great things for the community.

SPEAKER_02

Yeah, thank you. Excited to share our learnings as we uh like whatever we have done uh at OpenAI. So it was fun being part of the conference.

SPEAKER_01

That's that's fun. So I definitely know you. A lot of our program committee, and maybe some of the folks uh from Corinth knows you, but not everyone who are listening to us today do. So maybe you want to share a little bit about yourself.

SPEAKER_02

Sure, yes. Uh uh I'm Arvind. I'm currently at OpenAI. I lead the real-time infrastructure team. So we power the streaming backbone behind all AI products like ChatGPT, Soda. We also help a bunch of research workflows too. So I've been doing distributed systems and data platform for like close to eight years now. Uh and I've been like I've I found this space uh sometime. I can tell you about that in a little bit, but like the there was a moment in my life when I figured out that distributed systems was something that I was really passionate about, and I've been like diving deeper, uh building systems that scale pretty much my entire career.

SPEAKER_01

That's amazing. I mean, there's so much to know about distributed systems. I always feel like it's it's a whole world that when you start digging into it, when you're a user of a distributed system, let's say you're a consumer of an API, it's it's one angle. You try to optimize, you have a product mindset. But then when you dive deeper into the infrastructure, it's like, oh, things are not consistent. What do you do?

SPEAKER_02

Yeah, exactly. Like uh, I think this is a good point about infra and product, right? Like um, I've been part of infrastructure teams pretty much my entire career, and it has always been uh I like working with product teams or like different teams in the company because ultimately we are here to may build the foundations that make products successful. So we're co-building infrastructure with you. So in early stages of the team or the company, you end up co-building your infrastructure with the product team, and the product teams will be co-building infrastructure and their products with you. So I really like that experience. That is more fun, understanding their use case, partnering where we can help.

SPEAKER_01

That's amazing. So I know you've been, you know, you have many years' experience with data streaming and kind of data streaming is is a big passion and uh and the professionalism, you know, career that you chose. Um, I'm curious, like, what are some of the challenges that you've seen throughout your careers that, you know, and learnings that are people can learn from you?

SPEAKER_02

Sure. Um, so let me start with like um my recent experience at OpenAI. So the challenge here was basically how do you build a reliable streaming platform when you're scaling 10x every six, seven months? So your platform scales 10x every six, seven months, different sets of use cases, different shapes of workloads that we support. Now, how do you make sure uh the platform is also reliable when we do all of this, right? The core challenges uh involve something like balancing velocity and stability with like a such a small team, and how do you work with users? How do you hide all the infrastructure complexities? Um, so that just to give some context about uh OpenAI's use cases, right? The streaming infra is kind of at the heart of OpenAI in like different workflows. Like uh one of the core things that we use Stream Infra for is something called Flywheel. What that means is um we build great AI models, then we launch products using those AI models, then you get some usage data about that product based on what our users, how our users use their products. And insights from these are actually used to build better models later. And this kind of loop keeps spinning. Like you have more models, you get better usage data, then you get insights, you train better models, and so on and so forth. And guess what? The data pipe here is Kafka. So Kafka is used to store all of these events, and you use Plink a lot of stream processing to get all this data as real time as possible wherever it's needed, right? Um, and also in all uh distributed system settings, let's say experimentation is heavily distributed today. So the our researchers and engineers would like to know the latest status of an experiment and see whether it makes sense to continue that or not. And today's uh training is heavily distributed, right? So you have workers talking to each other, and Kafka has the data pipe to carry those events, and Flink is used to transform and join that to basically understand what to do. It's kind of bread and butter of uh researchers and engineers. And um, I've seen this in action. Like without streaming infra, it was very hard for them to build things. With stream infra, it was actually easier for them to do their day-to-day. So, in this background, um, this is basically kind of the result, right? But this is not how it was when I joined uh the company. It was like early stages, uh, it was still a scrappy uh piece of infrastructure. Like there was no standardization around Kafka clusters at that point. Like there was no blessed paths. Uh, Kafka was a single point of failure, uh, and there was a lot of friction and confusion on like how to use Kafka. So this is the state there, and like it was like close to a year or two um years of work to get us to this shape.

SPEAKER_01

Wow. So let me see, you know, collect all the all the pieces of information. Essentially, when anyone in the world decides to use OpenAI models, uh, we as humans can give some sort of feedback, like thumbs up, thumbs down, uh, or have some sentiment to uh uh to what we do. Uh if we're happy, not happy, asking for uh improvements and so on. And that essentially becomes a feedback to um to for retraining or rebuilding the models and improving them so the the whole system can become better. And that's as a feedback turns into Kafka events that are being processed later on with Flink in a streaming engine in real time. And that's kind of like the backbone for everything that you do. And then throughout, you know, as as you have these uh data data streaming pipelines, essentially it's not serving only that feedback loop with the model, but actually, you know, also researchers today that want to run different experimentation, or anyone who uh you know wants to tap into what is happening in the system needs to have access to that real-time data, essentially. Right?

SPEAKER_02

Yes, that that is cut, yes.

SPEAKER_01

Got it, got it. And then your team has to manage the cluster, right? Uh manage the large scale of this distributed system. Um you know, I'm guessing that's that means you know, a lot of machines. Um and then um make sure it runs with the right latency for the system, right?

SPEAKER_02

Yes, yes, that that is correct. The the emphasis on reliability is actually very high uh because all of this is mission critical data, it's important to us. So the we promise close to four, four, four and a half nines of Kafka uh availability. And also the core uh core principle behind all of this was uh building reliable systems was of course important, but another key principle that we kept in mind is how do we make all of this easy for our users to use? Because uh, when I say users, so we are an infra team. When I say users, I talk about all the other employees, engineers, researchers at OpenAI who would use all of this. So, how do we make that experience easy for them? So the ergonomics part was very important. Uh, it was like building AGI and building AI systems is hard in its own way. People did not want to become Kafka experts, right? Uh so how do we build the right abstractions? Uh, we made some early decisions. Like, for instance, instead of exposing Kafka clusters directly to users, we built layers of abstractions. So we built we built a bunch of proxies surrounding our Kafka clusters so that we can abstract Kafka away from our users. It helps us to do a lot of operations, scalability. It also helps in uh multiplex. So our Kafka topologies like Kafka messages get multiplexed across different clusters and different regions, and that helps in higher availability. That way, Kafka would not be a single point of failure. It does have some trade-offs here and there, but happy to add details there. But this actually simplified uh consumption of Kafka and like building workloads on top of Kafka easier for our users. Because tomorrow we can scale a cluster, we can add more nodes, we can rebalance them, we can move them from one region to the other, but having a layer of indirection in between actually helped us make those it made all those operations easier.

SPEAKER_01

Got it. So if essentially, um just to see if I understand correctly, different people in the organization, these are our users, different people have different skills and different expertise. Like ML researchers, they have you know probably the best knowledge and on how to build algorithms and how to improve them and so on. Uh, and they have knowledge of the tools that they're using, but we can't really ask them to learn everything about Flink and Kafka. And so we want to give them uh tools that are uh intuitive to what they do by building a layer that uh abstract away Kafka and Flink complexities uh so they can run faster. That's it, right? Okay, got it. Yes, and what are some of the challenges that you ran into? Because I can imagine, you know, people would have all kinds of requests.

SPEAKER_02

Yes, yeah, that's a good question. So our challenges here was mainly around Kafka being a single point of failure, but that's actually a very hard problem. Like uh if you deploy, let's say, Kafka clusters in a single region, and how do you make how do you make that work in a multi-region setup, right? Like either you can go down the path of having stretch clusters, which I've seen in some cases, uh, where you have Kafka itself distributed across regions. Uh, there are obviously latency considerations, stability at scale considerations here. Um, there are some other uh designs I've seen where you have mirroring from one Kafka cluster to the other, so that in case a region goes down. So we were solving about cases where even the cases where regions were going down, we still wanted Kafka to be up. So even in those cases, um we we have we basically had like a bunch of these design options, which were all kind of complex. Um we didn't know, we were thinking whether to do stretch clusters or whether to do mirroring. And that's when we decided, can we just make some simplifications here? We thought about uh, okay, we can have built layers of abstractions in between. And let's say uh you multiplex messages from one uh topic across multiple Kafka clusters. And when we do that, you obviously throw out some notions of like partitioning, ordering, which are kind of very, very important. Like uh for folks who come from a traditional Kafka world, they would be surprised, they would find it odd that we threw away these core primitives that Kafka became famous for, like partitioning, ordering. But we realized that a majority of use cases uh were okay with not having those guarantees. They were okay with, let's say, for partitioning, they were okay with doing that downstream, let's say the Flink app or the downstream consumer. And ordering, as you know, is a hard problem in distributed systems itself, right? So instead of relying on reliable ordering of events, why can't you just stamp your events with a logical clock and use that to uh sort in the downstream? So we made a bunch of these key trade-offs and we went back to our users and talked to them if all these things make sense to them. And they were okay with it. And that's when we actually went and built our first version of Kafka at OpenAI. It was basically a producer-side proxy which multiplexes messages across multiple clusters, which again there is a consumer-side proxy which pulls all of this for you and sends it to your end consumer services. So this was our first version. This came in sometime around mid-2024.

SPEAKER_01

Amazing. So essentially, it took a lot of complexities out of Kafka and some some a lot of uh failure guarantees, I will say. So it's probably you know, coming. I'm sure like the person that came up with that idea was everyone looked at it and was like, Who are you? Do you know Kafka? Um, but apparently, you know, it works because at the end of the day, you're serving a user, and these users are okay with these trade-offs.

SPEAKER_00

Now a quick word from our sponsor. Confluent Developer the Podcast is brought to you by Confluent Developer the Website, which has everything you need as a developer of data streaming systems. And it's completely free. We've got curriculum, hands-on exercises, executable tutorials, the online data streaming engineer certification, also free, a way to find a meetup near you, those are free. Everything is there. I really want you to be successful in your journey as a data streaming engineer, and this is the site that has what you need. Check it out at developer.confluent.io. That's developer.confluent.io. Now back to the show.

SPEAKER_01

It's it definitely shows like a flexible mindset of an engineering, because you know, many times in our career, and I my background is also in the in the um distributed system infrastructure. There's like the things that we're used to do, the way we forever did things. And um being user-obsessed is definitely um a value that can help us build systems that cater to our users versus trying to answer all the best practices in the world, I'll put it that way. Um so that's that's exciting. I mean, that's that's a big learning.

SPEAKER_02

Yeah, I totally agree with what you said, right? Like uh ultimately we are looking, we are building systems that solve problems. Uh, it doesn't matter whether we use a Kafka or a database or whatever tools are going to hold. Uh as long as we are able to build a system, it could be uh an abstraction that makes it easy for our users to use that system. And as long as it solves a problem, that's good.

SPEAKER_01

Right, right, right. No, at the end of that's that's what we're paid to do, right? Solve the problem. So I love it. Cool. Well, thank you for sharing. That's um that's really exciting. I know you build more things, right?

SPEAKER_02

Yes, uh, we built um so immediately after we shipped this, there were more this solved all the stateless consumption from Kafka. And then after that, uh there were use cases around stateful processing. So that's when like we started adopting Flink more at OpenAI. Yeah, and also the some of the problems are similar across both. Like, how do we make sure Flink apps scale? How do we make sure seamlessly with this multi-cluster Kafka? Like it's it's a different topology, right? If you have Flink working with a single Kafka cluster, there are in uh everyone has been doing that for years, but how would you make it work with multiple Kafka clusters? What uh how do you we had to build uh a custom flink source for it to work? Then we had to build a custom flink sync, then whether then we had to invest in control planes because it um it's only a matter of time. This is also something that I observed in the last year or so. It's only a matter of time until uh the system that you build, which is horizontally scalable, after a year or so, you end up with like a lot of clusters, a large split of clusters for you to manage, and it ends up with a lot of operational toil, on-call pages, et cetera, which is when we realize the importance of control panes and importance of self-healing systems, auto healing systems. And that's kind of where uh we went next. So we introduced control planes for all parts of an infrastructure, including Kafka, our consumer side of Kafka, Flink, all of these systems ended up having control planes just to manage the scale of clusters that we are dealing with.

SPEAKER_01

Yeah, that's that's smart. It's um from a certain scale, you have to have it. I remember uh back, well, it was uh when I was in Microsoft some years back, but it was uh post-acquisition of LinkedIn, they were building um different class different Hadoop clusters, and then there was a control plane that managed a couple of hundreds or thousands of uh mini Hadoop clusters just to uh support the scale. And um yeah, I believe uh today with uh cell architecture, um it is highly used. This is uh what we have in in Confluent, they call it the Cora engine. There is a paper out there, uh, but we essentially we build cell architecture so we can support as many Kafka clusters as needed.

SPEAKER_02

So I see what is a cell here?

SPEAKER_01

So the cell architecture is essentially it can be a collection of um well, a Kafka cluster that supports a collection of things in it. So you can read about the article, it's we you know, it's it requires it requires looking at the architecture, but essentially this is how uh AWS, for example, build S3. So this is the architecture behind that. Um so it's a multi-tenancy, multi-tenancy environment that keeps that takes care of scaling, security, uh, and a lot of other aspects. And that was a great uh, especially in the recent outage that was, um, confluence stayed up, which is important because you don't want to put everything in on one region. Uh you want to continue supporting our users. Uh so I'm very happy you're building a control plane and you have different control planes for what you're doing. Uh, and I'm curious because you know, adding Flink is something that you know you have to think about. Like, how do you how do you simplify Flink for the rest of the organization?

SPEAKER_02

Yeah, so it started off with building the right um uh so ultimately, uh yeah, rightly said, Flink has its own complexities. Uh like it's it has its own learning curve, right? It's not as easy, like it's easy to talk to Kafka and publish messages, but it's not so easy to write a Flink app. So it started off with um us working closely with users to explain Flink semantics to them and building um so most of the projects that ended up having complex Flink apps uh ended up like a lot of co-building with users, like working closely with them. But from an infra perspective, what we did is we made sure the Flink control plane handles all of the complexities around deployments, failovers, rolling back clusters, etc. Uh, like rolling back deployments that didn't work, etc. And on the user side, we made it very easy for them to actually create a Flink app. So we had like uh toolings that can generate all the boilerplate for you so that it's easy for users to actually just latch onto a Kafka topic and consume uh stream of data from that and write flinks up. And uh yeah, go ahead.

SPEAKER_01

So if if I'm a user and I want to create a Flink app but I know nothing about Flink, um do I need to use Python or Java, or do I have any specific requirements of how I write my code, or there's no code at all, maybe I can ask ChatGPT to generate the code for me.

SPEAKER_02

That's a good question. So we are getting there. So right now, what we do is so we adopt open is heavily, uh a lot of services are in Python. So we also adopted Pyflink. So the initial version of this was a good amount of scaffolded code that helps users to write Flink apps. The second version of it, which we are kind of going through right now, is uh giving users better tools. To understand their flink jobs. Like, for instance, coming up with um like giving tools for them to understand a health score behind their Flink app, whether they're following all the best practices and giving them a score like 80 out of 100, you're losing 20 points because of this, and these are things that you need to do, etc. Like you're managing your state. Another thing we notice that we are currently doing is our Flink state uh for some of these jobs are becoming very massive. And when our users end up uh debugging these apps, they need to look into their state and inspect. And actually, it's a very big challenge because uh once our app, if you want to debug your app, you need to look into the state. And without looking into the state, you're kind of blocked. And some of and most of these times, these states' files are very massive and they're stored in top store. So we are currently building a state file inspection tool. So basically, you can debug, look at all the state that's managed by your Think app, trying to understand what's going wrong, how to optimize things further. Some of these, even subtle bugs in your code, end up creating giant state files, even though you don't intend to. Uh, where we believe in the long term, majority of link apps can be represented through SQL or FinkSQL or a DSL. So we're trying to carve the structure of that DSL out and have like a managed platform where users can just give in a SQL or just codify their app in a DSL, transformations, joins, aggregations in a format, and then we manage the lifecycle of the app behind the scenes. So this is the no-code vision that we have. We can also integrate AI into it. You can have prompt engineering, like you can use natural language to define your Flink app, your Flint app DSL, and then you hit uh deploy and it gets deployed.

SPEAKER_01

Got it, got it. So if let's say I'm more on an engineering side, but I'm on more on the product side, I will be able to write a SQL and also have all the tools to inspect what is happening in my platform in case uh you know performance is not what I want, and I want to improve it. So I'll be able to actually have that uh access to the data. That's really interesting because one of the greatest challenges with Flink is state management. And you know, so many people get it wrong for for good reasons because you have to know so many things in order to uh make the right decisions, and it's just again, like we cannot expect everyone to know that. So that's that's really exciting. And the fact that you're using Python as well with Flink, is the infrastructure also in Python, or do you have the infrastructure side and more of Java world JVM?

SPEAKER_02

Yeah, so uh our infrastructure, at least for data platform, all of infrastructure is mostly in Java, but the end user abstractions are in Python. So Kafka client is in Python.

SPEAKER_01

Oh, got it, got it. Cool. All right, so it's um yeah, it's how if if there is an error that comes from the the Java's JVM stack, uh like how do you make sure the user can understand that there's uh I'm sure you kind of you have the tools and you you think through that, but uh that's always uh you know one of the challenges. Like let's say I'm you know, I never touched anything related to JVM in my life, only did Python, um which is not true. I I I was heavy in uh in the JVM space, but let's let's pretend for a second. Um I wouldn't know what to do with the Java exception, if I'll get it.

SPEAKER_02

Yeah, so actually we have faced a problem even before that. Um sometimes we noticed um the actual error inside your Flink app could be multiple layers deep, could be in some different file, some different place. Uh and uh we once we face enough problems of this shape, we uh we practice a workstam to basically build a UI that can just tell you exactly if there is a Flink app failure. This is what uh uh so basically debugging your Flink apps, getting the right exceptions. And also uh we are trying to go ahead with like uh we built an AI agent that looks into these exceptions and with enough context about your app, it will kind of tell you what the right next steps are and what we should do. Uh and you can keep uh there is a feedback loop here because if it doesn't work, you just give us feedback and then we try to fix the agent. But uh this is also something that we built.

SPEAKER_01

That's awesome. I love the usage of AI to assist in engineering work.

SPEAKER_02

Yeah, um I've seen like multiple examples. Like I think AI agent uh debugger is one example. Uh AI assisted uh Flink app creation through a custom DSL. That's like another example. Uh yeah, I'm seeing a lot of examples these days, making us more productive.

SPEAKER_01

Yes, 100%. It's um I see it how how it improves everything that we that we work on. And um, you know, even it's kind of like uh having a body tells you it's like, hey, you know, here and there you should take a look because it might be that the error is is you know, here are three potential reasons why why you got this error, start exploring rather than you know individual trying to uh figure it out on them on their own. Um so it kind of helps bring the collective brains of everyone in the company to uh to solve a problem without getting everyone involved necessarily uh in a meeting or in a uh you know gathering of other Slack or other tools. So that's uh this is great. This is great. So you've you've been working, that's you've been solving a lot of problems there.

SPEAKER_02

Yeah, um every day is exciting, a lot of scale, and um a lot of different shapes of problems too. So so uh as I said before, like uh like solving problems with the users because every um and trying to find the right abstractions that would actually generalize to a more two different use case shapes. It's kind of what I do on a daily basis.

SPEAKER_01

That's awesome. So if you would, you know, to give your younger self or someone that you mentor that now is entering engineering, just you know, finish their degree or want to enter the data streaming world, what what would it tell them?

SPEAKER_02

That's a good question. Um, a couple of things. Uh, one is I've noticed time and again that simple systems scale better. You could have different notions of uh any kind of like clever, complex core, but uh ultimately in the long run, although if the overall architecture is simple to explain to others, if uh only if that system is simple, it will stand the test of time. Learn this the hard way a couple of times. Uh and the other thing is uh you could uh the long-term reliability actually comes from like a lot of automation. This goes back to the control planes point which I made earlier. Your initial wins and reliability could just come from a nicely designed uh stack, but in the long term, you would uh you would notice a lot of manual operations, you would a lot of routine operations, rebalancing clusters, provisioning them. And so the long-term reliability would just come from automation. So building control planes, self-healing systems. So these these are some lessons that I learned. But uh understanding details behind every system was actually important. It's actually more important to understand what it is for and what it is not for, and getting that clarity because that is how you build the right building blocks um it for which other users can use and build great systems.

SPEAKER_01

Love it, love it. So simplicity, uh understanding the building blocks behind things and building uh what's necessary. So sometimes you don't have to uh over-engineer uh a solution. Arvin, thank you so much for sharing the different things that you work on and sharing your knowledge and expertise. Uh it was super exciting, and uh yeah, it's uh I'm sure people would have lots of takeaways from these.

SPEAKER_02

Yeah, thank you. I mean nice chatting with you.