Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov
Hi, we’re Tim Berglund, Adi Polak, and Viktor Gamov and we’re excited to bring you the Confluent Developer podcast (formerly “Streaming Audio.”) Our hand-crafted weekly episodes feature in-depth interviews with our community of software developers (actual human beings - not AI) talking about some of the most interesting challenges they’ve faced in their careers. We aim to explore the conditions that gave rise to each person’s technical hurdles, as well as how their experiences transformed their understanding and approach to building systems.
Whether you’re a seasoned open source data streaming engineer, or just someone who’s interested in learning more about Apache Kafka®, Apache Flink® and real-time data, we hope you’ll appreciate the stories, the discussion, and our effort to bring you a high-quality show worth your time.
Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov
The Future of Serverless and Streaming with Neil Avery
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
Neil Avery explores the intersection between FaaS and event streaming applications before taking a quick detour back in time to understand how we've gotten to this point in event-driven applications. He'll explain the pros and cons of FaaS, and cover how in its current state cold starts and latency concerns need to be part of the bigger picture when building streaming applications. Finally, Neil shares five rules that will help you understand how FaaS fits with the event streaming application.
EPISODE LINKS
- Journey to Event Driven – Part 1: Why Event-First Thinking Changes Everything
- Journey to Event Driven – Part 2: Programming Models for the Event-Driven Architecture
- Journey to Event Driven – Part 3: The Affinity Between Events, Streams and Serverless
- Journey to Event Driven – Part 4: Four Pillars of Event Streaming Microservices
SEASON 2
Hosted by Tim Berglund, Adi Polak and Viktor Gamov
Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
Music by Coastal Kites
Artwork by Phil Vo
- 🎧 Subscribe to Confluent Developer wherever you listen to podcasts.
- ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
- 👍 If you enjoyed this, please leave us a rating.
- 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
Serverless has been in the exciting part of the hype cycle for, I don't know, a couple of years, so it seems like we should start to be understanding it by now. We'll talk about serverless and streaming platforms on today's Streaming Audio, a podcast about Kafka, Confluent, and the cloud. Welcome back, everyone. I'm your host, Tim Berglund. I'm joined in the virtual studio today by my friend and coworker Neil Avery. Neil is in the office of the CTO at Confluent and recently, by which I mean, you know, in in uh the past many months, has been spending his time thinking about serverless. And I think, Neil, broadly, that's what I want to talk about today, but I think we want to take our time getting there. That's such a big topic. I don't just want to say, hey, let's talk about serverless, you know?
unknownYeah.
SPEAKER_00Anyway, welcome to the show. Thank you, Tim. Um happy to be here. Um yeah, serverless is is one of these topics that's really kicked off over the last 18 months, two years, maybe a bit longer, and you have to have been living under a rock if you haven't heard of it. And obviously, there's been a lot of um tooing and froing about what is serverless and what is FAS and backend as a service. Um so it's kind of really jumped onto our radar as something that we we're aware of, and it's something that we're embracing as we see technology move forwards to the next generation of thinking. You know, we're on this evolving path of technology. So it's quite interesting.
SPEAKER_01And uh yeah, as you say uh evolving path of technology, I think to think clearly about serverless, it helps for us to try to put it in some kind of context. Like, like, you know, we're always in uh some sort of developmental process in tech and in anything really, but uh I'd like to ask you kind of what do you think is happening now that's shaping the future, the future of information processing, streaming? I mean, I know that's a big question, but you you think about big things like that. So what what do you see that's happening right now that that is driving us towards something new?
SPEAKER_00Um yeah, I guess it's really interesting because I I guess we're seeing this convergence of everyone is moving towards events. You know, events are the way that people start modeling their domain. We've got domain-driven design, we've got event storming practices, and obviously we've got event-driven architectures at at the core of that. Um and serverless is is often toted or touted as, I should say, as an event-driven architecture. But in the world of streaming, for us, an event-driven architecture is one which runs through streams. Um, and and to give some clarity around that, when we think of an event, events they don't occur in isolation. An event is a reaction. So if an event has occurred, it's been triggered by something, and it's likely going to trigger something else. So, for example, if I am on a website and I'm adding items to my shopping cart, each of those is an event. And when I hit purchase, that's an event. The most important thing about events is the fact that we capture them in a temporal on a temporal basis, in a time-aware basis. Um, which means when we have a stream, we care about when the event actually happened. So when we build an event-driven application, we care about the time at which the event happened and the order in which those events occurred. So if we're going to build an event-driven architecture on serverless, then it makes sense to build it off the back of a streaming technology like Apache Kafka. And that's that's where we see this. Yes, it's a good example. Um, but that's where we see this affinity between where things are going. You know, large organizations are saying, let's move to this event-first thinking, you know, and us being um confluent, you know, event-first thinking is along a lot of the time it's where we start our conversation with customers.
SPEAKER_01Absolutely. And I think in our in our uh one thing we can say is that we were thinking of events slightly before it was cool. Yeah. Um at least a couple years before it was cool. So uh there's a couple things in there, uh, and I guess this isn't really a question, but I'm just reflecting on what you said about the relationship between event-driven architectures and serverless. Let me see if I can get this right. Um if if you are doing serverless, and hopefully we'll end by the end of this discussion, we'll have an account of what that means. If you're doing serverless, then um that's sufficient, uh kind of a sufficient condition for you to have built on top of a streaming architecture or event-driven architecture. If you're if you're doing serverless, there has to be uh, you know, you are by definition doing an event-driven architecture. But there are event-driven architectures which are not serverless, which we would not call serverless.
SPEAKER_00Yeah. Is that yeah. I guess just to clarify it a little bit, when when I'm talking about serverless, I'm talking about functions as a service where you can bring your own code and it runs, and it you'll have many instances of those running in parallel, um, depending on how many times you call on the API. But to rewind on that, you know, when serverless or functions as a service first came about, you know, the the use case that that Amazon had created it gave was it was to render a thumbnail. It was to render something that wasn't necessarily business critical.
SPEAKER_01And I'll I'll I'll have you know, so far, all of the conference talks about serverless that I've seen, and I guess I haven't really made it a point to see too many, but I don't want to say all of them, but there's a strong bias in conference demos and examples towards thumbnail rendering for serverless, right? Like that's the that's the that was like back when you were trying to explain Hadoop, it was always word count. And when you're trying to explain serverless, it's thumbnail rendering. And I I think thumbnail rendering is great, don't get me wrong. Like we need all kinds of thumbnails. But it's as you say, not a business critical process and maybe you know conceals some complexities elsewhere. So I interrupted you, but but thumbnails. Thumbnails are great.
SPEAKER_00Yeah, so thumbnails are the hello world of the serverless world. Um and fortunately, we move on from hello world to trying to build large-scale business critical applications, and that's where the streaming platform or event-driven architectures need to use the log or the streaming platform um as a nervous system, as a way of capturing those events and emitting those events at scale. And that's that's kind of where we see the relationship tie in between them. Um, I guess the other thing that's important about that is we're now seeing serverless used for machine learning, for all of these other things. Um, so it's just quite quite interesting in that we see them being used for you know some basic ETL pipelines, some data enrichment. But when we get to business critical applications, you know, we we have a few different aspects to consider.
SPEAKER_01Yeah, so I want to ask you what those are. Um I would I would love for you to drill down into that. But again, um you you keep saying important things. Uh the log, uh that that whatever this system is, it's built on top of a log. And events, whatever those are, maybe we can get to that. Uh, those are stored in time order durably. Um, and you kind of build your system on top of that. Expand on that a little bit. And you you were you were about to kind of get to the next step. Go ahead and and uh take this a step further for us.
SPEAKER_00Yeah, sure. Um so this is where we start seeing, you know, technologies like Kafka aren't just for moving data around. They're for building a rich ecosystem of interconnected data for streams of data. And at the ends of those streams of data, we have stream processors, maybe we have serverless functions or FAS. And then the output of those goes yet into another stream. And this provides us with the capability to do the stream processing against those streams, but also shape it in a way that the stream and the events within the stream are they become our persistence mechanism. Um, so a bit like when you have a database, an external database, and you want to capture all of the events within a table on a per-row basis, as those as those events change over time, they form a stream. And a bit like a database log or a transaction log, the log that that's transmitted into our ecosystem over is the source of truth. And the stream processor that sits over top of that provides us with a view, like the parallel to our table. And at the same time, the stream processor is not only doing view rendering, it's also processing logic. And this is kind of where we see a difference in the way that we used to build distributed applications to ones that are now uh event streaming applications, where we have events which are captured and stored within streams which are held within the log, um, and they scale out horizontally to our stream processes and our serverless functions. Um, and that's kind of where we're seeing the whole thing tie together.
SPEAKER_01Yeah. Okay, that makes a certain amount of sense. Restate it in case you're a listener and you're new to that idea. Uh I know the first time you hear it, it sounds crazy, but rather than you know, you you processing some event, and that's what your code does, is it whether the event is a form submission or a file gets dropped in a directory and you need to ETL it or whatever it is, right? You respond to events. Rather than taking those events and serializing them in database rows, we're saying, no, no, no, they're in vet, they're events, and we're gonna keep them in a persistent log in time order and then process them. And those processors are where our logic lives, and turn that otherwise rather unstimulating log into an interesting view of something. Right? Because you don't you can't you can't query a log, it's all you can do is read it. So to make queryable things, you've got these things that you call stream processors that that crunch on that log and and create some interesting view and expose it through some API to the world, um, which is a conventional stream processor. Uh, and that could just be some program running, you know, some some Java program that you write or or whatever. Um and it could also be a serverless function, which will we'll still come up with a definition of that. I we're trying to still lay some groundwork here, but uh you're saying those are kind of two possibilities for how to process data in a streaming system.
SPEAKER_00Yeah. And I guess the thing, you know, when when we apply, when we start designing our systems in terms of events and sending those events through streams, whatever receives that, the event starts to become, or the event is the API. Conventionally in distributed systems, we said, here's my API for doing stuff. When we're building streaming systems, the streaming system receives the event, it expands the event, and then it understands how to interact with it. Um, the nice thing about this is not only do we have a decoupled horizontally scalable system, we have something where the events can evolve and change over time if we think about events as APIs. Um, and this is the streaming world, but we can also do the same kind of thing when we come to serverless functions. So I think the streaming platform is a backbone for driving um event sourcing and um event-driven systems, when it's business critical, these become a natural way of thinking about it because you don't want to lose data, you want to guarantee ordering and correctness like you do in a relational database. Um, and as the log and the streaming platform, which is Kafka or something else, um has that parallel kind of business criticality to a database, except this is large scale and distributed.
SPEAKER_01Right, large scale distributed. And it's it's a different abstraction. It's the abstraction of a log rather than for a conventional relational database, the abstraction is tables at the at the level of the user or the architect or software developer who's building a system on it, you see tables. Um and when you're building something on Kafka, you see logs.
SPEAKER_00Yeah. Um and then yeah, with your stream processes, you build views, you can start turning to, you know, sorry, you can start building the the turning the database inside out that we've come to talk at, or the deconstructed database where you can mirror the parallel to what a relational database is, but you can do it with streams and stream processes. It's it's immensely powerful. I guess the other thing to note here is the way that we think about our streams and our stream processes. Anything that's consuming events from a stream is essentially a data-driven microservice. Um and the nice thing about this incarnation of microservices is you get the scale out, you get the service discovery-free topics, you get a whole bunch of other things that people have been trying to solve with conventional REST-based microservice frameworks for quite a long time. Um, a lot of this naturally gets solved with the streaming streaming platform.
SPEAKER_01I would at this point encourage the listener to Google that phrase if it's new to you, turning the database inside out. Um it's uh an incredibly powerful. It's not even an argument. It's just a notion that uh built uh microservices built atop uh distributed log actually delivers on the promises of microservices. We've been saying things about microservices for a good half a decade or so, and I think kind of stumbling around as always with a new paradigm in in the the half-light uh trying to figure out how to do things. And that that these these data-driven microservices sit on top of the a streaming platform idea actually delivers on the promise of microservices and the promise of evolvability that everybody said they would give us uh like that was not happening in REST-based microservices, I think, adequately. And now uh it seems like there's some hope that it will.
SPEAKER_00The other aspect to that is you know each stateful microservices, uh microservice instance is holding its fair share. So you're gonna run many of them, um, and together you get this horizontal stateful storage solution. The conventional problem with microservices is how do I handle state? That get that problem gets solved naturally by turning the database inside out.
SPEAKER_01Yeah, yeah, very elegantly. And you get at scale, at like organizational scale, when you start to see, you know, you have 100, 200 engineers working on the same system and it's all a gigantic microservices estate. Um, you start to see um uh an ecosystem develop, you know, where one group of developers can actually be unaware of services that another group is standing up because you've got this sort of fertile soil of the log uh down underneath it all. And so you actually get sort of a decoupled evolvable system because this is a big giant segue back to something we said before, because you have reckoned events as APIs. Now, uh I know exactly what you mean by that, uh, but I would love for you to drill down into it for somebody who has heard that phrase for the first time. Uh, what does that actually mean that might force you to give an account of what an event is? Sure. So, you know, if I'm writing code, how do I think of an event as an API?
SPEAKER_00Yeah, um I mean this is quite a fundamental um thing to reason about. So when we conventionally built distributed systems, I would have a hello world service and I would have a hello method against which you can call hello and pass it a parameter. And it might say hello Tim. Um and so that that's what we call the event command pattern. We will send an event to a remote um endpoint, and that event will call on a command, which is the hello um part of the API, that will process a response and then um send the response back to the caller. So obviously that's RPC, but it's event command because you're dispatching event to a command interface or you're calling against part of an API. Um the difference with the pure event-driven model is when you send an event, firstly, you don't know who's going to receive it. You send the event down a topic, and it might be the hello topic, and you might have 17 different languages of hello world processors sitting against them. But they don't necessarily expose an API, they consume events off of that hello topic. Um and the the event that's sent might not it won't be um submit, it might it won't be call hello, it will just be say hello. So it it emits a hello event. The hello world processors in all 17 different languages receive that event and then only known to them, they will process it with whatever language they need to in response. And if they need to send an event back to the caller, they will do, or they might send it downstream. But you've got this decoupling. Um, what this means is we can go and add other hello processors, we can add more languages, we can add other logic sits that sits on top of it, but the caller never actually needs to know who it's calling or what they're doing with it. So you have this separation of concerns. Um, and this model allows you to build large-scale decoupled systems that are purely event-driven where the events are captured, but the logic is also kind of encapsulated within the realms of the processor, which means you don't have this coupling between APIs. So the events themselves become a model of the domain. In this case, it's saying hello. Um, and the processes are basically reactions to those events. Right.
SPEAKER_01So the producer or the of the event, which historically we've called the caller, uh maintains a healthy disinterest in who might consume that event. Like I'm just I'm I it's it's just saying hello, and whoever wants to do something with that, those those services can exist or not, they can come and go, uh, and that just happens.
SPEAKER_00Yeah.
SPEAKER_01So that requires then um an API. That requires that we understand that this event in this log or this topic means something. Yeah. And specifically, I think that means we have to agree on how the data in there is formatted. So event uh events are the APIs is true and somewhat abstract. And if you haven't thought through it yet, you don't get it. But at a very low level, if you're still wondering what that means, that does mean we have to negotiate together uh what the fields in that thing are and what they mean. Right.
SPEAKER_00And normally that starts with a domain model. You know, how do I do domain-driven design, capture the domain through event storming and model those as events, which is the thing that's powerful about event-first thinking and events as APIs, because everyone speaks the same language. But the n the nice thing about that, that takes us nicely into the capability to drive serverless functions and where they fit and where they don't fit with respect to stream processes.
SPEAKER_01Before we get there, I I want to ask you I think we've kind of flirted with this notion, but how is this different from what we have been doing? So if I'm brand new to streaming and I've got, you know, I'm a conventional application developer using modern methods and tools and languages and blah, blah, blah. Um kind of differentiate this. Where is my mind going to break of what I've been doing?
SPEAKER_00Okay. Um so the way that we used to build systems, obviously it's changed since we had back when I first started in IT was Korba, common object request breaker architecture. And uh was quite a net quite kind of neat. Um we moved then to J2E where we had entity beans and session beans and and whatnot. Uh and then we then we were still modeling um APIs and things, and we eventually started to come around to the way of asynchronous distributed architectures could work with ESBs. And the problem with the ESB and messaging conventionally is that you kind of throw events through them and and then you're done. And the problem with that is, and and that this is where Kafka starts to change the way that we build systems, is once the event's in there, it's recorded. A bit like a database. Going back to, you know, this is a transaction log, this is the log, it doesn't get forgotten, it's kept there for the configured period of time, which might be three months or six months or something. But the main thing is because we start thinking of Kafka as a storage system for these events, it changes the way we build our systems. With ESBs, we didn't have this idea of long-term persistence of events, and we didn't think about events like we think about databases. The fundamental thing that's different about ESBs is that while it allowed us to interconnect decoupled system components and scale them, is the events that propagated through the ES through the ESBs or whatever messaging system we were using, we didn't necessarily want to keep. It was a turn it on, turn it off. There was no kind of real processing or logic around keeping these events. And that's where when we build streaming systems, is not only do we have this fundamental model of modelling events temporarily, temporarily. Temporarily, temporarily, temporarily, but more so that we can go back in time. And if we can go back in time, we can take a look at those events and we can reprocess them again and again and again, which is means that we think about and we build and we we uh operate Kafka brokers or street the streaming platform more like a combination of a message broker and a database. And we also come to regard the events in Kafka like those that we have in a database, in that we model them. We have APIs and we want those APIs to evolve. Um that's where we have Avro and things like that, which come from the big data world and they live in Kafka as your, you know, your core or your primary serialization storage mechanism. Right.
SPEAKER_01Okay, that's fascinating. That's a uh I also have an anti-ESB argument, and I have it because there's this one slide I use for the particular set of talks I give that feels like it's 2002 and I'm giving a pitch for an ESB. Just if you look at the slide, right? It's a sequence of two slides. Uh if you didn't know, then you'd think, oh, well, he's just selling an ESB. And you know, anybody of a certain age in the audience, uh like I do this at meetups, I have to ask them. They don't call me out on it right then, but oh yeah, they know that this looks suspicious. Like, wait a second, we've seen this before. You're just talking about an ESP. ESPs were terrible. I still walk with a limp because of what ESPs did to me or whatever, you know. Um, and so I've got my own anti-ESP argument that has more to do with the brittleness of what they did to organizations in pushing all of your spaghetti message routing into a opaque box. It cleaned up one diagram, but just pushed all the complexity into the box and put that box entirely in the responsibility of one often ill-tempered team. Um, and so like organizationally, they did some terrible things, and streaming platforms avoid that. But what you just said is, I think, even subtler, which is that, yeah, okay, they helped us think of messages and events and where events go, but because there was no persistence, they did nothing to challenge the paradigm that we had been using, which was relational databases are where we store stuff. And so you still store your, you you serialize your events and you put them in a place in a table, and ESBs came along, they're like, well, you want to route some events around, that's great, but uh your storage mechanism is the same thing. And with a modern-day streaming platform like Kafka or Confluent platform or any of these things, um you you uh that storage persistence mechanism is new, and that leads to a more profoundly different kind of system. So I'm just reflecting back your ESP critique fascinating.
SPEAKER_00So there's two aspects to this. One is you you know, ESPs will persist messages, but they won't persist them in uh a temporal fashion in the same way that you get with Kafka. And the the other aspect is which is often mentioned, is the ESB anti-pattern, where people build um logic that reacts to um events traveling over the ESP, and no one necessarily knows what it does or the consequences of changing it. Um so it gets lost, and you end up with a large-scale ball of spaghetti that you don't necessarily understand. And and people often say to me the same thing.
SPEAKER_01So it was terrible.
SPEAKER_00Yeah. And that and that in itself is a very good, subtle point there is the configuration, and I always go back to technology's moved on somewhat. Um, and I know Tim, you might still like to use XML a lot. Um for nostalgic, for nostalgic reasons. Um, but you know, that these days when we build and run stream processes, we're not manually deploying servers, or we shouldn't be, or processes. You know, these should be Dockerized processes which are done through automation, right? Which means we don't end up with these ESP anti-patterns because by the time you run them off the back of a CI CD pipeline, you've got your full test harness, um, your continuous delivery mechanism, and the whole thing is run in a safeguarded, automated fashion, which means you can completely avoid the whole ESB anti-pattern argument. Right, right.
SPEAKER_01Good.
unknownGood.
SPEAKER_01And it's not XML now, it's YAML, which is totally different.
SPEAKER_00Except so the other thing I heard the other day, which was quite interesting, um, is you know, YAML's great until you truncate a file. How do you know you don't have a truncated YAML file? Because there is no close. And it does happen. Right. Yeah, oh yeah.
SPEAKER_01So that's quite too long for the good old days.
SPEAKER_00Um XML.
SPEAKER_01Arguably not so good.
SPEAKER_00Yeah. Uh so shall we take this to cloud and serverless?
SPEAKER_01Exactly. That's where I was gonna uh ask you to go next. So please take it away.
SPEAKER_00So um so now, you know, to go back to baseline, you know, we're we're seeing large-scale distributed systems being built, which are data-driven applications, which we call event streaming applications, which are business critical, which means they need low latency, um, and we want them to scale elastically. How does this fit with the cloud? And why can't I just do all of this in FAS? Well, the answer to that is the typical IT answer is it depends. And I've got a blog coming out about this, which I'm going to plug, um, which goes through a lot of what we've discussed, but it also talks about the it depends aspect of functions as a service service or serverless functions, whatever you want to call them. So while FaZ is great in that I can I have a natural scale-out model, I don't have to manage these instances, they just I have to deploy them and they run. That's great. And if I can run them off the end of a FaZ connector, um, like the AWS Lambda connector or some of the other connectors that we've got for Confluent Cloud, um the question then is why don't I just do everything in FAS? Um, and the answer to that is twofold. Fazz is great for stateless operations, and it's great where you are aware of the fit. And when we talk about streaming and stream processing, we're normally characterizing applications which want to have throughputs in excess of 20,000 or 100,000 events per second. And when we start modeling a business system, then we need to understand the cost of end-to-end latency. So a typical stream processor might be processing up to 200,000 events a second. If we sent those 200,000 events into a FAS function or a serverless function, we're going to incur two things which have been quite openly discussed, and that is cold starts. Now, depending on the language, um your cold start could be one and a half seconds. Um, and then you've also got the cost of a hot invocation. So that's where you've made the initial invocation, which is a cold start, and the cloud provider is keeping it there. Uh, and someone like Amazon is smart in that they they bin pack your containerized serverless functions together on the same instance. Um and then the hot invocation, which is the subsequent invocation, you'll a lot of the time you get an invocation latency of 100 milliseconds. So if we were going to send a stream of events from a single IoT device where we want to guarantee the ordering through individual FAS functions and they're 100 milliseconds each, we're only going to achieve 10 events per second. Now, provided that that fits your workload requirement, then that's great. But for a lot of streaming applications, this is nowhere necessarily near what we're expecting to see.
SPEAKER_01And so it's like an even in the game, really. Yeah.
SPEAKER_00Exactly. And so that when we characterize a streaming application, there are three types of workloads to think about. One is what's my live streaming workload? You know, whether it fluctuates through the day. So if I have IoT devices coming on at nine o'clock in the morning and throughout the day, then they quiet down at the end of the day. What's going to be my peak? So 20,000, 100,000, that's great. That's my live streaming workload. Um, what is another thing that's not necessarily always considered when we build these systems is historic workloads. You know, we said Kafka allows you to go back in time and reprocess things. So maybe you have a machine learning um model which you want to calibrate or test or recalibrate, then you have a historic workload. Um, you might want to reprocess six months' worth of data, and you might want to reprocess that within 24 hours. And depending on the latency of your Lambda functions, then that may or may not be possible. Most it's unlikely that it would be possible. Um I'm gonna come come back to ways of solving it. And then you have elastic workloads, and these are you know your workloads that fluctuate according to demand. So, how do we consider the FAS functions in in the context of the streaming application? So if we model this, then we have an event. An event is stored within a stream, which is the temporal bucket or the temporal kind of mechanism where we we group these streams together by their relationship. So this might be an IoT device, a user session, or a shopping cart or something like that. And then all of these streams get stored within a topic. Now, if we want to reprocess some process something, then we need to understand what the latency throughput requirements are. And for business critical applications, if we can if those events don't occur too close together, oh yeah, there's lots of them, but they're not within the same stream, then we could potentially use FAS. Um, most of the streaming applications that we see, you generally use FAS for building thumbnails. So we're right back to the beginning. And so I've come up with some event-driven principles. Um the first one is they're great for in-band processing at the edge. And so an example of this is if I have an application and I have a user registration, then if I want to map that user um address or the GIP to a latlong cell or something like that, then do it on the way in because it's not business critical function. And that 100 milliseconds is going to be more than enough to validate the user, to do the latlong mapping, and then they actually make it into the core business system. So that's the first example, in band but edge. The second one is in band, but it's not latency sensitive and is stateless. And so this is the case I was talking about where you're just never going to have that many events within the same stream. Um, and this is quite rare. Um, the other one is in band, it's not latency sensitive, but you do want to do enrichment against external resources. So this might be an address lookup. You're doing it in band, you don't have that many of these events, so they're kind of low, low traffic. Um and the external resource might be you know a dynamo DB table or something. And the challenge of this is it needs to not be latency sensitive because if that FAS instance, so we with a FAS function, when you first build it, you get an opportunity to load state and you can periodically refresh that state against external resources. Um there's no guarantee that you're necessarily going to call against the same instance on subsequent invocations. So that's why another reason why it needs to be not latency sensitive. Um and then there's another one which is the third one which is out of band, but it's at the edge. So this might be an email notification. So maybe you purchased an item and it wants to email you to say, hey, it's going to be delivered tomorrow. Um, or do some post-processing analytic around BitOffice spread or something like that. And the final one are just ad hoc requests. So these are these are requests which are not time sensitive. They might be historic analysis, um, and they might be something like a Monte Carlo simulation, recalibrating machine learning models, but they're something that's not necessarily that time sensitive. And they're the five kind of FAS suitability principles that I kind of come up with as part of doing this kind of this thought process of understanding how FAS fits within the world of serverless and stream processing.
SPEAKER_01It strikes me that uh the domain of FAS suitability, and FAS is functions as a service, if you haven't heard that term before, but the domain of FAS suitability seems narrower than one might expect. Right? The hype is that, well, the economics of serverless are so compelling. Everybody is going to build everything in serverless in two years, and anything that doesn't fit, we'll just evolve our understandings of applications to make sure that uh we can build everything there because it's it's just the economics are dominant. Uh the latency, not just cold start latency, but but kind of the quiescent state latency is pretty high. Yeah and that that that matters a lot.
SPEAKER_00Yeah. Um I think the thing that we've seen that's quite interesting in so that the hype around FAS has started to subside. Um there's a ton of serverless conferences going on globally. I know a couple of people that um that organize them. Um and the thing that's most interesting about this is the hype's died and the dust is starting to settle, and people are starting to say, well, yes, it's cheap. But at the moment, if I handle have to handle guaranteed latency and I need low latency, then I need to think twice about whether that's suitable. And then I need to architect my application in a way that I can use FAS where it fits and not where it doesn't meet the requirements. And they're saying that it's naturally going to resolve itself over the next year or two.
SPEAKER_01With some maturity and people learning lessons and understanding, hey, here's a cheap way to do computation that fits into, in the nearly every schema, one of these five buckets. Yeah. And other kinds of computation. Well, you wouldn't do that because it's a super bad idea. Um, but it's it's a good way to save money on compute resources when you can.
SPEAKER_00Absolutely. I think it's got a great future ahead of it. Um, I think it'll fit nicely with us as we move forwards with you know the way we see our cloud technology evolving. So building FAS-based ETL pipelines which aren't time sensitive, they're more like housekeeping tasks. More build me a thumbnail. Um, maybe we can put little Tim Berglin uh logos on the thumbnails that we built.
SPEAKER_01I I think that's a fantastic product idea. And I don't know if we have a PM for serverless yet, but I'm gonna talk to the cloud PM as soon as we get done recording this and get that in the backlog. Uh we clearly need that. Uh but yeah, no, that you you can definitely see um and you know, we're both Confluent employees and we're talking about Confluent products, and so it's uh should be clear that we're just sort of speculating about the future, and this is sort of our you know, hashtag Safe Harbor. Uh we're not saying that any of this is a thing we build, but it makes sense if you look at a streaming cloud thing, um, the economics of being able to bin pack these little functions on compute resources when they fit in those criteria, you know, it it all of that goes together well.
SPEAKER_00The the subtlety here is um the cloud vendors must be loving this because it gives them orders of magnitude of efficiency of operation. It allows them to allocate, reallocate servers without actually taking them down or moving VMs or instances. It just happens naturally for them. So from a cloud vendor perspective, the operational efficiency is is insane. It's a massive upside for them.
SPEAKER_01And the consumer pricing is good too, but yeah uh it it's there's a lot of margin there for uh cloud vendors to be happy. Yes, increasingly so. Um and um a lot uh a lot to look forward to. So um any uh any final thoughts on this whole affair?
SPEAKER_00Um I think it's this is a a watch this space. Um we know it's going to evolve, and and we're doing uh more work in 2019 around FaZ connectors with Confluent Cloud. Um we see them as something that's necessary and something that we need to support because yes, you can use FaZ to interact with your cloud ecosystem in quite a flexible way. And you can also make them part of your your stream processing story where it meets the kind of the requirements um of your application. So there's definitely a place there. Um, and I think that's going to change significantly over the next two years.
SPEAKER_01My guest today has been Neil Avery. Neil, thanks for being a part of Streaming Audio.
SPEAKER_00Thank you, Tim.
SPEAKER_01And there you have it. I hope that was helpful to you. If you've got questions, you can ask me at TLbergland on Twitter. That's T-L-B-E-R-G-L-U-N-D. Or you can leave a comment on any of our YouTube videos. Your question might be featured on the next episode of Streaming Audio. And feel free to subscribe to our YouTube channel and this podcast wherever fine podcasts are sold. And if you subscribe through iTunes, be sure to leave us a review there. That helps other people discover the podcast and just generally helps us get the word out. We appreciate your support. See you next time.