Mark Saroufim

You know, at the time my thinking was like, okay, well, I really struggle to get people to care about my ideas for money. You know, but like, what if I could at least get them to care about it for like for free? And so that's kind of what got me into open source where I was like, okay, well, it's a good prerequisite for am I on the right track. So I got started with writing like an ebook called the Robot Overlord Manual, which is it's an interesting e-book in hindsight. It was like somewhere in between a robotics versus RL book versus math book. It's a very strange book. I don't think it's that good in hindsight, but I had a really good time writing it. And and uh and basically like a couple of those chapters like made it to the front page of Hacker News, and I think they all sort of had this sort of similar essence of the math was a bit different, but like the math helped motivate how to write the systems. One canonical example of this is like how like lie algebras can help you like represent rotations and robotics, and so that those are really fun for me, or using Lagrangian physics to sort of formulate physics and as an optimization problem that minimizes energy. So a lot of these problems for me were like very, very intellectually appealing.

Conor

Welcome to ADSP the podcast, episode 302, recorded on July 20th, 2026. My name is Connor, and today with my co-host Bryce, we interview Mark Seraphim. Mark is co-founder of both GPU mode and core auto, and in today's episode, we chat about his path into computing. It includes GraphCore, PyTorch, Meta, and more.

Bryce

I'm gonna have to take I'm gonna have to take off the B200 hat because hats are not super compatible with my uh with my hair. Well, we're here with Mark. I guess Mark, maybe you should intro yourself.

Mark Saroufim

Uh sure, yeah. Thanks for inviting me, Bryson. Thank you, Connor. Yeah, my name is Mark Sarufum. I've been working in like AI systems for like about like six years now. Most of it was spent like working on PyTorch, which is a pretty popular like framework for deep learning. But I guess like what I ended up being most known by was this like fairly large ML systems community called GPU mode, where we do things ranging from like GPU education to competitive GPU programming, and we have like a whole variety of like working groups. Most recently, like I'm no longer purely working in open source like I used to, but I helped co-found a new Neolab called Core Automation. Like our goal there is basically to build the world's like most automated lab and build the world's best model. And like sort of my research agenda there is like getting AI systems to be really good at writing, uh get getting like AIs to be really good at writing systems code. And I think this will be like fairly transformational, like really for everyone. So yeah, thank you guys, and I'm excited to be here.

Bryce

How did so what did you do before before you were working on AI systems? Like what did you do before PyTorch?

Mark Saroufim

It's a very good question. So yeah, I I I had like an interesting early career, like pretty bumpy. But the the way I got started in AI was like in AI theory. So I was a grad student at UC San Diego. I was fortunate to have like two amazing advisors with like Charles Elkin and Sandro Descupto. And they're like primarily, like I said, I was in AI theory. Like I taught like seminars on quantum computing. I was like trying to figure out if like you know, we could rewrite like axioms of probability theory using like game theory instead. And I had a really great time, but but but but basically, yeah, you know, at the at the at the time, like I didn't really sort of have like a good understanding of like how important AI would be. It was more like I just thought it was really cool. And I ended up like my first job out of school was uh was at Microsoft. So at the time, like I remember getting like the job offer about a year before I graduated, and so I just like spent a year doing math. But I joined as a product manager at Microsoft and I worked on like a whole bunch of like wacky things. Like I, you know, my first project was selling Microsoft's display ads business to AOL, which was like an interesting first project. SOC3U compliance as well for Outlook. It's like very exciting stuff. Wait, wait, what year were you selling things to AOL? So this was uh 20 like 2014, 2015, roughly. This is this is what I when I first started my career. But yeah, I mean it was a weird project because I remember at the time being like fairly AI pilled and like thinking like, okay, I don't know how like how big of a deal this will be, but uh, but I did get a sense of like this is a really important problem. But at the time, I would say there was this widespread view in the industry that this was sort of like like a luxury on top, like as in like it's sort of a further down the tech tree because what you needed to do is to get like big data, which was proprietary company data. And once you accumulated enough of it, the hypothesis was that like you'd be able to build very smart systems. But but funnily enough, this turned out to be incorrect because like private data turned out to be less valuable than public data, which was like very, I think, very strange conclusion. But but it was a reality, and so like it was really only in my last year and a half at Microsoft that I worked as a research scientist. So there I was like working mostly on like email intelligence. So think things like uh, you know, the AI systems that would figure out that this is a flight, and then given that it's a flight, correctly parsed information out of it. These systems had to be fairly conservative because if you got someone's flight information wrong, they'd be really angry. And so uh, but but at the time, like I mean, a lot of these were like fast rules-based systems, like more in the traditional AI. I myself had worked on like the first like deep learning system to to to to do this like contact extraction and email. But at the time it was considered too heavyweight because the model that I built was like around like 50 megs, and so that was considered at the time like unrealistic to deploy in prod. And primarily because like there was again this thesis at the time that like you would want like a small model to put on the client directly, that it wouldn't really make sense to have a central one because maybe people's preferences would be different. But it was interesting to me, like basically we were sort of correct in formulating the problem, but because we had certain ideological beliefs, like one about private data and two about local models, I would say that prevented us from building like really like a state-of-the-art model that people that people could use. So, yeah, it's it's a longer story, but that was uh the Microsoft days. And let me know if you want me to keep going, if you if you want to interrupt me here, because like the there's like maybe three for three more years basically worth of stuff that I can do.

Bryce

Yeah, yeah. I mean, that's keep going, keep going.

Mark Saroufim

I'll keep going. Okay. So after Microsoft, I remember at the time feeling like it was around the time when the Dota AI bots came out. And so for me, this was uh like this was kind of like in the same way I see a lot of younger folks feeling the same way about LMs today. Like that was sort of like my feel the AGI moment. Primarily because like before I got started and working on software, like I got started in software because of video games. Like my sort of like my first experience using a computer to do work was uh pirating Warcraft 3 and reselling it to my schoolmates. So that was kind of how that was my introduction to computers. I made a decent amount of money like doing this, but but yeah, like basically for for me, what what always felt like the most magical was like you see like an agent like doing things on a screen, and to me that always seemed like a like the real way to like look at the like because you could it felt like tangible in a way that like I don't know, grammar-based models or like parsing trees like never never really did. And so so the the my thinking then was that like that this work came out and I was like, okay, well, this seems really transformational. Like, what if we could do this for all of video games and we could have like really sophisticated game AI, and of course, like I myself had spent a lot of time playing games like StarCraft and Dota. You know, I was a competitive chess player in school, and so like I just sort of quit my job, and I'm like, okay, I'll figure this out, which turned out to be not like a very wise career choice for like a variety of reasons. One is like I had no idea how to sell, I had no idea how to market my work, I wasn't even that good of an engineer back then, and so like it was like struggling on all sorts of like skill issues. And most notably, and I think what ended up being like a really critical flaw of this work was that like it was like really expensive to actually get anything working. And so, like, here I was like basically running on the equivalent of like a 1080 Ti. And what I would notice is like running sort of simple experiments within Unity, which was like, you know, I would sort of change whether an agent is standing at like this coordinate system or like increment the X coordinate by just a bit, and like nothing would converge. And I was like very confused, like these techniques felt really brittle to me, and it wasn't really clear how to how to how to get out of it. And at the time I remember, you know, I was like mostly going off of like my savings from Microsoft days, and then COVID hit. Uh so very quickly, like my sort of money that was like in stocks was like evaporating towards zero, and of course, I wasn't making income either. So I really, really needed a job, and then I got a couple of them. A lot of them got rescinded because of COVID. Yeah, so it was like very, like, very, very like a weird time. And so, like, you know, at the time my thinking was like, okay, well, I really struggle to get people to care about my ideas for money, you know. But like, what if I could at least get them to care about it for like for free? And so that's kind of what got me into open source where I was like, okay, well, it's a good prerequisite for am I on the right track. So I got started with writing like an e-book called the Robot Overlord Manual, which is it's an interesting e-book in hindsight. It was like somewhere in between a robotics versus RL book versus math book. It's a very strange book. I don't think it's that good in hindsight, but I had a really good time writing it. And and uh and basically, like a couple of those chapters like made it to the front page of Hacker News, and I think they all sort of had this sort of similar essence of the math was a bit different, but like the math helped motivate how to write the systems. One canonical example of this is like how like lie algebras can help you like represent rotations and robotics, and so that those are really fun for me, or using Lagrangian physics to sort of formulate physics and as an optimization problem that minimized energy. So a lot of these problems for me were like very, very intellectually appealing. And yeah, and and I guess eventually I started applying, and like the only company that got back to me was this company called GraphCore, which was an old, I guess, now defunct competitor of Nvidia Core for 11 months, yes. It was a fun story as well. So GraphCore was great. I had like amazing colleagues, but basically the the sort of just like walking into it was we know we would basically go to these customers and we'd ask them, okay, well, we need to be deploying like you know, like do you want to buy our chip? And they'd be like, Well, does it make such and such model faster? And we'd run a whole bunch of experiments, and then you know, turns out with like a whole bunch of work, yeah, yeah, often the answer was like yes. The problem was that like it was like there was like two layers of problems. One was like NVIDIA wasn't a static target, like basically we would do all this work, and then like a year later, NVIDIA would just like double all the numbers we'd need to hit. That always like really sucked. Uh, and then the second thing was like the the the world basically changed. Like, I think GraphCore's like thesis at the time was that so it was like a small SRAM chip, basically, with like 300 megs and then later 900 megs. The idea was like if you can get a model to fit, it's incredible. But if it doesn't fit, you're sort of going distributed by default. And unfortunately, like at the time, like the BERT-based and large language models were quite large, and so it just like wasn't like a natural fit. We had to introduce like distributed programming like much earlier than the community basically was ready to understand these ideas. And so that that was hard. And but funnily enough, around that time basically was my first I remember at the time writing like a like a like a very popular post called The Great Stagnation of Machine Learning. I don't know if you might have read it. That post kind of put me on the map at the time because it was like sort of this like the very off-the-cuff like rant about how like most academic AI is like basically morally bankrupt because it sort of formulates itself as like this like innovative field that's like risk-taking, but a lot of it were sort of simple ideas like on scaling. And so, like, I think it was well written. I had a couple of like very catchy phrases in it, like graduate student descent, and there I was like attacking the auto research community, which is now defunct, but I was also attacking the optimizer community, which is turns out was defunct for a long time, but like later had the introductions of like Muan and Shampoo. Basically, they sort of salvaged themselves. And and it was like pretty interesting because like right after I wrote this, I remember sort of getting into trouble at work because like I was like also talking about like GraphCore about this. It was like it was like this very volatile post. But you know, I remember over like Christmas break, basically, there was like a job opening at PyTorch, and I had started using PyTorch myself at work, and and my framing there was that like you know, but like I started with Piano basically, you know, like I like that was sort of my first like deep learning framework, and I remember like none of them felt like a joy to use because it was just I had an idea, and then I would get bogged down with how to implement it. But at the time with PyTorch, I remember uh I was like reading the UNET paper. With the UNET paper is it's like architecture for like vision encoders, and it has like this like very interesting structure where like the encoders just get smaller and smaller, and then they up until a point, then they get bigger and bigger. And like I looked at this and it was like this fairly large architecture diagram, and I remember just reading the paper and then just like typing out the PyTorch code without any reference, and then it just like worked, and I'm like, oh my god, like this thing is insane! Like, like this, like what is this alien technology? And I felt like I had to work on PyTorch, and you know, I applied, and I got very lucky that I passed interviews because I think a good chunk of my interviews, people just want to talk to me about my blogs, which was like pretty funny. And yeah, and PyTorch was like one of the best places I've ever worked, that's why I ended up staying there for like a whole five years. And yeah, I mean, I I I just think for me, sort of like it taught me a lot about like building things that people want. I felt like the engineering culture of the team was quite strong, but like even which which was maybe rare, is like I also thought the product culture was quite strong. Like, people really understood what is it that researchers wanted, and they had a clear idea of like their customer base. This would change over time, but at least like in the early days, I felt like it this was like one of the like the best places in the world.

Bryce

Yeah, it's so fascinating to me the arc of how PyTorch sort of won as the dominant, you know, ML framework. Because like I remember when you know TensorFlow was the the the thing that we used for production, the big serious thing, you know, it's in C, it's fast, it's got this, you know, you know, this declarative model. And it, you know, it I think at the time it seemed like this kind of like crazy idea of like, well, just write it all in Python and you know be easy and simple and P that people would never do that at scale, but like that does like how do how did we get to the point where I mean if I'm if I understand correctly, like most of the big labs, or at least some of the big labs, do train in PyTorch at scale. Like, how did we how did we get there?

Mark Saroufim

Yeah, it it it it it's strange, and it's like something I've I've thought a lot about for what it's worth. But but basically, like my sense is that like and this is like something I felt like a lot of the maybe hardware vendors, with maybe the exception of NVIDIA, never truly internalized, which is that like if you ask a customer like what they want, they'll often tell you I want performance because it's like an easy way of like you put two bars on a slide and then you're like, Oh, I picked the bigger bar, therefore my decision is wise. But then, like, often like you'd realize, like, okay, well, people want performance, but like they also want print statements and like they want like dynamic control flow and they want like ragged shapes, and they want sparsity, and they want all these things, and and this is like very like sort of weird for a performance person to understand generally, is but and it's obvious in hindsight, but at the time I th I think people thought it was like generally weird. And so, like the way I think of it is like sort of our job as systems people today, and again, again, I'm a sort of a newer systems folks person than you are, Bryce, but like basically I sort of view my job as being at the service of AI researchers. Like, it's not to tell them you're you're you're the way you're doing things is incorrect, it's basically, well, they need to express their ideas and explore them, and it's like my job to figure out how to make this like class of people like productive. And turns out, like, it's not you write a training framework in like raw CUDA and it's and then you expect people to understand the all the all the details of the CUDA manual. And for what it's worth, like I think the reason why, like I would say this is no longer true today, to be clear, but I think the reason why this was probably true for a long time is because like memory bandwidth relative to flops, like there wasn't a sort of like this big dichotomy, and so you didn't really like fusion compilers were important, but they weren't like absolutely critical as they are today. And so like PyTorch basically hit a sweet spot where they basically figured out that like a certain trade-off didn't really matter for a long time, and then they figured out how to convert that into like a product advantage, and basically all the papers are written in PyTorch. And again, in a world pre-LLMs where there was like a big variety in the kinds of ideas people were exploring. And remember, there was things like nerfs and GANs and language models and image models and diffusion and like a whole bunch of like ideas that have been lost to time today. Like these wouldn't be possible if like every time a researcher needed to do any work, they needed to recruit like a research engineer to do their job. And I think there's like a deeply unpleasant way of working, which is like uh, you know, you go to the research engineer and then they're like, oh, grumble, grumble, like no, you shouldn't do this, and like this is a bad idea, and don't you spread and you know if that's not what you want to hear, actually. Yeah.

Bryce

I think that to some degree, like if if if the research had stopped, if we had figured out the answer in you know, 2019 or 2020, something like TensorFlow would have won. But the way that this industry played out is like the research continued, and that meant that the the people doing the work were researchers and the work being done was exploration, and I think to to some degree continues to be today. And so it was necessary to be able to move fast. Like I I I don't know. Maybe it maybe it's maybe when people are in a tech revolution, it always feels like this, but it feels like the industry moves faster now than it ever has in the past. And maybe people who are around during the internet revolution will tell us that it was like that back then, but it really does feel like, you know, like every every week, every month, you know, there's some big new development. I mean, I obviously there's some things have solidified, you know, it seems like Transformers just take over most of the world, but there's still, you know, the the idea of the researcher remains. It's not the production engineer, but it's the the researcher, and we're gonna explore a bunch of different models, and you know, that the the productivity continues to be key.

Mark Saroufim

I mean, yeah, like for for what it's worth, like I I have heard people say things like, oh, like since LLMs, this has sort of like sucked the oxygen out from a performance engineer's like perspective. But actually, like I think even the label of transformer is like quite a broad one as far as I'm concerned when people use it. Like basically, you know, is a transformer with DeltaNet like Kimmy, like is this a transformer or is this like an RNN?

Bryce

Have you seen like a alignment chart? Somebody tweeted the alignment.

Mark Saroufim

Yeah, that's exactly what I was thinking about. Yeah, and so when I look at this alignment chart, like what I'm thinking is like, okay, well, we found like an architecture that's really good, and then it's only natural that like people want to find some that are even better because like the amount of flops we're putting into models are remarkable. Like, yes, you can always get more GPUs, but there are sort of like economic limits, like you can't get seven trillion dollar, like a clutch you can't get a seven trillion dollar cluster easily, and there are like real bottlenecks in the system beyond just money. And so, like, I continue to feel like there's there's a big place for like architecture research. And I think for architecture research, I mean both like the kinds of models you pick or like variants of it. So think of things like a you know, a transformer with residual connections or like with uh kind of like the the the yeah, so like is this a transformer? Like, yeah, kinda like, and so what are the trade-offs there? And then you know, it turns out like those do have real trade-offs. Like, for example, you you use like a lot of VRAM for activations, you know. Turns out that doesn't matter for training because you need to save them for your backwards pass, anyways. But it really matters for inference where and now they use up a lot more memory. So, like, I uh like it's and I've realized this, like, especially since joining Core Auto, that like there is sort of a like a big spectrum of like ideas that people want to explore. And then every time the research team sort of comes to me with like a specific idea, or when we look at like papers and open source as to what people are doing, they have very subtle trade-offs for what makes them better, like worse or better for inference and and on on what on what dimensions. Uh but yeah, I mean I think ultimately the goal is capability, and so like this is why like you know, unfortunately, as systems guys, we have to accept that we're like second class because we want to see their markable capabilities, and then from that we decide how can we best make this fast. Because if it was up to us, how do we design architectures? We would just like, for example, send the same like weight matrix to the GPU, it would be as big as possible, and then we would just do like a mat mole versus itself over and over again, and there'd be no non-linearities, there'd just be like one giant mat mole that you do infinitely and we get AGI. Like, if researchers came to me with this, I'd be like, this is amazing. Unfortunately, like such a model would not converge or be useful. So like that's why I think inherently our opinions matter like a bit less than we would like.

Conor

And when you say you're on the PyTorch team, was that at Meta formerly Facebook at the time, or was you was that like contributing to open source and you were working elsewhere?

Mark Saroufim

Um yeah, so so so in my case, like I started as a user of PyTorch when I was at GraphCore and mostly spent a bit of time on the forums and stuff. I'd never sent a serious PR to PyTorch before joining the team. So when I joined it, it was like for five years, and I worked indeed, I worked on the on the PyTorch team, and at the time the company was still called Facebook. Yeah. And at the time as well, like it was sort of a like again, it's all funny for me to think through this in hindsight, but like it was pretty peak basically because it was like it was around the same time like PyTorch was sort of like this like dominant programming language where everyone loved. It was my first time working on a product that people adored, as opposed to a product that we're trying to get people to adore, which is like very different psychologically. And then Llama 2 was also like a big revolution because it was like also like uh it's like Llama 1 was a good model, but its license was very restrictive. Llama 2 had a good license, and so like now there was this community building on top of it, then you had things like local llama and like unsloth, and local alums were actually like becoming good, and PyTorch was sort of like the substrate behind that. And so I think for for us for like those like two, three years, like it it felt incredible because you really felt like you're like the good guys are winning in some sense. Yeah, and then the work was cutting edge, and that was like I think what was some of the the most rewarding time actually to work in PyTorch was like around this this period.

Conor

Yeah, did you overlap with uh Adam them or had he moved on to other things?

Mark Saroufim

Adam almost never worked at Meta, as far as I know, because he only worked at he only worked as an intern, which is like a funny story. So I like I I I met Adam for the first time last year. So like because he lives in Poland and uh Poland or Germany, I forget, but like he doesn't fly to the US very often. He's a bit of a hermit, but like one of the smartest people I've ever met, so I kind of understand. I kind of understand his perspective. But yeah, Adam basically was like the main person who wrote most of the code for the first prototype of PyTorch. But then later I would say like the project was mostly scaled up by like Sumit in terms of community and a Diank in terms of like the code, like basically CI and the code and and all the sort of like more hardcore features. And then of course, like the project ended up like, I would say there's maybe like 30 or 40 people that I think were very critical to the project's success. But like, you know, when I think of the founders, like there's basically, yeah, like there's like Greg Shannon, there's like there, there's there's there's Sumit, there's Adam, and then there's Ed. For me, these would be cons like these are the people who I would consider it to be the OGs. And there were, of course, like others, like you might recognize, like, you know, Natalia, who used to work at Nvidia, who did like the most of our like CUDA performance work, Horace and Jason Ansel, who did like the majority of our like newer compiler work with Torch Compile. So many, like, many great engineers like touch the project at various points in time. But yeah, I mean, like the the Adam and Sumit did something more remarkable, which is like create something from nothing. And I always like have a lot more respect for that because I think scaling a project in some sense is easier. It's hard, but like I think creating something from scratch is not something that like I don't know, it's not something I take for granted or that I necessarily know how to reproduce really well. Right, right.

Conor

Be sure to check these show notes either in your podcast app or at adspthepodcast.com for links to anything we mentioned in today's episode, as well as a link to a GitHub discussion where you can leave thoughts, comments, and questions. Thanks for listening. We hope you enjoyed and have a great day.

Bryce

Low quality, high quality, that is the tagline of our podcast.

Conor

That's not the tagline. Our tagline is chaos with sprinkles of information. Alright, and both and before we get started, you might have to do a lot of the heavy lifting, Bryce. My Wi-Fi is fine, but my deco mesh node keeps on like turning off, like failing every 15 minutes.

Bryce

What is a deco mesh node? It's a very fancy Wi-Fi. I have the same idea.

Conor

Yeah, it's just like it's not the best, but it's up there. And it's just like if you want to set up a mesh network, which means like, you know, a lot of people, if they're in a condo, you can just have your single rotor, that's fine. But if you have like a bigger space and you need multiple nodes, you can either like pay your Wi-Fi provider an extra like five dollars per little Wi-Fi extender, but I was on principle never gonna pay these companies more money unless if it's for faster internet. And so you can just buy these like I don't even know what the correct technical term for it. Like, do you know, Mark?

Bryce

Uh like if you live if you live in a New York pro in a New York apartment, you don't have any of these problems. You just they come in, they put in one Wi-Fi router, and it's always fast, except for blast.

Mark Saroufim

They're pretty amazing by the way. I looked into the stack a while back, but they they do things like they try to predict, for example, like like they can like measure which parts of your house like need the most bandwidth, and they will like just route more bandwidth towards it. So each of these routers is itself like this like interesting like statistical AI node. It's pretty like remarkable. I mean, I I get like 800 megs per second on Wi-Fi speeds because of it, and before it I was getting 30.

Conor

Wow.

Bryce

Okay, maybe I do need one of those.

Mark Saroufim

It's it's it's great. Like, I mean, if you play like especially like when I first set it up, I was just like download games and just like uninstall them for fun. You know, it was just like let's just like let's just waste capacity.

Conor

Anyways, all right, now we'll do uh we'll keep that we'll keep that in the edits, maybe uh, but we'll I'll throw it over to Bryce, who can do introductions here. But anyways, if I if I am silent for minutes on end to the listener now, you know it's because my my little node is acting up. Over to Bryce.