Benchtalks
Benchtalks is Snorkel AI's podcast series at the intersection of AI evaluation, data quality, and real-world impact. Hosted by the Snorkel team, each episode brings together researchers, practitioners, and leaders to dig into the questions that matter most as AI benchmarks grow more sophisticated, dynamic, and reflective of the complexity found in real-world deployments.
We explore the full stack of what it takes to build AI that actually works — from the design of rigorous, open benchmarks that close the gap between what we measure and what we encounter in production, to the expert-in-the-loop data creation and curation pipelines that make reliable evaluation possible. Along the way, we get into reinforcement learning, reward modeling, and the evolving science of data quality that underpins it all.
Whether you're building agents that operate over long horizons, crafting rubrics that go beyond pass/fail, or trying to understand what "good" looks like for a multi-artifact deliverable — this is the conversation for you.
New episodes drop regularly on YouTube and wherever you get your podcasts. Follow us @SnorkelAI (LinkedIn, X, YouTube) to stay current as the field moves fast.
Benchtalks
Benchtalks #2: John Yang (SWE-bench, ProgramBench) - The future of coding benchmarks
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
For our second Benchtalks, the series dedicated to the researchers building the measurement toolkits that frontier labs hill-climb on, Snorkel AI co-founder Vincent Sunn Chen sat down with John Yang, a Stanford PhD student and creator of the SWE-bench franchise, SWE-smith, CodeClash, and most recently ProgramBench.
This interview covers:
- Why every frontier model scored 0% at launch — until GPT-5.5 cracked the first task (cmatrix)
- Why ProgramBench grades the runnable artifact, not the implementation — and lets models build in any language
- The post-training tell hiding in plain sight: how much models (especially GPT) love Python, even when it looks like a handicap
- The reward-hacking problem — models with internet access cheated up to 36% of the time, and why nine LLM judges still couldn't agree on what counts as cheating
- The lineage from SWE-bench to SWE-smith to CodeClash, and what ProgramBench needs from expert contributors to grow
Full interview/transcript: https://snorkel.ai/blog/benchtalks-john-yang-programbench/
Honestly, when SWE-bench was released, I got some comments from I supportive, wonderful people, but at the release they were like, oh, this seems impossible. I don't know if models will ever be able to do GitHub, you know? And funnily enough, I think you get the same wave of comments think this kind of skepticism I think is incredibly helpful. And I think it's almost healthy. But to me, I think the reason I've gone in the direction that I And I think the reason I say why it's a big pivot now is I think we've very very much entering this realm of before you could sort of find things where it's like, well, you know, Vincent and and humans in general can do this task in models. And now, I mean, I can't implement FFmpeg from scratch in one And I think that's a gear shift that's meaningful. Yeah, it's challenging to get right, and I'm sure we'll get it have faith in that direct.
Narrator:Welcome to Benchtalks, Snorkel AI's podcast on benchmarks, quality, and real-world impact. Today, Snorkel co-founder Vincent Sunn Chen sits down with Yang, the creator of iconic software engineering the SWE-bench franchise, SWE-smith, CodeClash, and most ProgramBench.
Vincent:I'm here with John Yang. John is a PhD student at Stanford and creator of our favorite software engineering benchmarks, including the SWE-bench SuiteSmith, CodeClash, and most recently ProgramBench. we're super excited to have you. Welcome to Benchtalks, John.
John:Thanks so much for having me.
Vincent:Last week you dropped ProgramBench. this was a really exciting new benchmark focused on end-to-end program level output, similar to Carlini's work on the C When you first launched it, every Frontier model had 0% pass I guess just this week, you know, GPT cracked that, you know, 0% what has the release been like? what what has the last week in general been like for you?
John:Yeah, it's been a fun time. I think a lot of people were quite excited because in many ways, works that kind of have investigated this problem at a case And to some extent, I feel like it is sort of a natural to sort of the SWE-bench setting when you have sort of this just here's a GitHub issue, and then you create a PR for it. And that was meaningful for two years. But I think now, now that these models are so good at solving together a large application? And so, you know, Nicholas Carlini and Anthropic's blog on the inspiration. there's a blog post from Cursor that was talking about kind using to put together a browser. and then Epoch AI along with METR. they had RE-Bench where they did this case study on four that sort of captured a lot of the things that we were thinking So that was super exciting. to kind of TLDR, I feel like the role of ProgramBench is really to add sort of enough task instances such that we can study domain with meaningful statistical power. Because I think when you have these kind of one-offs, these don't necessarily transfer. People use different things. There's certainly sort of different incentives that people So I think as with all benchmark builders, the hope is like this kind of equal ground, equal footing where we can study task, you know, with purpose and sort of with a lot of clarity.
Vincent:Yeah, no, it's super exciting launch and overall kind of for where these software engineering style tasks are going. Talk to me a little bit more about the style of the eval. in this case, you're evaluating the artifact, not just you know, you're up-leveling the interface and Help me understand that a little bit more. What is the what is the goal with that specific task
John:Yeah, I really love the description you just gave. I think I couldn't have said it better. one of the key things, I think, for me was there have generation benchmarks that existed before. so there's like really great work that Sasha Rush and his student which was released like close to around when SWE-bench was and so that paper was sort of a very good reference point and also there was we're gonna generate a Python library from scratch. They, you know, took I don't know, something like marshmallow or implementation, but they still kept kind of the function headers class headers and everything. And the way they evaluated was the models effectively asked to use the existing unit test suite. I think sort of the big thing I wanted to do was how do I just the model? Meaning that, you know, with CommitZero, you're still imposing a dictates what language the model should be implementing in, you what are the classes and the functions and the relationships and the typing. And so it's certainly not everything, but it's a very good of like system design and language selection and how do you about decomposing modules? And I think these are all things that we want to meaningfully So the point, sort of, sort of the small eureka moment was when I and he said, Oh, you know, like it's not so much the it's really the artifact, the deliverable that it creates. And so that was when we kind of made this decision of, all right, let's just completely put aside implementation as sort of a principle for the benchmark, and let's just do everything around the artifact. Or, you know, that's why it's called ProgramBench, the program. but specifically in this case, we do kind of executables and
Vincent:Yeah, I find that super interesting in your characterization know, the actual ways that coding agents are working. I think in the paper you had a few interesting findings, at to me, you know, there were flatter directories in general. I find it funny that it brings back up the kind of monolith, know, type of debate in software engineering. interestingly, you know, we saw GPT one-shot, you know, almost What were some of the surprising findings on your end of and more end-to-end types of problems?
John:One I particularly love was that so in the default inference for ProgramBench when model solve it, the key constraints are you know, we don't allow them to use Ghidra or these kind of like engineering of the binary. So those are the two key constraints amongst the smaller but we do allow the model to implement its solution in any language. So instinctively, you might think that, oh, if the model, it's written in Rust and the model chooses Rust to kind of re-implement solution, it must have a bit of a leg up, you know. But what we found is that's not always the case. So first and foremost, models don't just kind of pattern match and FFmpeg. I'm definitely going to implement it in C. we have this kind of nice language confusion matrix in the shows this. I think the key takeaway that was quite fun for me was that models And I think this that that artifact of post-training is very, But what's more so is with GPT 5.5, which was the model that you mentioned that solved the very first task instance out of the 200 That that cmatrix animation was not written in Python, but the Python. And if you look at sort of the proportions, so I think the Claude models tend to kind of be more like sometimes I'll use Go or Rust, rarely C or C++ for anybody, which I think is interesting. I'm not sure how to interpret that takeaway, but yeah, I mean GPT Python, which initially you think maybe it's shooting itself in much not. So yeah, that's fascinating.
Vincent:I think one higher level reflection I've had, I mean, the headliner here is hey, you've built this really effective, you know, first of its kind benchmark, right, that that had a lot of headroom, right? The 0% you know, frontier model score at launch. but I will I really see this all as almost a research tool, We're able to ask a lot more questions about how to fuzz or you know, a lot of these repos as we're talking about. We're able to ask questions about, you know, cost, cost versus quality trade-offs, you know, in terms of implementation, different types of tasks, right? You know, how does performance different differ from features? how how do you think about this as a research tool or something that you know the community should be building on top of more broadly?
John:Yeah, for sure. I mean, I think in the same way that when I read the cursor blog gave me like so much insight and information. we do 200, and I think that's the right way to go to sort of scale But I think even just the experience of like looking at what do just for this task instance, and you run multiple scaffold to mini sweet agent, right? which is sort of the justification being that this is the SWE-bench era. Right. And now we want to stress test it. And it's actually good, I think, that it starts a conversation scaffold can be better, right? But sorry, to answer your question very directly, I would just instance, like over-indexing on that, just looking at sort of this model does this, like what does the actual trajectory look What proportion of time is it spending actually probing the How often is it writing it? That's kind of where that insight of like, you know, cloud is probe a little bit, implement a little bit, right. And then kind of the next round of probing is to understand of what is the gap in what I have in progress in the reference. Versus GPT says, oh, you know, let me just do like a really job, right? Flesh out this very sort of thorough specification for myself. Right. And then let me one-shot it against the specification. And so I think it's just the tip of the iceberg in terms of the, you know, the way we can understand these models as what we want to garner and how we want to affect this kind of
Vincent:So what is your view on how these benchmarks should be used by, in dictating, you know, the changes in the field or or how folks are going to be approaching this discipline moving forward with in the loop, of course. Yeah, definitely.
John:I think something that a lot of prior benchmarks, especially Bench, I think has done a really fantastic job of is, you community effort, right? Like I think this tagline that you can add your task instance to something that the frontier is evaluated on. big credit to Mike and Alex and Ludwig and obviously the folks So the way I kind of hope this message carries over, and certainly I'll try to operationalize it, is that yeah, whatever your favorite command line tool is, put that in a ProgramBench. You know, we'll we'll open up a leaderboard, but also I think add tasks. and I think just the whole pipeline of understanding how to for it, right? How to actually sort of, whether it's a prompting problem or training problem, understanding the failure points point it at recreating sort of or developing different programs. I just think there's a lot of general takeaways that will be sort but then also things that you can really that are really valuable in terms of the nuances of like, wow, for this application, you cmatrix, which is this kind of nice terminal matrix animation. Yeah. And you can see, wow, it actually deals really well with kind and I think that's a surprising finding that I wouldn't
Vincent:I'd love to talk a little bit about the lineage you know, the future of it. Yeah. you know, walk me through how you got here, right? Like a lot of the band is back together from ProgramBench as But you know, you started with SWE-bench, the SWE-bench of data sets. the community kind of really took that on and stride, and date. You know, you did a bunch of work on SWE-smith that we're synthesize tasks against environments, tried different well, the tournament style eval, and most recently kind of What does that journey in lineage look like for you and where do you see that going?
John:Oh man, a lot of luck for sure, and a lot of thanks to the I've been able to collaborate with. But yeah, I guess if I were to construct a narrative in hindsight, yeah, the this journey kind of started back in when I was a and the first person I had really worked with, who I think I did ReAct and Tree of Thoughts and a lot of this really fundamental and so I guess to answer this question in terms of sort of, you lineage and where it's going and kind of the things that still fundamental principles. I think the fundamental things are that, you know, there's rigor, reproducibility. I think this is something that when I collaborated with Carlos on together the task instances and finding those kind of the 2,294 task instances. So finding those PRs and issues, it wasn't you know difficult We kind of got together in like June 2023 and we just kind of one day. Right. And we spent like the next month and we had it. Really, what ended up taking us the next three or four months was sounds like amazing to people now, right? But when we released the SWE-bench originally, we didn't use that ended up being kind of a nightmare for reproducibility. this is totally my fault, like because Carlos was like, maybe we I was like, oh, but it's so heavy, there's so much, like you to download gigabytes of images. Let's just do conda, it's gonna be fine. Yeah, he was 100% right at the end of the day. so that was kind of a driving theme. And I think like Kilian and I invested so much time in making sure people could download it and run it. So that's kind of sort of what's remained the same. I think what has remained different is to be willing to our notions of what language models can do. And I think ProgramBench in particular to me represents a fairly big pivot. Right. So the very, very kind of first project I did was called Inner And this was just kind of taking a lot of like NL to bash, and MBPP and HumanEval, and these the a lot of these code benchmarks be sort of input-output, no interaction at all, right, and it into an interactive format. Right. And so the question there was can the model benefit from this? With SWE-bench, it was like, look, this stuff is not realistic, passing LeetCode questions, but this is not what we do day to Can models do this? And right honestly, when SWE-bench was released, I got some from I won't name who they're all very supportive, wonderful But at the release, they were like, oh, this seems impossible. I don't know if models will ever be able to do GitHub, you know. And funnily enough, I think you get the same wave of comments think this kind of skepticism I think is incredibly helpful. And I think it's almost healthy. But to me, I think the reason I've gone in the direction that I And I think the reason I say why it's a big pivot now is right, I realm of before you could sort of find things where it's like, and John and humans in general can do this task, can models. And now, I mean, I can't implement FFmpeg from scratch in one And I think that's a gear shift that's meaningful. Yeah, it's challenging to get right, and I'm sure we'll get it have faith in that direction. Yeah.
Vincent:I mean I completely agree. I think one of the things I find really exciting about the you and other, you know, benchmark builders have taken is for where the where the field is going, right? You're putting forward these challenges, you're you're a path where, hey, this is where we think capability should or or can go. And you know, the skepticism is often a good sign that orthogonal to what people believe is possible. But I think that ethos is really, really inspiring and an way, you know, for all people building benchmarks to you think about their work.
John:Definitely, definitely. Yeah.
Vincent:I'd love to talk also a bit about the methodological differences of the benchmarks that you've worked on. You've explored a number of different ways to grade model from unit tests to you know tournament style eval, now to kind of this awesome fuzzing and you know output-driven validation approach. what have you learned about how to grade these models? What are the trade-offs between them?
John:how do you think about this now? The evolution of verification itself, I think, is a very problem. And in a lot of ways, kind of drives the way I think about it. And I think a lot of people think about it. So I guess with SWE-bench, I think the verification, it didn't do MBPP, HumanEval apps, like all these kind of LeetCode style things, they also had unit tests as well. The change there was like, okay, with HumanEval, even in the actually have people write the tests. And right, Carlos and I were like, we don't, we're not open AI, How do we look for tests in the open, in the wild? Right. You know, so I think that was kind of the key thing there. Otherwise, from like a formulation standpoint, I think it why I'm sure it got adopted. With CodeClash, I think I was very kind of inspired by the Ella people have done. Right. and I thought, okay, what does it mean to kind of bring this into a There were some prior works that had done things like, you know, Copilot Arena or even in the products with VS Code and Cursor some voting mechanism of proposed version A, proposed version Right. there was something where I felt like that made sense, but I scale because like how many times are you gonna be able to ask A or B, where it's like two different code snippets? Right. Before they kind of throw up their hands and they say, I don't As long as it works. Right. You know, right. With CodeClash, it was kind of testing this idea of like, all code artifacts and make them compete against each other. Right. and so that would that then it landed itself into this idea very competition. I think something I still like about CodeClash, even though it ProgramBench, is that technically it deals with the little bit better. Right. Right? Like it's very open-ended. It's it's truly faithfully open-ended in that way. You're not gonna run into like, if a model is able to create solution that's better than another, that's all that's needed Right. so I think like I'm hoping that maybe it's like a little bit It's also possible that because the arenas are all games, sort of, I want more economically valuable things, which But I think that idea of making artifacts compete against inspiration for ProgramBench, which is okay, it's still sort of factor-wise written as a pytest, but they're all invocations it's all calls of the program. It is no longer touching anything about the implementation. So now you completely disentangle the evaluation from sort Right. You know, the overarching theme is sort of like fundamentally I think with SWE-bench, it was can you implement correctly? With CodeClash, it's hey, can you kind of evolve your code base Right. And then with ProgramBench, it's can you create this artifact? Yeah. And I think these like from that sort of what's the end goal of view, I think then that will kind of naturally sort of with Inform sort of what the verification should look like. Yeah.
Vincent:That that resonates. The hey, what is the outcome that I'm trying to actually And then, you know, let's let's find a clever way to actually, you know, inject that reward or learning signal as a Where do you see as the role of human knowledge or human in driving these evaluations or steering these types more broadly? I know you had a position paper, you know, around how humans these coding benchmarks and research in general. Where do humans need to be in the loop as you're thinking about these types of mechanisms?
John:I've had a couple conversations with my advisor, Diyi very into like human AI collaboration. And I kind of came into the PhD program being sort of more But in the two years since I've been very convinced that like And so to really sort of ground it, the inspiration, Program SWE-bench for sure. and ultimately it's exciting that the solution space is so large. But in some sense, like even as human software developers, like have certain characteristics to them. And it's kind of conditioned on the purpose of the software. But generally, you know, we want them to be styled well, we want them to be readable, we want them to, you know, be portable. I think there's some general heuristics that are true, where I'm sure some of these labs are probably already post-training like this, but something about sort of like empowering the individual to say things and then very quickly steer the model effectively their preferences. Right. I think generally it's just important to me this idea that want it built. I don't understand hardware. I don't understand what goes into like HDL. Those people are experts. But as you and I, as the people who are putting these tools it's not sort of my. I think a better way to kind of solve the problem is not to kind of figure out what they do and then decide kind of like, all right, these are the things we should inject, but more to just make the model itself fundamentally steerable. So in the position paper, like me and Zora, who's a wonderful Neubig and the OpenHands folks who are great, she talks about this idea of steerability in the paper. And I think that's something that feels very compelling to me of the day where we're at like 80% on ProgramBench, that's very But then what does that mean for the people that ultimately use these things? Right. If it can get 80% on ProgramBench, but it's just and it's not really using a file system and the code is not Yeah. that's like it's a meaningful end state, but I think we can ask Yeah.
Vincent:Yeah, I agree with that. I think that that end state is more ambitious and harder to kind general. But yeah, we're really glad that you put out that position paper charting of again, kind of where we think things might go. I'd love to talk a bit about quality as well. I think one of the, you know, pieces that we you know pay a building data sets or environments or benchmarks is hey, overall aggregate quality, you know, is at a really high individual task level quality is at, you know, high you know, having models or agents in the loop is really It can be, you know, a double-edged sword in some cases. Yeah. in the case of you know, ProgramBench and you know, where of steer, you know, generate a bunch of the kind of instances. You've had a lot of learnings from you know, SWE-smith where you know generated tasks from well-curated How do you think about these settings where you know agents loop? but you know, you also want to manage really, really high
John:So far, to be honest with SWE-smith and ProgramBench, it's a funny thing with benchmarking where it's like you effectively models can't quite do yet that you want and you want to help them get there. But along the way, if there are certain portions that the model engaged, you know. Of course. So it's it's kind of funny. But with ProgramBench, I think the big thing was that test suites are not available, like behavioral style test suites that invocations of these programs are not available. empirically, we could ask a person to kind of help with this, and to some extent, like I'm not sure that a person without true can really write a good test suite. To concretely answer this question, I think models these days that there's probably three vectors I would imagine you could using the model. Right. One is to assist you like with generating the verification. Right, right, right. One of them, so the SWE-smith-oriented take is to task instances, which is like you have the code base, it's and then you do things like you know, you ask the model to you know, kind of a funky implementation that breaks tests. Right. And that's another thing. The last factor I would say is actually constructing the itself. So this is something where for both SWE-smith and ProgramBench, it Where for ProgramBench, it was, okay, if I want to get the reference executable, clone the code base and then ask the model to actually write the compilation script with all of the With SWE-smith, it was here's the Python library, make sure you sure you can run the unit tests as well. So I think sort of it's hard for me to come up with like a good mainly because like I think people are just experimenting with and it's just empirically in effective. Yeah, yeah, yeah, yeah. but yeah.
Vincent:No, I don't think there's a one size fits all solution either, right? I mean, in practice, a lot of what I and our teams find is you expertise and steering and and you know, high-level guidance in right places efficiently is it's tricky, it's case by case, you know, it's pipeline or workflow dependent. Yeah. but you know, having an intelligent way to say, hey, this is you know, to steer this part of the program. This is where actually, yeah, let the agents rip and start take it over. I think makes a lot of sense. so you know, you're you're certainly at the frontier of that, which is why I thought it was a yeah.
John:I mean, I think like in some sense, we can afford to kind of be small scale experimentation. Right. Because unlike before, once you get that right, it's much easier to scale up with agentic systems.
Vincent:In the ProgramBench paper, you stated that models that had hacking up to, I think, 36% of the time. You've also found, you know, back back in your SWE-bench days, right? let's call it leaderboard hacking when it came to you know how do you think about leaderboards, open benchmarks? and as a benchmark builder, you know, how to address this that we're measuring model performance in an honest and
John:I think there's like two parts to this really great question. The first half, which is like models are cheating, and the second, which is like, how do you set up a leaderboard that's open? And they're both really meaningful questions. When ProgramBench was launched, I think a big topic online internet is certainly useful. so why disallow it? and I would say like the central thesis that we were really anytime there's X percent improvement on ProgramBench, we want be sort of like undoubted, that there's no question, there's no model actually got it. It would be really cool to see a model sort of go on GitHub and source code itself. Right. and sort of do these things. Maybe it checks Stack Overflow, maybe it checks like a I agree that the internet is useful, but not at the cost of the rigor of the benchmark. And so the finding that you pointed out, it was twofold. We early on in the experimentation process, we allowed And then what we would do, and this is big credit to Kilian, is run at the trajectories and then they decide like did the model cheat Right. And you have these kind of cheating criteria specified in language. The problem wasn't just that there was, I mean, like a third is very tricky for us to release a benchmark where we're saying zero percent, but also like they cheat on like the messaging is very Right. And Kilian also found that the judges would disagree. And we have like a couple examples in the paper where like saying, like, oh, it's not allowed to look it up on SourceForge And the four is saying, well, it was only GitHub that was disallowed, so SourceForge should be fine. Right. And it becomes this kind of crazy cat and mouse game. So I think this is a constant tug of war for benchmark builders, but the typical advice I would go for is that rigor is king, is king, reproducibility is king. Everything else is kind of cherry on top. But if any of those things start to really affect, you know, the rigor is kind of really core to being able to then, you have people come together and hill climb and not point at each didn't do it right. You know, that's important.
Vincent:For the second half of setting up the leaderboard, yeah, I
John:curating submissions for SWE-bench was certainly like a a wonderfully interesting experience. like as of today, and I we've kind of you know tailed off on I think we ended up accepting like over 300 on the leaderboard. I think the experience there is to set up leaderboard submission pipelines that enable members of the community to kind of be by and also sort of like double check each other's work. Right. So there was a big moment in SWE-bench where because of sort of has been repeated in later benchmarks, that we said you must And I'm really glad to say I think this is something that Bench has adopted now, and we'll certainly see that adopted for But the moment we ask people for all those trajectories, one, I submit, I think naturally, you know, you kind of like remove those solutions came to be. And then the second thing is when the trajectories get uploaded, that people can look at it and they they can kind of it does inspire new approaches, but it also does help with sort of questions of is this cheating? You know, like like the credit to the fair coding team, they were behavior of SWE-bench like of some of the models fast forwarding, like going to future commits to solve the problem. And thanks to them, because after that we realized, you know, a reporting these things, right? Not necessarily out of malice, maybe because they just hadn't but yeah, I think just as open as possible because it just helps reproducibility, with kind of trust that you can build with, you you as a leaderboard host and the people participating in that. Yeah. yeah, never bad to be too open.
Vincent:That makes sense. How do you go about that error mode, failure mode analysis? I can only imagine, right? As these trajectories get longer and longer horizon, as the as you pointed out, get more and more creative. how does one try to understand and even kind of dig into
John:I think I there's like two sides of the coin to me for me on One side is like, we should definitely invest in sort of better Using language models themselves. I think there's been great research around this. Like you know, for instance, the Transluce folks have the Docent inspect SWE-bench trajectories using Docent. They're all loaded up. You can ask questions about it, you know, and it'll it'll do of give you an answer aggregated across all of the So tools like that, absolutely worth investing in. But then on the other hand, nothing beats just like sitting down and like manually scrolling through one of them. Yeah, yeah. So I think it's kind of like a balance. If anything, it kind of echoes the thing we said earlier about like human involvement in these things of just, yeah, take a little time, figure out how it looks for one or two instances manually, and then you know, scale to you know above and beyond.
Vincent:Yeah, I'm a I'm a big believer in both as well. We used to have as an onboarding exercise at snorkel, you gotta dig into the data, dig into the trajectories. Because to your point, there's nothing, nothing beats personal firsthand intuition for what the shape of the like, you know, what artifacts might look like, right? And you know, then translating that to more You know, I agree with that method. Yeah. It makes a lot of sense.
John:Yeah, there's just the threat of drift if it's always kind of one Exactly.
Vincent:I'd love to move to a lightning round, if you don't mind. So you know, hot takes are welcome. you know, no need to go super, super deep, but let's just these high-level ideas. You tweeted that we're we're moving from a regime where you know, the X things that humans can do, but we're shifting to we're asking the question, can language model actually do that were previously impossible? what are the types of tasks that fall into that second category? Yeah.
John:I've talked with Ofir Press, who's like a longtime collaborator some really excellent thoughts. He has this one diagram which is called like the five steps or where like one, two, three are kind of one, two, three are all kind of in the realm of like, can humans do this, you know? but then four and five are kind of like, oh wow, humans can't So I think ProgramBench, and this is kind of very much stealing So thanks, Ofir. But four is sort of in the realm of like there is kind of a solution that exists. So for example, for ProgramBench, the FFmpeg source code Right. you can generate tests on top of it. There's continuing to be progress on it. Basically, we're now compressing the timeline with which we can do that. So I think there's almost sort of models can effectively do the span than people can do. And so I think that's sort of superhuman in the sense of like, this off, but it just took us a way longer time. Right, right. so I think that's one way to think about it where it's like, what general, you know, representations of, you know, intelligence or, you know, you know, things that took a long, long that you can sort of now have the model try to recreate, like such a good proxy where it's like if a model could meaningfully that maybe it's gonna create equally revolutionary game-changing software that's novel, that's not FFmpeg that we haven't seen. Right. You know, but it's hard to grade it on those unseen things. So I think the seen things are good. The yeah, I think the fifth realm, and I think this is maybe and there's maybe people with like better formal ways to approach this, is like, yeah, just things that we know that are it's kind of those crazy Millennium Prize things, like and physics, like really challenging the boundaries of kind And I'm sure this is kind of where a lot of the excitement around you know, recursive self-improvement and AI scientists And I believe with that, I agree with that. I think my personal approach is more like having models be able that exist is a good stepping stone to then having meaningful something then new that's you know, just as, if not more
Vincent:Yeah, I subscribe to that mindset as well. I think there's a real value in empiricism, right? That's what benchmarks give us, right? I'm yeah, I am as much a fan as the kind of vibe eval or kind of you know, future looking more per projection-based eval, but scientifically study some of these models in the current you know, I think there's nothing that beats that. Exactly. Exactly.
John:Couldn't agree more. Yeah. Awesome.
Vincent:so GPT-5 just cracked, you know, one of the tasks in when does the first model hit 80%?
John:80%. Yeah, I guess for our sakes, hopefully not too soon. Just joking. maybe like a year, maybe like a year and a half from now. the way I'm thinking about it is with with ProgramBench, there is of nicely happened to have a gradient of difficulty. Right. So, you know, cmatrix, we have this kind of like difficulty scoring thing that we put in the appendix. If you look at cmatrix, it's it's literally one of the easiest And our difficulty rating is not empirically based. It's purely just based on looking at lines of code and the of dependencies. It's purely grading with respect to the characteristics of So it's totally independent of performance. And why I think that's meaningful is yeah, you know, the it's solved, it's fantastic, but you know, it's one of the easier ones, you know. So I think there's this nice grading of difficulty really at the of the toughest task instances, SQLite, the PHP there's this one really amazing dev called Fabrice Bellard, really fantastic software, like FFmpeg, whatever. I think there's just, I mean, there's a lot of just like true and I'm pretty, I have faith that the models will sort of become, one day. I think 80% to me on ProgramBench, which is like 160 out that sort of like start are quite meaningful at that point. Right. in terms of, okay, maybe it gets things like ripgrep or search in their utility. the remaining 20%, I think it I'm not sure. Maybe it doesn't take that long. My my sort of hunch, at least for a human, is that you know it's a little bit of a long tail in that, you know, something like they're tricky. Yeah, but I mean compared to FFmpeg, this is kind of a different leak. So I think, you know, one year is kind of my my more bullish take, but I'm happy to be proven wrong. How how would you characterize?
Vincent:You said there's a meaningful difference between the newer and I like do do you have words to capture what what that shift actually looks like? Oh, for sure.
John:I mean, I think some of these executables, like if you just see some of them have like a lot of subcommands. And so, in some sense, like even putting aside the lines of code and dependencies and all that, they just encompass different of sort of functionality, you know, like that FFmpeg handles so and it does all sorts of different kinds of, you know, you you know mp4, but you can apply this kind of like encoding to a model, you know, they're gonna really have to do some heavy downloading different assets. These assets vary by how long that video is, kind of what the person? Or, you know, I mean, I'm just coming up with these dimensions on things I've missed that FFmpeg accounts for. and so I think to some extent, ProgramBench beyond just correctness and models as software developers, I think it's extent of like, did you probe enough? You know, like because if you didn't probe enough, it doesn't And I think something about that almost extends to models then better sort of, you know, innovators in that so much of what there's no specification, there's no kind of well-written doc it. It's just someone really kind of poking endlessly at this thing that wasn't realized before. and they formalize it. You know, this is the whole like it's not a matter of just you enough. Like he just literally had to try a bunch of different metals know. So I think this kind of like persistence I think is exciting, you
Vincent:I love that notion that a bunch of software is actually the contacting reality and and really understanding. Understanding their own preferences, being curious and to really capture all the different ways that they might be to use it. And for something like ProgramBench to be able to extend that form of, you know, as you put it, curiosity or I think is a really exciting way to think about, you these models are going. Yeah. What benchmark should more people be paying attention to?
John:I think these days my personal preferences and focuses are on long horizon things within the realm of kind of 20 to 30 to 40 SWE-bench. I think we have the formulas in place to kind of solve that so in terms of long horizon, I think it's meaningful because it know, I guess continual learning is a hot word these days. So I'll lean on that. but very much in the sense of like, how do models construct memory And that memory, it doesn't just have to be natural language. It could even be sort of like this model and this model. They both solve this task instance. But which one provides a solution that's more extensible to picks it up next? You know? so I think concretely, kudos to you guys, like the continual that your your team did, that feels like really great. I almost kind of like it in that the way I see it is that you like a hundred turns or you know, just as a random number, but learning bench, which is also effectively a hundred turns, but you middle depending on kind of like where things pick up. Like one of the tasks I really like was it's kind of resembles like where you have issue A, and you can kind of find these naturally occur online of like B was blocked by A, but then C was blocked and then obviously you can keep going there. And so once you unblock A, when you solve B, I mean, it's one that function in a way where it's kind of at the right layer in such that B can correctly invoke it, right? Versus, oh, you know, the solution is too high up in the call So then when you implement B, you end up redundantly repeating below. that seems really fascinating to me. Yeah. Yeah.
Vincent:Well, appreciate the kind words. Yes. A lot of the goal there was, hey, how do we measure an ability to learn from experience, learn from, you know, interacting with an environment, multiple instances environment. So it's something that humans can do quite efficiently and And I agree that that's that's you know one of the ways that know, autonomy long term. definitely. What is a benchmark that you wish existed? Oh man. I think there's there's a lot of them.
John:One I think I'm fairly excited about is using coding generally to aren't just coding. So what I mean by that to expand is for instance, SWE-bench, I grade models on their ability to write code for the sake of better You know, we fix this bug. But I think as we've seen with like Claude Code and just like the proliferation of these kind of SWE-agents that people are code, right? Like in some sense, like I mean, Claude Cowork is fantastic, an engine where underlying the engine is very much just a coding agent, you know, with a couple extra tool calls and whatever. So I think I would just be really excited about leaning on the outside of. I mean, this is this is maybe speaking in too general, you know, too many generalities, but I truly believe like, hey, if you had sort of law problems, could you recast that into something you know, it's manipulating code to some extent, it's like through a bunch of documents, but you know, it builds its own So I guess this is like a very, very high-level answer. if I had to kind of sit down today and brainstorm and start the biology or medical industry, I think you know, like science of like, hey, there are these experiments, there repeatedly do. I guess I don't want to kind of say too much out of distribution because I'm not a biology person, but I would just be really about understanding more there and understanding like so many of truly realistically long horizon. Right. what's the blocker there for why we can't deploy an agent, you easily operationalizable. What do we still need to solve to make it then be able to do those things? Yeah, yeah. Yeah.
Vincent:Two two reactions. One, I think one thing I talked to Alex about also on the last one on Terminal-Bench and harbor. I think the bet on the CLI, the bet on code in general as a an awesome one, right? Because to your point, it's underlying a lot of the affordances that agents are using inherently to interact with the actually move it. two, I 100% agree with this higher level notion that, hey, we into the benchmark building process, right? A lot of you know what we as scientists or engineers or have contact with are problems within our domains or ones, right? But there's a whole universe of you know use cases and and ways that we could be accelerating and you know taking think needs to be represented in the broader family of so lots of agreement there. I think the biomedical space is a super interesting one.
John:Yeah, with you 100%.
Vincent:Five years out, does a benchmark still look like a
John:If we asked this question around when SWE-bench was created, like not that different, you know, from like 2018 to like 2023, but I don't want to give I don't I don't think so. And I think like a lot of it has to do with kind of my my kind of that we discussed earlier, right? Which is that we're entering this realm of benchmarks and sort of technically are like borderline, or if not absolutely, possible for people to do. And I feel like a lot of benchmarking that has at least has been sort of like silos of you know what people can do. we did GPQA, we had question answering, we had sort of, you know, the T5 paper came out, uniting all those things. We had machine translation, blah, blah, blah. And these are sort of like different cuts of what people can do and so fundamentally, I think having tasks that are inspired by, machine do this? I think it was, I mean, it was a very, very good source of task And we kind of calibrated which benchmarks to pay attention to models were in terms of their performance, right? So I think SWE-bench in 2018 probably wouldn't have been that fairly timely. so going forwards, I am excited to see how it evolves. I just think like the human inspiration will be a lot, like the maybe is gonna be less and less the actual source of it, in my Yeah. And more so along like how do we engage with it, how do we work Yeah, being more of the core premise than just sort of like hill climbing what a person can do. and in terms of just like the form factor of it, you know, I feel environments, whether they're digital or maybe they become environments sort of in the coming years, I think that sentiment will stay around. I think the verification is where I'm kind of more curious how least in the regime of code, verification has evolved so much in the past three, four years. Yeah. And I don't expect that trend to stop at all.
Vincent:Last question for you. how can how can people leverage ProgramBench? Yeah, definitely.
John:programbench.com, you know, the URLs there. It was a little pricey, but you know, we got it. so we're gonna set up a leaderboard very soon. We purposely kind of are encouraging people to just pick at we use mini-SWE-agent. You use your own scaffold if you think it's better, you know. single agent versus multi-agent scaffold, give that a I mean, as with SWE-bench, at some point I would be really excited to see people train their own models on this and see, like, hey, model? you know, well, is there really sort of, you know, given that what is training on that? You know, what how does that affect sort of how we train on code tasks when they get longer and sort of wider and bigger with larger solution spaces? So anyway, that's for setting up the leaderboard. I would also love to set up a way for people to add tasks because I think just ProgramBench as a formulation is meant to be very You know, there's very we want to impose as little constraints on You know, we just do executables, but there's no reason or an iOS application or a website or you know, anything that rendered software, you know. And so those would be, I think, the two call to actions. And, you know, Kilian and I are very sort of, you know, looking looking at the website traffic. So we're we're very looking forward to hopefully growing this just as excited as we are about sort of why this benchmark is
Vincent:Well, I'm excited as well. I hope you know, folks can contribute to the leaderboard, add and yeah, thanks again so much, John. This was an awesome chat.
John:Yeah, this was fantastic. Thanks so much, thanks, thanks.