Integrity Matters
We work on complex and challenging issues around the globe.
Integrity Matters
Beyond BASIC Episode 1: Methodological innovation in the evaluation
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
We’re excited to launch Beyond BASIC – a new Integrity Matters podcast mini-series sharing learning and innovation from Integrity’s endline evaluation of the UK Foreign, Commonwealth and Development Office’s (FCDO) £23.5 million Better Assistance in Crises (BASIC) Programme.
Episode 1: Methodological innovation in the evaluation, dives into how the evaluation was designed and delivered, and how innovation helped ensure rigour, transparency, and practical learning in volatile contexts. Niki is joined by two guests: Integrity MEL expert, Ed Small and Georgia Plank from o-x.ai, who both were part of the evaluating team.
Hosted by Niki Wood, Integrity’s Head of Monitoring, Evaluation and Learning (MEL), the series takes listeners behind the scenes of evaluating a complex global research and technical assistance programme – from the methods used to the insights generated.
Look out for the rest of the series, launching here every week during July:
- 2 July – Evaluation findings: Relevance and outputs
- 9 July – Evaluation findings: The future of social protection programming
- 16 July – Behind the evaluation: Relationships, transparency and trust
Hello and welcome to Beyond Basic, a podcast series showcasing learning and innovation from evaluating the UK Foreign, Commonwealth and Development Office's £23.5 million Better Assistance in Crises or BASIC program. My name is Nikki Wood, and I am the head of Monitoring, Evaluation and Learning, MEL for short, at Integrity. I will be your host for this podcast mini-series where we will be talking about integrity's endline evaluation of the BASIC program. We will be exploring the methods used, diving into some of the key findings, and reflecting on the relationship between the evaluators and the program team. In today's episode, we're delving into how the evaluation was designed and delivered. And here's the story we're telling today. Evaluating a complex global research and technical assistance program like BASIC required methodological innovation, not for novelty, but to protect rigor, transparency, and usability. By combining contribution analysis, expanded country casework, structured evidence grading, and responsible AI-supported document review, the evaluation strengthened both credibility and learning in volatile contexts. I'm joined today by two members of the evaluation team. Could I please ask you both to introduce yourselves?
SPEAKER_00Thanks, Nikki. It's great to be here. I'm Ed Small. I'm a MEL specialist at integrity. I've been designing and operationalizing MEL systems for humanitarian and development initiatives for more than 10 years. And I was an evaluator on the basic end line.
SPEAKER_01Hi Nikki, I'm Georgia Planck, and I was also an evaluator on the project. And in fact, I've been involved at all three stages of this evaluation, from the baseline to the midline and end line. I have a background in social development, originally at DIFID working first in Yemen and then the London-based governance policy team. And one of the things I did for this evaluation was carrying out a large-scale document review using AI and a particular tool that I co-created, which I think we'll get into a bit more shortly.
SPEAKER_02Thank you so much. It's wonderful to have you both with us. Let's begin with context. BASIC is a £23.5 million centrally managed FCDO program, and it was running from 2018 to 2026. It worked across three work streams: technical assistance, research, and knowledge management and learning, with the aim of strengthening the use of social protection approaches in crisis settings. It operated globally and at the country level, often in fragile and conflict-affected contexts. Importantly, it sought to be catalytic, aiming to influence systems, policy, and coordination, not to deliver services directly to end beneficiaries. It was a complex program and needed an evaluation that could capture more than outputs. It needed an evaluation that would allow FCDO to understand how influence and system strengthening actually worked, how the components interacted, and how the need to adapt affected results over time. Ed, can I ask, what were the methodological implications for designing an evaluation that would capture all of this?
SPEAKER_00As you say, this was a complex program. It wasn't one where you could measure impact through experimental methods or direct beneficiary surveys. The scope did not include measuring final impacts on end beneficiaries either. Instead, we were assessing contribution to policy shifts, coordination improvements, capacity strengthening, and system level change. So overall, we were dealing with a catalytic program, multiple modalities, a global and country footprint, adaptive implementation, and political and fiscal volatility.
SPEAKER_02That is a lot. How did you describe an approach that could handle such complexity and uncertainty while also being structured and credible, Ed?
SPEAKER_00One distinct feature of the evaluation was that it followed a longitudinal approach. We evaluated BASIC across three phases baseline, midline, and endline, which allowed us to track how the program evolved over time, how adaptation took place, and how evidence on performance strengthened over the life of the program. Building on this approach, we adopted a blended theory-based and case-based design. At the core was contribution analysis. Rather than trying to prove attribution, we tested the program's theory of change against evidence. We asked whether the assumptions in the theory of change held, whether alternative explanations were plausible, and whether BASIC's activities plausibly contributed to observed outcomes. Because BASIC operated through policy influence, technical assistance, and research rather than direct service delivery, a non-experimental approach using mixed qualitative and quantitative methods was most appropriate. The country work was also particularly important. Four countries, Jordan, Nigeria, Somalia, and Yemen, were selected for in-depth study in every phase of the evaluation. And at the end line, we added three light touch cases, Lebanon, Ethiopia, and South Sudan, to improve geographic representation and test where the patterns held in new contexts. Finally, we also benefited from a fantastic team of evaluators, all of whom have backgrounds in social protection, including our quality assurance lead.
SPEAKER_02That's so interesting. Could you give an example of how the case-based approach strengthened the analysis?
SPEAKER_00Sure. In some cases, program reporting suggested strong optake of technical advice. But when we triangulated this with interviews in country, we found that while advice was technically sound, political constraints limited implementation. For example, in Nigeria, basic technical assistance helped inform the national cash and voucher assistance policy, which was approved in 2025. But interviews noted that implementation remained uncertain because political commitment to social protection was low, with respondents pointing to the absence of predictable government funding and recent corruption scandals affecting the sector. That didn't negate contribution, but it qualified it. Adding additional country cases helped us test whether that pattern was context-specific or more general. That's where the case-based logic added depth.
SPEAKER_02I see. And you also had to manage data limitations.
SPEAKER_00That's correct. For example, we conducted a survey, which we had also done at baseline and midline, and the response rate at endline was around 25%. Which is not unusual in busy policy environments, but it required careful interpretation. We had to adapt our approach slightly by offering structured online interviews where we walked respondents through the questions rather than relying purely on self-completed online surveys in order to increase the response rate. Then, from an analysis perspective, we treated survey results as one data source among many, rather than as standalone evidence. This structured triangulation was critical. We never relied on a single source to answer a sub-question, and we used open-ended questions to collect qualitative responses. A good example would be system level effectiveness, where the survey suggested relatively limited evidence of change. But when we looked at the country case studies, we found concrete examples. For instance, reforms to the National Aid Fund in Jordan that were informed by basic analysis.
SPEAKER_02That's so interesting. But let's talk about judgment. Evaluations of complex programs often face mixed or incomplete evidence. How did you handle that transparently, Ed?
SPEAKER_00We introduced rubrics at two levels. One to assess the overall strength of evidence, and another specifically for the case studies where we graded strength and significance of contribution. For each sub-question, we assessed the scale, consistency, and saturation of evidence. That meant asking, how much evidence is each sub-evaluation question based on? How consistent is the evidence across sources? Does it reach a point of analytical saturation? Where findings were mixed. We reported them as mixed. We didn't force binary conclusions.
unknownRight.
SPEAKER_02And how did you maintain independence while engaging closely with stakeholders?
SPEAKER_00Well, transparency was central to this. We were explicit about methods and limitations from the outset. FCDO and suppliers were given opportunities to comment on delivery progress and factual accuracy prior to submission. But conclusions remained the evaluator's responsibility. Independence wasn't achieved through distance alone. It was achieved through clarity of process and explicit criteria for judgment. And of course, by design, we had a team of people with a mix of technical and thematic understanding to support engagement while maintaining independence.
SPEAKER_02Let's turn to something innovative in this evaluation, the use of AI in document processing. Georgia, where did AI sit in the methodology?
SPEAKER_01Thanks, Nikki. So AI was used to support a large-scale review and analysis of project documentation. So we had over 130 documents, and that included quite a lot of scanned and non-machine readable files. So reviewing that volume manually is possible, but expensive, and maintaining consistency in coding against, I think we had 21 sub-questions would have been quite challenging. So the automated document review was carried out by staff from OEX, which is a research agency specializing in AI and data science for evaluation that I co-founded in 2022. And this helped ensure close integration between the AI workflow and the wider evaluation design and the work of the core evaluation team. The process operated within a secure Microsoft Azure tenant, so there was no public internet connectivity. And the workflow followed five stages. So in simple terms, we first uploaded and digitised the documents. Then we used natural language processing to identify themes in the text. Those themes were mapped against our evaluation questions. And finally, everything was reviewed and validated by the evaluation team. While existing qualitative software packages do have some AI functionality, we preferred this approach because it gave us much more flexibility to produce more useful analytical outputs that were tailored directly to our analysis. We also made sure that findings emerging from the document review could be traced back to the source text by making sure that all analysis was thoroughly referenced. And what did AI actually improve? Consistency of the analysis, efficiency, and also scale. So, for example, AI-supported theme extraction helped us to identify recurring references to system strengthening across multiple years in different geographies, and also highlighted where certain criteria, such as gender or inclusion, for example, appeared less frequently in the documentation. So that prompted us to take a closer look. We conducted spot checks, sampling checks, and cross-validation throughout. And we were also attentive to bias risks such as overweighting frequently used terms. So really the principle was about augmentation, and by that I mean that AI supported evaluators' reasoning and it did not replace it.
SPEAKER_02Sounds like a really valuable approach. And as we close, I'd like to ask both of you for one key lesson from this methodological approach. Let's start with you, Ed.
SPEAKER_00My key lesson would be design for the nature of the program. Contribution analysis is a great approach, especially when you take the care to operationalize data collection with contribution analysis in mind. For example, by designing tools that interrogate contribution claims without biasing the data, tools that can draw on a wider range of sources to understand and contextualize changes, and of course, tools that engage a broad range of stakeholder groups.
SPEAKER_01Georgia? When using AI, I'd say the importance of using secure infrastructure, staging prompting, traceability, and human oversight are all essential. AI can strengthen consistency and transparency, but only when it's embedded in a robust governance framework.
SPEAKER_02Thank you both so much. If there's one theme running through this episode, it's that innovation in evaluation isn't about novelty. It's about handling scale. It's about handling complexity. It's about making uncertainty explicit rather than obscuring it. And it's about strengthening credibility in politically and operationally challenging environments. In the basic evaluation, innovation supported rigor. It strengthened triangulation. It improved consistency, and it protected the integrity of judgment. Innovation in evaluation isn't about replacing expert reasoning, it's about strengthening it. And that's all for today's episode of Beyond Basic. Stay tuned for three more episodes in this series where we explore the key findings from the evaluation and delve into how the relationship between the evaluator and evaluan developed over the course of the evaluating process. They'll be available on the integrity website as well as on our LinkedIn. Hopefully, see you soon, and thank you for listening.