This is the third blog in our reflective series on co-designing evaluation of holistic Community-Led Development (CLD) for the Full Spectrum Evidence and Learning Initiative (partnering with Tostan, One Village Partners and IDinsight).
In our first blog we named the mirage we were chasing as the idea that somewhere out there sits one design that is simultaneously the most credible, the most appropriate, and the most ethical way to evaluate a holistic, community-led approach to development. In our second blog we took you into the room of our co-design workshop where that mirage dissolved into something messier. We shared our reflections from a group of partners with genuinely different views on which causal designs were appropriate for this evaluation, plotting our distinct positions and reasoning rather than glossing over it.
We left that blog with some open questions. Did we find a middle ground? What would an appropriate and collective design look like in practice? How do we hold the absence of community members from that room, when their experience of change is what the whole evaluation is meant to be about?
This blog picks up where we left off and explains where we have landed after a six month theory-building phase.
We didn’t compromise. We recombined.
The honest answer to “did you find a middle ground” is not in the way that question implies. We didn’t take the counterfactual design and complexity-aware design and average them into something blander than either. Instead, the design that emerged treats the tension itself as a guide about what the evaluation needs to do.
The anchor for the whole design is contribution analysis (CA) which is a theory-based approach to evaluating complex interventions where many factors interact to produce change. CA doesn’t try to isolate a programme’s effect, rather, it builds a plausible, evidence-based account of how and why a programme contributes to the changes we observe, alongside everything else going on in a community’s life, and it gathers evidence both for and against that account, revising it as the work proceeds
Choosing CA as the backbone for our design means we have opted for a framework flexible enough to hold genuinely different kinds of evidence and methods, brought together for the specific question or ‘causal hotspot’ they best suit. In evaluation circles we increasingly refer to this as methodological bricolage, a practice of recombining parts of methods in an intentional, iterative, co-design process that responds to user needs. It’s a concept some of us have written about before, and it is, in a real sense, what disrupting the “performative dance” we described in blog one actually looks like in a finished design, not just in a deliberative workshop.
Anchoring an unwieldy design through defining causal hotspots
Holistic CLD is, by definition, holistic. Tostan and One Village Partners’ (OVP) programmes touch nearly every dimension of community life, from how decisions get made to how people see their own capacity to act together to the material gains they seek through their actions. No evaluation could responsibly claim to cover all of it. So rather than being blinded by our internal debates about methods, we spent November 2025 to May 2026 in a theory-building phase, which aimed to anchor the design in what really matters to partners and the sector.
We reviewed programme documentation from three organisations, holding structured conversations with Tostan, OVP and Legado teams to surface their explicit and implicit theories of change. Two rapid evidence reviews and analysis of existing data added to our understanding of what evidence and theory exists and where there are gaps. Rather than staying faithful to one methodological school, we explored what partners (and the field) need, where we need to build theory and what is realistic and feasible to build what is (in our opinion) the best design possible.
That work produced two things the rest of the design now hangs on:
- A shared Theory of Change for holistic CLD, agreed across partners, built around nine interconnected dimensions that reinforce one another rather than unfolding in a simple linear sequence. We theorise how individual and collective confidence enables collective action, and how collective action builds social cohesion, and how, through multiple reinforcing dynamics – which create positive feedback loops– through time, together, they contribute to social transformation, good governance, and the human development outcomes communities are ultimately seeking.
- Six agreed and prioritised causal hotspots: these are the specific points along that theory of change where the evidence base is weakest and the learning opportunity is greatest. These became the anchor points for the entire design, the places where we decided it was worth spending the evaluation’s limited time and resources.
These anchor points have enabled us to navigate the “how do we evaluate everything” problem to land on what is field building, and, possible.
What “fit” looks like in practice
Methodological bricolage is easy to say and harder to show, so here’s what it actually looks like in the design.
For our first research question on whether holistic CLD contributes to outcomes in governance, social transformation, and human development, we will have two samples side by side rather than choosing between them. A baseline-to-endline household survey will include all participating communities, generating quantitative evidence broad enough to see patterns and benchmarkable against standard sources. Alongside it, a smaller set of case-study communities will be followed in much greater depth through longitudinal qualitative work: micro-narratives collected from community members over time, participatory methods like PhotoVoice and Ripple Effects Mapping, and Outcome Harvesting to capture changes the programmes didn’t necessarily predict. Neither sample alone would give us what we need. Together, they’re wide and deep at once.
For our second research question, which asks what design principles matter and under what conditions, the method is defined by the hotspots. One hotspot focuses on how social norms shift and take hold over time, and here we will use Realist Evaluation because we expect the same programme component to work differently for different people and in different contexts, and we need a method built to explain that variation rather than average it away. Rather than relying only on reconstructing change after the fact, we will track 9–12 case-study communities longitudinally across the implementation period, returning repeatedly to the same ‘tracer’ participants so we can observe norm change as it unfolds. Quantitative longitudinal tracking of attitudes, behaviours, and both descriptive and injunctive social norms will run alongside, to see whether the patterns we observe qualitatively hold at scale.
The point isn’t that any one of these methods is ‘better’ than the others. It’s that appropriateness gets decided hotspot by hotspot, against the specific causal dynamics in play, rather than settled once for the whole evaluation. That is what “fit” means in this design.
A golden opportunity to follow change over time
One piece of the design we are particularly excited about sits within our third research question, on sustainability and scale is an ex-post evaluation. Tostan has been implementing across Senegal for decades, and many current staff once worked as hands-on facilitators in the very communities we want to revisit. Very few holistic CLD organisations have the institutional memory or capacity to let an evaluation ask what their work looks like twenty or thirty years on, and we get to. We’ll start with staff workshops using Ripple Effects Mapping to trace how change has rippled outward over time, then sample communities ranging from recent (2–3 years post-programme) to very long (20–30 years post-programme), deliberately including cases where things didn’t stick. Returning to those communities with Innovation Histories, we’ll work with community members to reconstruct what changed, what lasted, what reverted, and where the turning points were. This is, as far as we know, a fairly unique feature of our design, and one we think holds real promise for understanding what holistic CLD leaves behind and how underlying system dynamics might have been shifted once the programme itself has moved on.
What stayed true, and what’s still unfinished
We want to be honest that landing on this design didn’t resolve everything blog two raised. We will continue to fine tune as we deepen checks on operational viability with field teams in Senegal and Sierra Leone as the study moves from paper to practice. And the question we asked at the end of blog two “how do we manage the reality that community members, the people at the heart of holistic CLD, were not in the room when this design was built” remains open but not forgotten.
We engaged with this challenge through piloting the use of micro-narratives with Tostan and OVP’s MEL teams. Between February and April 2026, OVP and Tostan teams collected dozens of stories from community members and programme participants of critical moments that stuck with them. The storytellers then made sense of their stories through answering signifier questions about what changed, how, and why it mattered to them.
The early findings have surfaced some patterns we wouldn’t have predicted, and others that confirm what we expected to see. In the OVP pilot, for instance, storytellers overwhelmingly framed positive change in financial terms, such as: improved hygiene reducing medical costs, better agricultural practice reducing waste, financial independence freeing them from needing to ask others for help. They tied that financial change closely to a sense of individual confidence. What the stories did not show, at least not yet, was storytellers explicitly connecting that confidence to a belief in their community’s collective capacity to act together, which is a useful and humbling check on an assumption built into our theory of change. We are continuing to refine the method with both teams, and expect micro-narratives to sit alongside the case studies and other methods as a further, distinct way of letting communities’ own framing of change shape what we learn, not just confirm what we already believed.
What has stayed true since blog one is the underlying belief that drives the partnership: that evaluation design is a deliberative act, not a technical one, and that naming the values and biases we each bring into the room is what lets a genuinely fit-for-purpose design emerge, rather than a performative one.
If you’re grappling with similar design dilemmas in your own evaluation work, particularly across multiple partners and contexts, we’d like to hear from you.
For the full technical paper on the causal hotspots, or for more detail on the evaluation design, contact Marina Apgar or Mariah Cannon.