Skip to main content
Verify a survey
Probatio Africa

Services

Evaluation

We test whether something worked. An evaluation asks whether people are better off than they would have been without the programme, who gained most, and what it cost.

Questions this answers

  • Did it work, for whom, and at what cost?
  • Is it being delivered the way it was designed?
  • Is this programme ready to be evaluated at all?

The kinds of evaluation we run

Most evaluations combine two or three of these, at different points in a programme’s life.

Baseline, midline and endline studies

Measures a programme before it starts, partway through, and after, so change can be tracked over its life rather than guessed at the end.

Process and implementation evaluation

Checks whether a programme is being delivered the way it was designed, before asking whether the design itself worked.

Impact evaluation

Experimental, quasi-experimental and theory-based designs, used to ask whether the programme itself caused a change, not just whether a change happened alongside it.

Gender-responsive and inclusive evaluation

Asks who gained, who was left out, and why, not only what happened on average.

Cost-effectiveness and value-for-money analysis

What the result cost to achieve, and whether a different design could have achieved it for less.

Rapid and evaluability assessments

A quick assessment where time is short, or a check of whether a programme is actually ready to be evaluated at all before a full evaluation is commissioned.

How we design an evaluation

The design changes what the evaluation can prove. We state ours in every proposal, before you commit.

Start from the decision, not the method.

Whether a programme should continue, expand, change or end is what decides the questions we ask, not the other way round.

A theory of change, agreed first.

What the programme assumes will happen, and why, written down and agreed with you before we design how to test it.

A counterfactual for any impact claim.

To say a programme caused a change, not just that a change happened, requires knowing what would have happened without it. We state how we estimate that counterfactual, and we say plainly when a design cannot support a causal claim.

Evaluability checked first.

Some programmes are not yet ready to be evaluated: the objectives are unclear, the data does not exist, or it is too early to see an effect. We say so before proposing a design that could not answer the question.

The six standards an evaluation is judged against

Most development evaluations are judged against six criteria set by the OECD’s Development Assistance Committee. They were revised in 2019, when coherence was added. We state which of the six a given evaluation covers, and why not the others, in every proposal.

Relevance

Is the intervention doing the right things?

Coherence

How well does it fit with other efforts?

Effectiveness

Is it achieving its objectives?

Efficiency

How well are resources being used?

Impact

What difference does it make, for better or worse?

Sustainability

Will the benefits last?

From our own work

Real work of ours that used this kind of research, not an invented example.

A cross-border energy infrastructure operator

Real example

Measuring community impact along a cross-border energy corridor

308 households in each of two rounds, 2019 and 2022. 33 communities in Lagos and Ogun States, along a 56km right-of-way.

Reading an evaluation against the six criteria

Illustrative
CriterionFinding
RelevanceMatched the state’s stated priorities at design
CoherenceDuplicated an existing programme in 2 of 6 example local government areas
EffectivenessReached 71% of its enrolment target
EfficiencyCost 18% more per participant than budgeted, mainly on transport
ImpactNo credible counterfactual was available; effect on outcomes not established
SustainabilityDepends on continued donor funding; no domestic budget line agreed

This is what a criteria-based finding looks like: a short, checkable statement against each standard, including where the evaluation could not draw a conclusion (impact, here) rather than a single overall verdict.

Questions worth asking an evaluation provider

There is no single published checklist for buying an evaluation the way ESOMAR publishes one for survey samples, so we built five questions from the OECD criteria and the independence principle on our Standards and ethics page, and we answer all five up front.

01

What is the counterfactual, and how is it established?

Stated in the proposal, in plain terms, including when no credible counterfactual is possible and the design says so.

02

Which of the six criteria does this evaluation cover, and why not the others?

Named explicitly. Not every evaluation needs all six, but the choice is stated, not left implicit.

03

Is the programme actually ready to be evaluated?

Checked before the design is proposed, not assumed.

04

Who decided the evaluation design, and could pressure change the findings?

Scope, data access and publication terms are agreed in writing before fieldwork begins, and we report findings whichever way they point.

05

What can this design prove, and what can it not?

Stated in the report itself, not only in a methods appendix.

What you receive

An evaluation report with findings split by sex, age, disability and location, and a plain statement of what the design can and cannot prove.

The findings

Organised against the criteria the evaluation covered, split by group and place wherever the data allow.

What the design can prove

A plain statement of what can, and cannot, be concluded from the method used — including where a causal claim is not supported.

What it means for the decision

Whether the evidence supports continuing, expanding, changing or ending the programme, stated directly.

Questions people ask about evaluation specifically

What’s the difference between monitoring and evaluation?

Monitoring tracks whether a programme is running as planned, continuously, using routine data. Evaluation asks a deeper question at a point in time — whether it worked, and why — usually using data monitoring alone cannot supply.

Can you prove our programme caused the change, not just that it happened alongside it?

Only where the design supports a counterfactual. Where it does not, we say so plainly rather than implying causation the data cannot support.

Is our programme too early to evaluate?

Possibly. An evaluability assessment checks this before a full evaluation is commissioned, and can save you from paying for an evaluation that cannot yet answer your question.

Will you evaluate a programme you also implement or advise on?

No. We are a research firm and do not implement programmes. If we have advised on the programme being evaluated, we say so in our first reply, and may decline.

Deciding whether to continue, expand or end something?

Tell us the decision and the date you need to make it by. We will say which evaluation design can answer it in time, and what it cannot.