AnyLearn
All lessons
Businessbeginner

Knowing Whether It Worked, on a Small Charity's Budget

Every funder asks for impact and few organisations can afford real evaluation. This lesson covers what you can honestly claim, the biases that make feedback flattering, where AI genuinely helps with qualitative data at scale, and why generating impact claims is the one thing that ends an organisation.

Updated · AI-authored, review-gated · how lessons are made

Not signed in: your progress and quiz score won't be saved.
Progress1 / 8

The impact question, honestly

Funders ask what difference an organisation makes, and the honest answer for most small charities is that they do not know with any rigour. That is not a criticism, and pretending otherwise is where the trouble starts.

The difficulty is structural. Establishing that your programme caused an improvement requires knowing what would have happened without it, and you do not observe that. The people you helped might have improved anyway. The ones who improved most might have been the ones already most likely to. And the people for whom it did not work may simply have stopped coming.

The machinery for handling this properly is covered elsewhere in this catalogue, in the cursus on knowing what actually works: comparison groups, randomisation, and why before-and-after misleads. That machinery is real, and it is mostly out of reach for an organisation running a programme for forty people on a grant of thirty thousand.

So the practical question is not how does a small charity run a rigorous evaluation. It is what can a small charity claim honestly, and how does it get better information without pretending to a rigour it cannot afford.

The answer has three parts, developed through this lesson. Be precise about what you actually observed. Be explicit about what you cannot separate. And collect the qualitative information you can, properly, because it is genuinely informative even though it is not causal evidence.

That is a defensible position, and funders who know the sector recognise it as more credible than a confident impact figure with nothing behind it.

Full lesson text

All 8 steps on one page, for reading, reference, and search.

Show

1. The impact question, honestly

Funders ask what difference an organisation makes, and the honest answer for most small charities is that they do not know with any rigour. That is not a criticism, and pretending otherwise is where the trouble starts.

The difficulty is structural. Establishing that your programme caused an improvement requires knowing what would have happened without it, and you do not observe that. The people you helped might have improved anyway. The ones who improved most might have been the ones already most likely to. And the people for whom it did not work may simply have stopped coming.

The machinery for handling this properly is covered elsewhere in this catalogue, in the cursus on knowing what actually works: comparison groups, randomisation, and why before-and-after misleads. That machinery is real, and it is mostly out of reach for an organisation running a programme for forty people on a grant of thirty thousand.

So the practical question is not how does a small charity run a rigorous evaluation. It is what can a small charity claim honestly, and how does it get better information without pretending to a rigour it cannot afford.

The answer has three parts, developed through this lesson. Be precise about what you actually observed. Be explicit about what you cannot separate. And collect the qualitative information you can, properly, because it is genuinely informative even though it is not causal evidence.

That is a defensible position, and funders who know the sector recognise it as more credible than a confident impact figure with nothing behind it.

2. Why your feedback is flattering

Most organisations collect feedback and most of it overstates how well things went, for reasons that have nothing to do with dishonesty.

Survivorship. The people who complete a programme are the ones for whom it worked well enough to keep coming. The person who found it useless left in week two and is not in your data. If you survey completers, you are surveying the subset selected for satisfaction.

Who responds. Among those you do ask, the ones who reply skew toward the engaged and the grateful. People who were disappointed frequently do not bother.

Who is asking. A beneficiary asked by the staff member who supported them, whose funding may depend on the answer, is under obvious pressure to be positive. This effect is strong, and stronger where there is any dependency in the relationship.

How you asked. Did you find this helpful invites yes. Leading questions produce the answer they lead to.

And hindsight. People asked afterwards how they felt before tend to misremember in the direction that makes the change look larger.

None of this makes feedback useless. It makes it systematically optimistic by an unknown amount, which is a different thing.

What improves it, cheaply. Ask people who left, which almost nobody does and which is the single most informative group you have. Have someone other than the delivering staff member ask. Use neutral wording. Ask what was not useful, explicitly, and treat a nil return as a sign the question failed rather than that nothing was wrong.

And report response rates alongside results, because a glowing result from six responses out of forty is a fact about six people.

3. What you can claim

A ladder of claims, from what any organisation can say to what almost none can, with the evidence each requires.

At the bottom, activity. We ran twenty sessions. Requires delivery records, and every organisation can say it honestly.

Above that, reach. Sixty people attended, of whom forty came more than once. Requires attendance data counted consistently.

Then observed change. Of those we assessed at both points, most improved on this measure. Requires before and after measurement, and note the careful phrasing: of those assessed, which excludes everyone who left.

Then reported experience. Participants told us the programme helped them, in these specific ways. Requires feedback collected with the biases above acknowledged.

And at the top, attributed impact. Our programme caused this improvement. Requires a comparison group and is out of reach for most small organisations.

The practical position for a small charity is to claim honestly at the third and fourth levels and be explicit about not having the fifth. We cannot separate our contribution from other factors is a sentence that costs nothing and buys credibility.

The damaging move is claiming the top rung with evidence from the middle, which is what a generated impact statement does by default.

flowchart TD
A["Attributed impact: we caused this"] --> B["Needs a comparison group: rarely affordable"]
C["Reported experience: participants told us"] --> D["Needs feedback, with its biases acknowledged"]
E["Observed change: of those assessed, most improved"] --> F["Needs before and after measurement"]
G["Reach: sixty attended, forty returned"] --> H["Needs consistent attendance data"]
I["Activity: we ran twenty sessions"] --> J["Needs delivery records"]
A --> C
C --> E
E --> G
G --> I

4. Where AI genuinely helps: the open text

There is one evaluation task where these tools change what a small organisation can do, and it is worth being precise about it.

Most charities collect qualitative material and then cannot use it. Open-text survey responses, feedback forms, case notes, exit interviews, comments in sessions. Analysing that properly means coding it: reading everything, identifying recurring themes, categorising each response, and counting how often each theme appears.

Done by hand, that is days of work for a few hundred responses, which is why most organisations read a sample, pick three quotes that sound good, and put them in the report. The rest is never analysed.

Assisted thematic analysis changes the economics of that. A model can propose themes across a full set of responses, categorise each one, and produce counts, in minutes.

Why this is a genuine and unusually safe win. The input is your own data. Nothing is being invented, because every theme is grounded in responses you collected. And the failure mode is a miscategorisation you can check by reading the responses in a category.

The disciplines. Read a sample of each theme yourself to confirm the categorisation means what you think. Watch for the model collapsing distinct concerns into one comfortable theme, which is its characteristic error and which loses exactly the specific criticism you needed. Look explicitly for negative themes, since summarisation tends to smooth. And anonymise before processing, per lesson two.

The result is that a small organisation can actually know what its beneficiaries said, in aggregate, for the first time. That is a real capability gain rather than a time saving.

5. Quotes must be real

A short step, because it is a bright line rather than a judgement, and it is the highest-consequence rule in this cursus.

Beneficiary quotes are the most persuasive element in any charity communication or report. A funder reads one and it does more work than a page of statistics. Which is exactly why they must be real.

A generated quote attributed to a beneficiary is a fabrication used to obtain money. In a fundraising appeal it is a donation obtained by deception. In a funder report it is a false statement to an organisation that gave you money on the strength of your honesty. In a regulated charity it is a matter for the regulator.

And the practical exposure is severe, because charity work is local. Funders visit. Journalists ask. Beneficiaries read what is written about them. A quote nobody said is discoverable in a way a rounded statistic is not.

The rules that follow.

Quotes come from a recording, a written response, or contemporaneous notes. If you cannot point to where it came from, it does not run.

Light editing for length or clarity is normal practice and should be signalled where it changes meaning. Composing a quote from several people's comments is not a quote, and if you present a composite you say so.

Consent is separate from accuracy. A real quote used without permission is a different failure, and for beneficiaries in vulnerable circumstances it can cause direct harm.

And when you need an illustrative example and have no usable quote, describe the pattern in your own words. Anonymised case study, no direct quotation is honest. An invented sentence in quotation marks is not.

6. Measuring what matters, not what counts

A failure mode worth anticipating, because tooling makes it easier to fall into.

When measurement becomes cheap, organisations measure more. That sounds good and it introduces a known hazard: what gets measured gets managed, including when the measure is a poor proxy for the goal.

The general form of this is well established. Charles Goodhart's observation, popularised as Goodhart's law, is that a measure which becomes a target ceases to be a good measure. Donald Campbell made a similar argument about social indicators. The mechanism is that people optimise the indicator rather than the thing it stood for.

In this sector the versions are specific and damaging.

Counting attendance rewards recruiting people who were easy to reach rather than people who needed it most.

Counting completions creates pressure to enrol those likely to complete, which selects against the people with the most chaotic circumstances.

Counting outcomes on a short timescale rewards programmes with quick effects over ones with durable effects.

And reporting only successes teaches an organisation not to look at failures, which is where the learning was.

What protects against it. Choose few measures and keep them stable, so you can see change over time rather than shopping for a flattering metric. Pair every quantitative measure with qualitative material, which resists gaming because it is not a number. Explicitly measure who you are not reaching. And separate the numbers you report to funders from the numbers you use internally to improve, because those have different purposes and conflating them corrupts the second.

The deeper point. Cheap measurement is only an improvement if it makes you better informed. Measuring more things badly makes an organisation more confident and no wiser.

7. Proportionate evaluation

What a small organisation should actually do, given that rigorous evaluation is out of reach and doing nothing is not acceptable.

Start with a clear statement of what you think happens and why. A theory of change need not be elaborate: this is who we work with, this is what we do, this is the change we think it produces, and this is why we think that. Writing it down is valuable on its own, because organisations frequently discover their staff hold different theories.

Then measure the few things that would show your theory is working, and crucially, the few that would show it is not. Most evaluation only looks for confirmation, which is why it always finds it.

Collect consistently over time. A modest measure tracked the same way for three years is worth far more than an elaborate one-off evaluation, because it shows direction and it survives comparison.

Talk to people who left. Repeatedly, because it is the highest-value and least-collected information you have.

Be honest in reporting about what you cannot separate. Contribution rather than attribution is a legitimate and widely understood position.

And where a funder wants rigorous evaluation, ask them to fund it. Evaluation is a cost, small organisations cannot absorb it, and funders increasingly recognise this. An organisation that says we would need resources to evidence this properly is being professional, not evasive.

What the tooling contributes across all of that is real but narrow: analysing the qualitative material you already collect, drafting the reporting, and structuring the theory of change. It does not produce evidence, and any use that appears to be producing evidence is producing text.

8. What the cursus adds up to

Pulling three lessons together for someone running or working in a small organisation.

The realistic gain is capacity, not efficiency. You are not going to run the same organisation with fewer people. You are going to do the reporting, the supporter communication, the volunteer induction and the feedback analysis that were previously being dropped. Judge it on that.

The repetitive layer is large and safe to compress: reformatting organisational material across applications, assembling reports from records, drafting routine communications, analysing open-text feedback. All of it operates on material you produced.

The evidential core cannot be generated, and every part of this cursus converges on that. The local specifics that win a grant. The numbers you will be held to. The beneficiary quotes. The claims about impact. Each of these is something only your organisation has, and each fails in a way that damages trust rather than merely being wrong.

The data risk is disproportionate. You hold information about people in difficulty, and budget pressure pushes you toward the tools with the weakest terms. That inversion needs an explicit decision rather than drift.

And the strategic warning: do not respond to cheaper applications by submitting more. Fit decides outcomes, mass applying degrades the sector's success rates for everyone, and it pushes funders toward relationships, which favours organisations that already have them.

The unifying idea. In a sector whose entire asset is being trusted, by funders, by beneficiaries, by supporters and by a regulator, the tools are safe wherever they express what is true and dangerous wherever they supply what should have been true. That distinction is worth more than any list of approved uses.

Check your understanding

The lesson ends with a 5-question quiz. Take it in the player above to see your score.

  1. Why is programme feedback systematically optimistic?
    • Beneficiaries are reluctant to complete surveys
    • Survivorship, response bias, who asks the question, and leading wording all push the same direction
    • Small samples are inherently biased upward
    • Charities edit responses before reporting
  2. Which claim requires a comparison group?
    • We ran twenty sessions
    • Of those we assessed at both points, most improved
    • Our programme caused this improvement
    • Participants told us it helped them
  3. What is the characteristic error when a model does thematic analysis?
    • Inventing responses that were not submitted
    • Refusing to categorise ambiguous responses
    • Producing too many distinct themes to be usable
    • Collapsing distinct concerns into one comfortable theme, losing the specific criticism
  4. What does Goodhart's law predict when measurement becomes cheap?
    • Measures that become targets cease to be good measures, as people optimise the indicator rather than the goal
    • Organisations will measure fewer things
    • Measurement error decreases with volume
    • Funders will require fewer metrics
  5. What is the unifying distinction across the cursus?
    • Free tools versus paid tools
    • Internal documents versus external ones
    • The tools are safe where they express what is true and dangerous where they supply what should have been true
    • Quantitative work versus qualitative work

Related lessons

AI
intermediate

What This Teaches About Measuring Anything

The exchange is a case study with transferable rules. A conclusion resting on failures needs a failure taxonomy. Every instance must be verified solvable before anyone is scored against it. Output format is a confound whenever answers get long. And when two explanations fit the same data, the productive move is to find the prediction on which they differ, then test it.

8 steps·~12 min
AI
intermediate

The Rebuttal: Three Ways to Score Zero Without Failing

The response disputed none of the data and argued the experiment measured something other than reasoning. Models had to print move lists exceeding their output limits, and said so in the transcripts. Some instances had no solution and were scored as failures anyway. And asking for a program instead of a move list produced high accuracy on instances reported as total collapse.

8 steps·~12 min
AI
intermediate

The Experiment: Puzzles With a Difficulty Dial

Apple researchers built an evaluation designed to fix a real problem with benchmarks: puzzles where difficulty turns up smoothly while the logic stays identical, and every step can be checked. They found accuracy collapsing to zero past a threshold, and not improving when the solution algorithm was handed to the model. This lesson covers the design and why it was a good one.

8 steps·~12 min
AI
intermediate

What a Percentage Does and Does Not License

A model went from 27 percent to around 57 percent, so it is more than halfway to AGI and the rest arrives shortly. That inference is wrong in at least four ways, and working through why is more useful than the score itself. This lesson covers the linearity assumption, construct validity, contamination, and what the framework is good for once you stop reading it as a progress bar.

8 steps·~12 min