Paper
Björkman, Martina, and Jakob Svensson (2009). Quarterly Journal of Economics, 124(2): 735–769
Read the original paper → Opens at the publisher; use the DU library / JSTOR login if it asks for access.
A randomised evaluation of community-based monitoring of primary health providers in rural Uganda. The explainer below covers the intervention, the randomisation, and how the paper reads treatment effects on effort, utilisation and health.
An interactive, section-by-section guide to Power to the People: Evidence from a Randomized Field Experiment on Community-Based Monitoring in Uganda by Martina Björkman and Jakob Svensson, Quarterly Journal of Economics (2009)
Who watches the clinic?
In 2005, 25 of 50 rural Ugandan communities, chosen at random, got a report card on their local dispensary and a series of meetings to decide what to do about it. A year later, this is how the two groups compared.
144under-five deaths per 1,000 live births
661outpatients a month at the dispensary
Control: 2005 averages in the 25 control communities. Program: the treatment-group figure where the paper reports one (deaths, equipment, waiting time, suggestion boxes), otherwise the control average plus the paper’s estimated program effect.
Health workers in poor countries are often absent and rarely face consequences, because supervisors seldom visit and pay does not depend on performance. Björkman and Svensson ask whether the people a clinic serves can fill that gap if they are given information and a forum. The short-run results were remarkable. Later studies tell a more mixed story, covered at the end of this guide.
Who holds health workers to account?
Covers Section I, the introduction
About eleven million children under five die each year, almost half of them in sub-Saharan Africa. More than half of these children die of diseases such as diarrhea, pneumonia, malaria, measles and neonatal disorders that a small set of proven, inexpensive services could prevent or treat. Why are those services not provided? Anecdotes, and increasingly systematic evidence, point to weak accountability: health workers who are absent, and public funds and drugs that go missing.
A useful frame comes from the World Bank’s World Development Report 2004, which the paper cites for evidence on failing services. The report contrasts a “long route” of accountability, in which citizens choose politicians who oversee providers, with a “short route,” in which clients hold providers to account directly. This paper is about the short route: can the users of a rural clinic make its staff serve them better?
Earlier attempts were discouraging. Olken (2007) found only minor effects of encouraging community participation in monitoring corruption in Indonesia. In Rajasthan, India, Banerjee, Deaton and Duflo (2004) studied a project that paid a community member to check whether the nurse-midwife was at the health center. Attendance did not improve, perhaps because one person’s information did not move the wider community to act. Björkman and Svensson designed their intervention against these problems: it was structured to limit capture by local elites, it used basic facts about services drawn from the community’s own experience, and it tackled two constraints at once, missing information and too little participation, by involving many people in agreeing on a plan.
The paper also differs from the medical field trials it compares itself with. Those trials ask what a treatment achieves when health workers competently do their jobs. This one asks how to get health workers to do their jobs at all.
A year after the first meetings, how much lower do you think under-five mortality was in the program communities than in the control communities?
A year after the first round of meetings, the treatment communities were more involved in monitoring their clinic, health workers were more often at work and examined patients more carefully, outpatient numbers were 20% higher, infants were heavier by 0.14 of a standard deviation in weight-for-age, and under-five mortality was a third lower. Using variation across districts, the authors also show that bigger gains in monitoring went with bigger gains in outcomes.
Bottom line. Give citizens information and a forum, and they may succeed where distant supervisors fail. The rest of this guide asks how solid each link in that claim is.
Uganda’s dispensaries
Covers Section II, the institutional setting
Uganda had a functioning health system in the early 1960s. It collapsed in the political upheaval of the 1970s and 1980s, and health indicators fell sharply until peace returned in the late 1980s. Since then the government has been rehabilitating the public health sector. There are four types of facility, hospitals, health centers, dispensaries and aid posts, and each may be government-run, private for-profit or private not-for-profit.
The experiment concerns public dispensaries, the lowest tier at which patients meet trained staff. Most are rural, and by the government’s standard they should offer preventive and promotional care, outpatient care, maternity services, a general ward and a laboratory. Public health services have been free since 2001. A typical dispensary in the sample had an in-charge or clinical officer, two nurses and three nursing aides or other assistants.
Who can discipline a dispensary worker?
Tap each body in the chain of oversight.
As the paper notes in Section III.A, health workers also have few financial reasons to work hard. Public money does not follow patients, and hiring, pay and promotion depend mainly on seniority and qualifications rather than performance. A worker may still work hard because shirking goes against her own ideals, and community members can add social rewards for good work and social sanctions against shirking. Those social levers are the main tools a community has.
Bottom line. Formal supervision is thin, and the people with the power to sanction are far away. The clinic’s users, who have the most at stake, have only a few representatives on a committee that cannot sanction anyone.
The experiment
Covers Sections III.A, III.B and III.C: the overview, the design and the data
The project began in 2004 as a pilot of citizen report cards, designed by staff from Stockholm University and the World Bank and run with Ugandan practitioners and 18 community-based organizations (CBOs). Its premise was that two constraints held communities back. The first was information. A mother knows whether her own child died and whether the clinic helped, but not how many children in the community died before age five, where people usually seek care, or what the community can reasonably expect from its clinic. The second was coordination. Monitoring a public clinic is a public good, so each person is tempted to leave it to others, and people had not agreed on what it was reasonable to demand.
If the program got communities monitoring, the change the authors expected was in staff effort: more of it, encouraged by social rewards and sanctions.
Design
Fifty public dispensaries in nine districts, covering all four regions of Uganda, took part, all of them rural. A facility’s community, or catchment area, is everyone living within five kilometers of it.
What a catchment area looks like
Illustration: the teal square is the clinic, and each dot stands for 10 households, placed at random. The counts are the paper’s averages: about 2,500 households within five kilometers of the clinic, 350 of them within one kilometer. The inner circle holds 14% of households on 4% of the area, so settlement clusters near the clinic.
The facilities were grouped first by district and then by population size, and within each group half were randomly assigned to the program. That gave 25 treatment and 25 control facilities, with each district contributing to both groups.
Data
Two surveys were run before the program, and their findings formed the report cards. Both were repeated a year later. The facility survey took its data from the clinics’ own records, such as daily patient registers and stock cards, rather than from administrative reports that staff might have reasons to misstate, and enumerators made visual checks. A stratified random sample of households in each catchment area, roughly 5,000 households per round, answered questions about their health and their experience of the clinic. In all, 88% of baseline households were interviewed again, and the rest were replaced. Where possible, patient exercise books and immunization cards backed up the answers. The follow-up survey added under-five deaths and the weights of all infants.
Bottom line. A stratified randomization of 50 clinics, with outcomes taken from records and checks that are hard to manipulate. The limitation is the sample: 50 clinics is not many.
Three meetings and a contract
Covers Section III.D, the intervention, and Figure I, the timeline
Each treatment facility and its community received a report card summarizing the baseline surveys for their area, translated into the main local language. Staff from local CBOs, trained for seven days in interpreting data, participatory methods and conflict resolution, then facilitated a series of meetings. Step through it:
From report card to community contract
Timeline redrawn from Figure I. Filled circles mark the current step; the shaded bar is the year in which communities monitored their clinic on their own.
The 18 CBOs mostly worked on health, including health education and HIV/AIDS prevention, and before the program they had been active in 64% of the treatment communities and half of the control communities. Section V.F asks whether they, rather than community monitoring, could explain the results.
Bottom line. The program added no money, drugs or staff. It added information, a forum and an agreed plan, and then left the community in charge.
Measuring impact
Covers Section IV: outcomes and the statistical framework, equations (1) to (3)
The paper looks for effects at every link of the accountability chain. Did communities monitor more? Did staff behave differently? Did more people use the clinic? Did health improve? It also tests alternative explanations: spillovers to control areas, changed behavior by other agents such as the health subdistrict, and effects of the meetings or the CBOs that bypass community monitoring. Spillovers would bias the treatment-control comparison itself. The other channels would not undo the causal effect of the program, but each would change what it means.
Equation (1), term by term
Tap any part of the equation.
Tap a term to see what it does.
Where an outcome was measured both before and after the program, the paper also stacks the two rounds and estimates a difference-in-differences model, equation (2): yijt = γPOSTt + βDD(Tj × POSTt) + μj + εijt. The estimate βDD is the change in treatment areas minus the change in control areas. The facility fixed effect μj removes everything constant about a clinic, so the estimate does not rely on the two groups having started out alike.
Families of outcomes
Many outcomes come in families, such as four monitoring tools or four kinds of clinic visit. Following Kling and coauthors, the paper estimates each family jointly as a seemingly unrelated regression system, equation (3), and reports an average standardized effect: the average of β̂k/σ̂k across the family’s K outcomes, where, in the paper’s words, β̂k and σ̂k are “the point estimate and standard error.” You can check what that does to the numbers:
What is in an “average standardized effect”?
Dividing each estimate by its standard error turns it into a t-ratio, which mixes the size of an effect with how precisely it is measured and grows with the sample. An average of t-ratios is still a reasonable test of whether a family of outcomes moved, and its significance stars stand, but its size is not a number of standard deviations. In the toolkit the paper cites (Duflo, Glennerster and Kremer 2007) and in Kling, Liebman and Katz (2007), each estimate is divided by the control group’s standard deviation of the outcome, which makes the average an effect size. The standard errors the paper reports for these averages fit that reading. Each t-ratio has a standard error of about 1, so an average of them has a standard error of at most about 1, smaller the more outcomes it averages and the less they move together. Every one the paper reports lies between 0.31 and 0.89.
Standard errors and confidence intervals in this guide
Facility regressions have at most 50 observations and use heteroskedasticity-robust standard errors; difference-in-differences regressions on clinic records stack two rounds (100 observations) and cluster by facility; household regressions cluster by catchment area. The 95% intervals drawn in this guide are the estimate plus or minus 1.96 standard errors, and rows are colored by the paper’s own significance stars, with solid dots for results significant at 5% or better.
Bottom line. Three tools: a cross-section with district effects, a difference-in-differences where baseline data exist, and family summaries. Read the summaries’ sizes as average t-ratios, not standard deviations.
Were the groups alike?
Covers Section V.A, differences before the intervention (Table I)
Randomization should make the two groups alike on average, but with 25 clinics in each group, chance differences can be sizable. Table I compares key baseline characteristics and, for eight families of baseline measures, average standardized differences.
Before the program (Table I)
No difference is significant at the 10% level, and a joint test that all eight family averages are zero gives χ² = 4.70, with p = .79. The largest raw gaps go in opposite directions: control clinics saw more outpatients (675 against 593 a month), treatment clinics more deliveries (10.3 against 7.5). One imbalance surfaced later: a replication found that children in treatment communities were already more likely to be vaccinated at baseline (Section V.C).
Bottom line. On the measures reported, randomization produced comparable groups, with one later caveat about immunization.
Did communities start monitoring?
Covers Section V.B, processes (Table II)
After the first meetings it was up to each community. To avoid influencing them, the researchers sent no outside agents, so they could not document everything that happened. Reports from the CBOs and a survey of local councils describe a process run mainly by local councils, health unit management committees and community members. A typical treatment village held about six local council meetings in 2005, and in 89% of villages those meetings discussed the project facility, mostly the community contract or parts of it, such as staff behavior.
Committees seen as ineffective were replaced. More than a third of the health unit management committees in treatment communities were dissolved and new members elected, against none in control communities. According to the CBOs, community members also used their visits to the clinic to reward staff, or to question them about parts of the contract that had not been met. Enumerators’ checks a year later show the tools of monitoring in place:
Program impact on monitoring and information (Table II)
The household survey points the same way: more households in treatment communities discussed the clinic at local council meetings, and more had heard about the management committee’s role. Both the amount of discussion and its subject changed, from general discussion to the specific terms of the contract.
Bottom line. Communities did act. Suggestion boxes, waiting cards and posters appeared, and more households discussed the clinic at village meetings.
Did health workers change?
Covers Section V.C, treatment practices (Tables III and IV)
If communities monitored more, did staff work harder? Table III looks at how patients were examined, how long they waited, whether staff were at work, the state of the clinic, preventive advice and drug supplies.
Program impact on treatment practices (Table III)
Supplies of drugs did not differ between the groups (Section V.F), yet control clinics ran out more often while treating fewer patients. The authors read this as evidence that more drugs leaked out of control clinics.
Immunization
Using the national immunization schedule, the authors code whether each child had received the doses due for its age of DPT, BCG and polio vaccines, vitamin A and, from age one, measles. They then average across these for each age group:
Immunization by age group (Table IV)
Bottom line. Staff were more often present, used equipment more, kept clinics cleaner and gave more preventive advice. Several estimates are imprecise, and much of the immunization gap may predate the program.
More patients
Covers Section V.D, utilization and coverage (Table V)
Better service should draw more patients. Clinic records give monthly numbers of patients, and the household survey shows where people went when they fell ill. Where baseline data exist, for outpatients, deliveries and the household measures of where people sought care, a difference-in-differences model is also possible.
Program impact on use of the clinic (Table V)
With baseline data, the outpatient estimate is larger both in absolute terms and relative to its standard error (a t-ratio of 2.8 against 2.1). Deliveries are a different story: treatment clinics already had more of them before the program (10.3 against 7.5 a month), and the difference-in-differences estimate, +38%, is smaller and significant only at 10%. Households in treatment communities also cut visits to traditional healers and self-treatment, with no significant change in their use of other providers: they switched to the project clinic.
Bottom line. Outpatient numbers rose by a fifth or more, deliveries by 38% to 58% depending on the model, and people moved from self-treatment and healers to the clinic.
Fewer deaths, heavier babies
Covers Section V.E, health outcomes (Table VI and Figure II)
Better and busier clinics could improve health in several ways: more sick people treated, a switch away from self-treatment, better care for those treated, more immunization and more preventive advice. For a country with Uganda’s disease profile, 73% of under-five deaths are estimated to be preventable with proven interventions. Community-based medical trials have cut under-five mortality by 30% to 50% within one or two years, among them home treatment of malaria in Tigray, Ethiopia (40%) and pneumonia case management in India (30%). Those trials introduced a new treatment, though. This program introduced none and left the supply of health inputs unchanged.
Deaths, births and pregnancies (Table VI)
Which children were spared? Chance of dying in 2005, by year of birth
Each row is the program’s effect on the chance that a child born in that year died during 2005 (control-group mean across ages: 2.9%). Children under two drive the result. For children born during 2005, the paper’s estimate implies a 35% lower chance of dying.
Weight
A weight-for-age z-score compares a child’s weight with the median for children of the same age in a reference population (the 2000 CDC growth reference), in units of that population’s standard deviation. Ugandan infants are far below the reference, and the gap widens with age. Among 1,135 children under 18 months, the program raised weight-for-age by 0.14 (0.14 again with controls for age and sex; 0.16, with standard error 0.09, when outliers are kept).
Weight-for-age, program against control (Figure II)
The gap is clearest among underweight children. The authors read this as better treatment of sick children rather than a general rise in nutrition: underweight children have weaker defenses against infection, fall ill more often, and so need health care more. How much could 0.14 matter? The paper takes the control group’s shares of mildly, moderately and severely underweight infants and applies estimates that their risk of dying from infectious disease is about two, five and eight times that of other children.
From weight to survival
lower average risk of death
underweight (below −2), before and after
Bottom line. Under-five mortality fell by a third and infants weighed more, especially those most likely to need care. The mortality estimate is imprecise, and a much smaller effect is also consistent with the data.
Is monitoring the reason?
Covers Section V.F, getting inside the box and robustness tests (Figure III and Table VII)
Large effects on both monitoring and outcomes fit the community-monitoring story, but do not prove it. Following Kling, Liebman and Katz (2007), the authors ask whether districts where the program raised monitoring more also saw bigger gains in outcomes. They build a monitoring index, the first principal component of the six measures in Table II, and estimate
by two-stage least squares, using the treatment-by-district interactions as instruments for the index M and controlling for district effects. If M is the channel through which treatment works, δ is consistently estimated. With only district indicators as controls, δ is simply the slope of a line through the district-level points below.
Bigger monitoring gains, bigger outcome gains (Figure III)
Two-stage least squares estimates of δ (Table VII)
Adding a treatment dummy is a stricter test: the index is then identified only by differences across districts in how much monitoring rose. The dummy itself is small, with the wrong sign, which suggests the program worked through monitoring rather than some other route. Adding the in-charge’s knowledge of patients’ rights, a proxy for staff responding directly to what they heard in the meetings rather than to community pressure, leaves the monitoring coefficients largely unchanged.
Other explanations, and what the paper finds
Bottom line. The cross-district pattern and the robustness checks favor community monitoring as the channel. With nine districts and 50 clinics, the evidence is suggestive rather than decisive; the authors themselves call the staff-knowledge test “not conclusive.”
Worth the money?
Covers Section VI, the discussion
The program ran in nine districts with about 55,000 households in the treatment catchment areas, so in that sense it has already operated at scale. The authors stress what is still unknown: long-term effects, effects across sectors, whether bottom-up monitoring works better combined with reformed top-down supervision, and a full cost-benefit analysis. For the last, they offer a back-of-the-envelope calculation for child deaths.
Cost per under-five death averted
deaths averted in a year
program cost per death averted
The slider’s marks at 8% and 64% are the ends of the paper’s 90% confidence interval. The paper does not show its arithmetic. Multiplying 55,000 households by the control group’s 0.21 births per household (Table VI) and by the fall in deaths per 1,000 reproduces its figure of about 550 deaths averted at 33%.
Two caveats cut in opposite directions. The calculation sets one year of averted deaths against the whole program cost, so if the effects last, as the follow-up study below found, the cost per death falls. But it rests on the point estimate: at the low end of the confidence interval the program costs more per death averted than the $887 benchmark.
- Communities monitored more
- Yes. New tools in clinics, the clinic on village agendas, committees replaced (Table II and Section V.B).
- Health workers changed
- Mostly. Less absence, more equipment use, cleaner clinics, more advice; some estimates imprecise (Table III).
- More people used the clinic
- Yes. 20% to 29% more outpatients and 38% to 58% more deliveries, depending on the model (Table V).
- Health improved
- Yes, imprecisely. A third lower under-five mortality, heavier infants (Table VI).
- Monitoring was the channel
- Suggestive. Cross-district 2SLS and robustness checks (Table VII, Figure III).
Bottom line. A cheap program that gave communities information and a forum produced large short-run gains in a randomized trial of 50 clinics. The gains in care look robust; the gains in health are less certain, and a much larger later version did not repeat them.
Test yourself
Eight questions on the whole paper
Glossary and citation
Key terms, and how to cite the paper
- Absence rate
- Share of staff on the clinic’s baseline employee list who were not physically present at an unannounced visit, leaving out staff reported to be on outreach.
- Average standardized effect
- A summary of a family of outcomes: the average of each estimate divided by a scale. In this paper the scale is the standard error, which makes it an average t-ratio; in the methods it cites, the scale is the control group’s standard deviation.
- Catchment area
- The community a clinic serves, defined here as everyone within five kilometers of it.
- Community-based organization (CBO)
- A local nongovernmental organization. Eighteen of them facilitated the meetings in this experiment.
- Community contract
- The shared action plan agreed at the interface meeting between community representatives and health workers: what to do, how, when and by whom, and how the community would monitor it.
- Difference-in-differences
- The change in an outcome in the treatment group minus the change in the control group, which removes fixed differences between them.
- Elite capture
- When local leaders or better-off groups take control of a participatory process. The community meetings invited a cross-section of residents to limit it.
- Health Unit Management Committee (HUMC)
- The committee of health workers and nonpolitical community representatives that should oversee a dispensary’s daily running. It can monitor but not sanction staff.
- Principal components analysis
- A way of combining several correlated measures into one index, the weighted combination that captures as much of their joint variation as any single combination can.
- Seemingly unrelated regression (SUR)
- Estimating several regressions jointly so that their estimates’ correlations are known, which allows tests about the family as a whole.
- Spillover
- An effect of the program on the control group, for example through news of better service spreading to nearby communities.
- Stratified randomization
- Randomizing within groups of similar units, here districts and then population size, so that treatment and control are balanced on those characteristics.
- Two-stage least squares (2SLS)
- An instrumental-variables method. Here it relates outcomes to the part of the variation in monitoring that comes from random assignment in each district.
- Under-five mortality rate (U5MR)
- Deaths of children before age five per 1,000 live births. In this paper, the sum of 2005 death rates for each one-year age group in a community.
- Weight-for-age z-score
- A child’s weight minus the reference median for its age, divided by the reference standard deviation. Below −2 is the usual definition of underweight; the paper also treats −2 to −1 as mildly underweight.
Björkman, M., & Svensson, J. (2009). Power to the People: Evidence from a Randomized Field Experiment on Community-Based Monitoring in Uganda. The Quarterly Journal of Economics, 124(2), 735–769. https://doi.org/10.1162/qjec.2009.124.2.735
This guide paraphrases the published paper and redraws its results from its tables. Values for Figures II and III were read from the vector drawings in the paper’s PDF, so they match the printed figures to within plotting precision; the weight-category shares and the mortality-risk calculation use those readings. The catchment illustration, the cost calculator’s range and the risk calculator for shifts other than 0.14 are illustrations, not results from the paper. Points a careful reader may notice: the paper’s formula for its “average standardized effects” divides each estimate by its standard error, and the printed values match averages of t-ratios (our reconstruction); a later replication, which describes the error as the use of treatment-group rather than control-group standard deviations, confirms that the published sizes are too large. The text refers to absenteeism as “row (3)” of Table III when it is row 5, to the cross-sectional waiting-time estimate as “column (4)” when it is row 4, and to the household summary effects as “specifications (8) and (9)” when they are (8) and (14). It also calls the cross-sectional equipment estimate “less precisely estimated,” though its standard error is smaller than the difference-in-differences one; it is the estimate that is smaller. The 33% fall in mortality compares raw group averages (144 and 97), while the regression estimate is −49.9 with standard error 26.9, from which a 90% interval of roughly a 4% to 65% reduction follows, a little wider at the bottom than the 8% to 64% the text reports. Figure III’s lower panel labels its axis “infant mortality rate” though its caption and the text describe under-five mortality, and its outpatient line has a slope of about 0.64, against 0.77 in Table VII, which also controls for baseline outpatients. Later research cited: Björkman Nyqvist, de Walque and Svensson (2017), “Experimental Evidence on the Long-Run Impact of Community-Based Monitoring,” American Economic Journal: Applied Economics 9(1): 33–69; Donato and Garcia Mosqueira (2019), “Information Improves Provider Behaviour: A Replication Study of a Community-Based Monitoring Programme in Uganda,” Journal of Development Studies 55(5): 967–988, and 3ie Replication Paper 11 (2016); Raffler, Posner and Parkerson (2026), “Can Citizen Pressure Be Induced to Improve Public Service Provision?” Journal of Politics.
