I think this could be a great intervention, but I disagree on the RCT front. I strongly believe that in the medium to long-term, this intervention should, and will only go to scale after high quality RCTs.
Yours and RP's reasons for not doing it aren't compelling to me. I'm almost a bit confused.. This might actually be one of the easier RCTs to do in the development world. If I was a funder, I would be asking why you aren't in the middle of a big RCT at the moment, or at least why you aren't getting started on one. I agree with the funders.
"Funders often tell us to come back with an RCT, which we estimate would need around 72 hospitals across three countries, US$5 million and three to four years." If this is what's needed, then yeah do it. You can get the 5 million - have you asked GiveWell? But I'm skeptical you really need 72 hospitals. Mortality rates are high enough in massive hospitals I would imagine if you are in national referral hospitals in capital cities, or similar, then this number would drastically drop. Maybe you are considering a bunch of smaller hospitals?
From your article "conventional RCT is hard to run for this intervention. It can't be blinded, because health workers know when monitors are on the ward. It has to randomise whole hospitals, because the central patient overview and the positive impact on workload affects care for every child on the ward, not only those attached to a monitor. Death is a relatively rare outcome, and baseline neonatal mortality across our sites ranges from 4.3% to 27.8%, which makes a trial harder to power. Ministries are also reluctant to keep hospitals in a control arm for years while neighbouring hospitals have monitoring. So is an RCT needed before scaling? If the effect is real, waiting three to four years for trial results means children in hospitals that could have had monitoring go without it. Last but not least it is important to emphasize that continuous monitoring is already considered standard of care in most parts of the world, so perhaps it is less about if it works but rather to what extent it should be prioritized amidst many other priorities and gaps.
From RP "Note that a traditional randomized controlled trial may be genuinely hard to run in this context: a monitoring system cannot be disguised as a placebo, and it may change how the whole ward operates. For example, staff know when the monitors are present, and the intervention is bundled with training and supervision. The observed effect size therefore reflects both the device plus the attention surrounding it (i.e., the Hawthorne effect[5]),
1. Your current evidence is fairly weak, and low on the evidence ladder. Before and after studies are notorious for being unreliable because there are so many other factors involved. Hospitals are always improving their systems and there are many reasons why mortality could drop. Seeing consistency between studies is encouraging but evidence needs to be better.
2. Blinding is not very important - the staff seeing the monitors is part of the intervention no? I'm not sure why RP and you think this is a problem? If it changes how the whole ward operates that is part of the benefit of the intervention. In fact that might be where you main benefit is! I'm confused about this comment by RP here and not sure what point they are making.... The training and supervision front could be a big deal though, especially when you believe your main impact to be ward spillovers. That's why I think training and supervision can't really last more than a few months, then you need to largely leave the hospital be with the monitors with minimal supervision and see if mortality reductions are sustained over 2-3 years. This isn't that long an RCT!
4. Randomising whole hospitals isn't an ethical issue. Ministries will not refuse to allow hospitals being in a control arm (they are getting free monitors) and the evidence isn't nearly good enough yet to argue an RCT is unethical. If the difference in mortality between hospitals is big after year 2 say, the study will be triggered to stop for ethical reasons anyway.
5. The range of mortality rates doesn't make powering a study harder I don't think? (might be missing something) Its low mortality rates and low sample sizes that might be the issue here. But there are lots of hospitals where multiple kids die every week (I have one 500 meters from my home) If you are in the busiest hospitals where 150+ kids die every year, then I would think you wouldn't need neadly that amount (I might Claude this later). If you need to do a multi-country RCT so be it, that seems likely needed.
6. The urgency argument seems unrealistic. Because of the low level of evidence intervention isn't "obviously" better than other things we could spend money on. I don't think your case is nearly strong enough here to make a "bypass the RCT" argument. I think actually even framing it like this could make people lose confidence in the intervention. It feels more like a sales pitch than a good argument.
7. What is considered "Standard of care" in rich countries is close irrelevant in healthcare interventions in low-income countries. I don't think it should even be a factor really. There are 100 "standard of care" things you could introduce to a pediatric ward in Malawi (Better antibiotics, Dr. nos, nursing ratios, frequency checks, infection control, blood cultures, blood gases etc.). This doesn't mean you don't study their benefit in a low-income setting.
As an unrelated point, I feel like the presentation of the intervention here might be too branded. The intervention is a monitoring device, yet you are at pain to emphasise how your intervention is different from other devices. If your plan is to scale through government, its the monitoring system that matters not the branding. It might be better even to brand machines as "Ministry of Health" machines and think of yourself as a support org as much as an innovation org? This is a tricky one though I know.
I absoutely love this intervention - I'd be keen to get on a call and chat about it even if you might be. Sadly someone ffrom your org was pegged to meet me at EAG New York, but I can't go because of travel restrictions due to ebola :(
Hi Nick, thanks for your message. Happy to jump on a call myself as well to dive deeper feel free to email ([email protected]) or text via the chat function of the forum (if it has that)
as a first quick reflection/response, hopefully I have more time later this week but here are a few reflections, and apologies for not following your exact numbering.
A) we have been trying over the past 4 years to get trials funded (EDCTP, givewell, EU funding, Gates) without success. There are some ways to make it lower cost (lowest we got was about 2,5M, but that also reduces quality and strength of evidence). In between time we have done what we could through different smaller sources of funding to get the best evidence possible. Agree by the way that blinding is not needed and not relevant and that a large effect actually comes from the behavior change due to its clear presence on the ward. Also randomization is not an issue as long as control sites also get the intervention.
B) trials can be great, but if we wait for those results to come out before scaling we would likely not survive as an organization. This is not unique for us, but for all innovations and a major bottleneck for innovative solutions in general. Moreover, 3 years of waiting to scale and finding out that the results are same, would create a huge amount of deaths not averted, happy to provide more precise estimates at a later stage. Or plea is not to skip the trial, our plea is to do it in parallel.
C) While I like trials I actually think real world evidence is better or at least as important. With monitoring and many other interventions it is not about if they work, but how do they work in a very complex environment that is under-resourced and understaffed. In other words, if I would be the minister of health, I would value evidence from 100% of 50 sites in a real world setting a lot higher than the evidence from the intervention under trial circumstances. Just to be clear --> we are doing all this (we are currently running a controlled interrupted time series in multiple countries and multiple other impact evaluations are still ongoing) and still want to do the trial if we can get the funding for it.
D) A trial as the definitive answer to answer the question if something should be scaled is not rational from a perspective of levels of uncertainty. In my view the lower bound of the margin of certainty just needs to be above (or close to) the treshhold because if it is above that level it would not be rational to say we need a trial first before scaling it up, because the statistics already say it is highly unlikely that it is. I sometimes use the picture at the bottom to describe that (NB if evidence changes the direction of the arrow may also change). Of course it is good to discuss if you agree to the estimated impact and the margins of uncertainty to that which Rethink priorities found, because I do think that is important to agree on (or disagree and find ways to close the gap)
E) too salesy --> I think you may underestimate how important the technology, service-model and business model are compared to anything else that is out there. There are dozens if not hundreds of changes we have done and are continuously doing to improve the intervention (at GOAL 3 level) and at the facility (to improve adoption and usage at facility level) all embedded in the business model which allows us to do continue to do this over time. We do not describe this here in this post to sell or brand our solution, but also to make clear that just putting a monitor down will not have the same effect.
Last but not least: it is also a product of passion and enthusiasm that comes from building 8+ years towards this solution. I just can't be fully objective in how I describe our solution. Firstly because it is the answer to the problems I experienced myself when working in the field, secondly I think there are a lot of lessons to learn from our success that can be used by others. I hope you can accept that (and maybe even appreciate it ;-) )
F) thanks again. Really appreciate the open discussion. Would be great to connect
Nice one, some quick replies and lets set up a call. So great you are so passionate about this, and I for one think that your org should be well funded and an RCT should be done.
A) Sorry you can't get the RCT funded. Have you tried the new DIV? You're right randomising at hospital level makes it really tough. I feel like you probably need national referral hospitals for the required power. From Claude if you had hospitals roughly with 1,000 admissions a month at 3% mortality, looking for a 20% reduction in mortality you'd need 4-5 massive (national referral) hospitals, or maybe 2 massive ones and 15-20 smaller? 72 feels like too much but I obviously haven't looked into the details, just asked Claude. As your current results seem to show 30ish percent mortality reduction this seems on the difficult but feesable end?
B) I think you are in a great position continue to grow steadily if not scale without better evidence and survive - your organisation is still so young there's plenty of room to grow and then scale. You are in a pretty good position for philanthropy to scale the setup costs (rather than government) if an RCT shows it is this amazing.
C) This one confuses me I don't really understand what you mean by "Real world evidence". Your RCT should ideally be basically just what you do anyway, so wouldn't that be real world evidence? The problem is that you're looking at these mortality drops as evidence enough, where I would see them more as higher level M&E of your program. Your own M&E can never touch an external study for importance. Orgs like ANSH used mortality drops as a marker for their progrram, but they already know the intervention works from RCTs. I just dont' think it can fly for before/after studies or your own M&E to be the evidence which carries you to scale.
D) I like your chart a lot and I agree with the concept completely! But I don't agree with the levels of uncertainty RP cites, although at least they are pretty big. Given the variation even in your own data, relying on before/after studies, the nature of the effect of these kind of interventions to reduce over time (I think a 7 year horizon is too long for this CEA), and the general trend of hospital mortality improving over time I would put your error bars much wider. Claude told me Kamuzu hospital mortality dropped fom 9% to 3% between 2012 and 2015 of its own accord, so these kind of drops seem not unusual for a bunch of reasons. Also Claude and RP cited a bunch of studies in Western pediatric wards which showed benefits in before/after studies, but then when high quality RCTs were pooled the benefit disappeared. I would have anchored on the best quality research we have on the topic much more heavily than RP did, even though its in a different context. I think if the evidence isn't clear in high income countries, it increases the burden for you somewhat.
A) We want it to be representative so only selecting referral hospitals does not do the job as it will not inform a ministry of health what will happen in rural areas. The challenge is that hospitals vary quite a bit, so you have to match them first for certain characteristics and then randomly assign them. The larger the variation the more hospitals you will need, hence we came at 72 through our power calculations. The challenge is most funds only go up to 2 or 3 million USD
B) I think you are right from where we are today, but there have been many moments where it was likely we would not survive. We have our own development team, registrations to maintain all requiring significant investments. If it takes too long, likelihood of dying in the valley of death will increase. There are very few innovators who are actually successful in public health systems for this reason.
C) In any research setting there will be research staff and a lot more data collection, this by itself has a huge effect on how care is provided. Moreover not all sites will be suitable for it, creating a significant selection bias. Therefore the performance in (optimal) trial circumstances often has limited meaning for what it means in real life. Moreover variations across countries, will likely influence success of the intervention. So generalizability is also limited (Why would someone in West Africa, South-East asia or the middle east believe the trial from East Africa) Real world evidence solves a large part of this, whether if it is implementation research or derived from the real world evaluation. We look at each installation for usage, adoption and impact on health workers stress and burnout, where possible (if funding allows) we look at impact on mortality. While the quality of the evidence may be lower, it says a whole lot more about what is actually happening within the specific hospital/country and provides opportunities to intervene (either through us, or the hospital itself) Seeing this over multiple hospitals in one country will help a lot more with decision making then a (potentially outdated) trial from 5000km away (as a figure of speech).
Again, not opposing a trial here just wonder if it is the right instrument for the stage we are at. Just to provide a different perspective. A 5 million implementation program, would allow me to go to 200-300 hospitals in different countries and perform a thorough monitoring and evaluation program in parallel which can look at before-after differences, but also at differences between other hospitals. Those results would be available within 1 year after implementation, so likely 2 years after a grant was awarded.
D) Don't think Claude is accurate. Without a verifiable source on the mortality drop that answer is pretty useless, beyond that it is not wise to assume a conclusion that this is common based on one site. Data quality is notoriously bad in this setting and there is severe under reporting. We have solved this by diving into individual patient files hand by hand (not us, but our research partners)
The study you cite is hardly comparable in my view - general wards versus ICU/HDU - High resource settings only - high levels of staffing and no lack of monitoring equipment - Only adults - Vast majority is surgical patients - total number of patients only 1284 over all studies combined - relative rare occurance of the events making them statistically underpowered to proove significant impact
Moreover if you dive into their discussion you also find that over multiple systematic reviews all findings are consistently positive favoring monitoring, yet not significant. That is relevant to our discussion, because one of the key conclusions is that even at 1-2% mortality reduction IMPALA would still be cost-effective against any bar.
Full disclosure: I’m also a co-founder of GOAL 3, so definitely not an unbiased observer here.
What I find an interesting dilemma is how we should balance further RCT/evidence generation vs. scaling a proven intervention (for medical professionals)? What are the ideas on this for the forum?
76x cash transfers: too good to be true or a (missed) opportunity? An invitation for dialogue on a recent Rethink Priorities Report — EA Forum
76x cash transfers: too good to be true or a (missed) opportunity? An invitation for dialogue on a recent Rethink Priorities Report
GOAL 3 is a social enterprise that created IMPALA, a patient monitoring and digital platform system for low-resource settings. There is also a GOAL 3 foundation that raises funds to ensure underserved communities will be reached.
A recently completed report from Rethink Priorities calculated IMPALA to have a cost per DALY averted of $7–17 in paediatric care and $12–30 in neonatal care. This translates to around 76× and 28× unconditional cash transfers respectively for paediatrics and neonates against GiveWell's bar, and roughly 11,700× and 4,400× on Coefficient Giving's scale against their 1,000× bar.
RP corrected for study design and robustness of findings by applying a 60% internal-validity and a 30% external-validity discount to our measured mortality effects, arriving at roughly 12% for paediatric wards and 7% for neonatal units.
RP concluded that even at 1-2% relative reduction, the intervention would still be cost-effective.
The post concludes with an invitation for discussion on what these results mean for further research and scaling of IMPALA.
Introduction
GOAL 3 commissioned the Rethink Priorities Global Health and Development team to provide an external analysis of the impact and cost-effectiveness of our product ‘IMPALA.’ IMPALA is a continuous patient monitor combined with a digital health platform designed for low-resource settings (links with more info in the appendix). The report estimates that IMPALA is highly cost-effective, far surpassing every funding threshold considered. We would like to use this post to discuss the studied intervention, the results, the questions that may arise as a result, and how we can address these now or in the future.
Understanding the problem in its context
Sub-Saharan Africa accounted for 2.8 million of the 4.9 million deaths of children under five worldwide in 2024.[1] Access to care alone does not prevent these deaths. In low- and middle-income countries, an estimated 58% of all amenable deaths from conditions treatable by health care occur among people who did use the health system but received sub standard-quality care.[2] Important root causes to this challenge are lack of adequate equipment and shortage of staff.
In many low-and-middle-income countries (LMICs) there is a huge shortage of staff. WHO uses 4.45 doctors, nurses and midwives per 1,000 people as an indicative minimum for progress on the health-related Sustainable Development Goals.[3] A survey of 47 countries found that the WHO African Region had an average of 1.5 doctors, nurses and midwives per 1,000 people, about a third of WHO's minimum (Malawi was even below 0.5, less than a ninth).[4] Moreover, staff lack the equipment they need to monitor and manage their patients with up to 40% of equipment malfunctioning. Qualitative studies in paediatric high-dependency units in Malawi describe monitoring as intermittent, with nurses forced to focus on the sickest children while staff shortages, power cuts and too few working devices get in the way. [5]
In these conditions, health workers are often unable to recognize and act upon patient deterioration, forcing them into reactive care; leading to late stage response to emergencies when treatments are more costly and more time consuming while outcomes are poor. This sets off a vicious cycle whereby scarce time and resources are used on late-stage escape treatments, further hampering the ability of the health systems to respond in time to changing patient conditions.
Theory of Change
IMPALA enables health workers to monitor patients continuous with health workers receiving alerts when a patient deteriorates. As a result, health workers can recognise patient deterioration earlier and prioritise care more effectively. Monitoring also reduces staff workload. By automating the process of checking vital signs, IMPALA helps save hours of work for health workers every shift[6], allowing them to focus their attention where it is most needed.
Additionally, when patients are treated earlier, complications can be prevented and treatment demands fewer resources. This means less medicine use, less patient costs, and reduced time spent on treating the child by the health workers. The result is reduced health expenditure. [7]
Importantly, IMPALA is not an intervention that increases the burden on already overworked health workers but is embedded into their existing workflows and makes this easier. Each implementation includes on-site training for the ward team, a group of trained IMPALA champions among the hospital's own staff who support colleagues and onboard new starters, and ongoing follow-up from our in-country team through IMPALA Care.
Figure 1: GOAL 3 Theory of Change of IMPALA
What the RP report tells us
Rethink Priorities' Global Health and Development team took on the challenge of evaluating our evidence, commissioned by GOAL 3. RP reviewed nine mortality estimates from hospitals in Malawi, Rwanda and Tanzania. Eight of these point towards lower mortality. [8]
RP's reading is that our evidence consistently points to a mortality benefit but does not establish its size, so its model discounts heavily. RP combined our evidence into weighted averages which led to the number of 42% mortality reduction for paediatric wards and 26% mortality reduction for neonatal units. It then applied a 60% internal validity discount and a 30% external validity discount, in line with how GiveWell treats non-randomised evidence.[9]That left a modelled effect of 12% mortality reduction in paediatric wards and 7% in neonatal units.
RP then estimated cost-effectiveness using its full model (see full model parameters and results here). It puts the full cost to the health provider at $6,100 per ward per year (annualized over IMPALA's lifetime), and assumes 3,000 admissions per year on a paediatric ward and 1,000 on a neonatal unit. On that basis, it estimates that IMPALA costs $7–17 per DALY averted in paediatric care and $12–30 in neonatal care. This translates to around 76× and 28× unconditional cash transfers against GiveWell's bar, and roughly 11,700× and 4,400× on Coefficient Giving's scale against their 1,000× bar for pediatric and neonatal care respectively. The full table from the report can be seen below in figure 2.
Figure 2: Cost-effectiveness of IMPALA from different funder perspectives.
Interpretation, strengths and limitations
A model that returns 76x cash transfers could make any reader sceptical and we understand that. It is important however, to root this in understanding of the intervention and its context. The intervention works (in contrast to many public health interventions like malaria nets or vitamin A) within an environment in which already many resources are deployed. IMPALA allows hospitals to optimize the use of these scarce resources for the patients' benefit by automating repetitive tasks and ensuring essential information is readily available. Because patients have less complications and go home faster it further relieves pressure on the health system which benefits quality of care but also contributes to workload reduction and cost savings. This was found in a cost-effectiveness study recently published in the BMJ Paediatrics Open.[7] We like to think of IMPALA as adding oil to a rusty machine. All components are already there, we just allow them to work more effectively together.
However, we want to be clear about what the evidence does not yet show.
There is no randomised trial as part of our evidence. Most estimates compare a ward before and after IMPALA, so anything else that changed between those periods can look like an IMPALA effect. In our only controlled study, economic and health effects (incl. cyclones, drought, and economic crises) outside of our control may have influenced the results. As the study had a single control per site, RP notes, a shock at that one facility is hard to separate from a real trend.
Several estimates also show little or no effect of IMPALA. Their confidence intervals cross zero, or, the lower bound sits at a 1-2% reduction. Furthermore, large effects measured in small samples tend to shrink as more data comes in. GOAL 3's service model includes ongoing support, which may preserve some of that, but we cannot yet say how much. GOAL 3 has also published its non-significant results and has been transparent with all the data, including one hospital where mortality did not change, which RP took as a sign that they were assessing the full body of evidence rather than a favourable selection.
Reasons the estimate may be too low or too high
May be too low
May be too high
Mortality is the only benefit counted. Averted morbidity is excluded, though on our figures it is 3–4% of DALYs.
The model contains no counterfactual adjustment: no allowance for wards that would have obtained monitoring anyway.
Hospital and family cost savings are excluded. Avoided bed-days are directionally favourable at 8,400-10,920 per hospital per year against an annual system cost of $6,100, though not statistically significant.
No effect decay is modelled across the seven-year equipment life. Our longest follow-up is 24 months. So far we see increasing adoption and usage, but this has not been systematically assessed.
Costs are taken at maximum. No at-scale pricing is assumed, though we expect 20–30% reductions when implementing at volume.
Post-discharge mortality is unmeasured in every study. Some deaths may be deferred rather than averted.
The validity discounts may be too steep, particularly the 30% external validity discount. RP notes that quasi-experimental data is a reason not to discount external validity too heavily, because the intervention was observed in ordinary wards rather than trial conditions.[10]
The largest paediatric result depends on spillover effects. At Zomba, roughly 93% of the children in the population showing the benefit were not on a monitor.
Impact on health worker stress levels and potential health worker retention is not valued. Attrition is a first-order constraint in these systems and we have satisfaction data but no retention data.
Several individual estimates are statistically compatible with no effect.
The results that were currently found were mostly from 2023 and 2024, since then we have improved both the product and the service model to further enhance adoption
Early hospitals may have had an above average level of innovation readiness compared to scaling hospitals (likely already accounted for in applied discounts).
Please note we are not claiming the pros and cons balance. We are setting out both so that readers can apply their own weights.
Simultaneously, we still think the case holds up. As RP notes: “A consistent direction across independent implementations is harder to explain by chance than any one result considered alone.” The body of evidence is wide enough to show that IMPALA consistently points in the direction of mortality reduction across settings.
The starting point also matters. Standard care in the Malawian HDU in these studies is manual observation every six hours. A study in Kenya found hospitals completing only about half of the intended manual checks.[11] Against that baseline, a modest improvement is far more plausible than it would be in a well-resourced hospital.
Finally, the cost-effectiveness calculation from RP does not depend on a large effect. Implausible cost-effectiveness figures usually come from implausible effect sizes. Here the modelled effects of 12% and 7% sit well below every favourable site estimate. As RP also noted, “We believe this is a robust conclusion because it is driven by low cost rather than a large effect: the break-even mortality reduction is only ~1% (pediatric) and ~2% (neonatal), so the conclusion holds even if there is only a very minor mortality benefit.”
For IMPALA to fall below GiveWell's bar, it would have to save almost no lives, and the evidence (not just our own, but also evidence from similar interventions) points the other way.
Discussion: What would it take to fund scale-up now? How much more evidence, and of what kind?
We are publishing this because we want to hear what you as the EA community have to think about this result.
The dilemma. We are in more than 60 hospitals, with about 1,000 monitors. The main binding constraint on starting in new hospitals is capital: governments in Malawi, Rwanda, Tanzania and Kenya are asking for this and will fund recurring service costs (Malawi's Ministry of Health has committed to this in writing), but cannot fund the upfront investment. We also see demand in our own work: hospitals in our installed base asking for paediatric coverage, the Beginnings Fund's neonatal rollout, and Queen Elizabeth Central Hospital asking to expand its use of IMPALA. RP's model shows continuous monitoring to likely be extremely cost-effective. We believe this evidence is enough to answer the question of whether patient monitoring is a cost-effective intervention in low-resource settings. However, funders often tell us to come back with an RCT, which we estimate would need around 72 hospitals across three countries, US$5 million and three to four years.
Additionally, as RP also notes, a conventional RCT is hard to run for this intervention. It can't be blinded, because health workers know when monitors are on the ward. It has to randomise whole hospitals, because the central patient overview and the positive impact on workload affects care for every child on the ward, not only those attached to a monitor. Death is a relatively rare outcome, and baseline neonatal mortality across our sites ranges from 4.3% to 27.8%, which makes a trial harder to power. Ministries are also reluctant to keep hospitals in a control arm for years while neighbouring hospitals have monitoring. So is an RCT needed before scaling? If the effect is real, waiting three to four years for trial results means children in hospitals that could have had monitoring go without it. Last but not least it is important to emphasize that continuous monitoring is already considered standard of care in most parts of the world, so perhaps it is less about if it works but rather to what extent it should be prioritized amidst many other priorities and gaps.
Our take. The opportunity cost of waiting is too high. Decision-relevant gaps are not whether patient monitoring is effective. We see gaps in areas like: heterogeneity across settings, durability beyond 24 months, and post-discharge outcomes. We recognise that funding IMPALA before an RCT is a higher-risk bet. In a recent post Jack Lewars argues that donors don't always need an organisation to have run its own RCT.[12] They can instead look for monitoring, evaluation and learning plans calibrated to the intervention, with outputs, outcomes and accurate costs measured against targets set in advance. He is also clear that this makes a grant riskier, and his first test is whether the intervention itself has RCT evidence, which continuous monitoring in these settings does not yet have. We think our evidence meets his middle-ground standard, and we are working to close the remaining gap. We do think a stepped-wedge RCT, led by the Amsterdam Institute of Global Health and Development (AIGHD), is still a valuable long-run design, but nested in a funded rollout as we scale, not before. We see a trial as one route to stronger evidence; further real-world evidence studies, such as the planned evaluation of our installed base, are another.
Our questions
How much weight would you put on this report from Rethink Priorities? Do you see critical gaps in the analysis (positive or negative)? How much more evidence and what kind would you like to see to say with confidence that IMPALA is a cost-effective health intervention?
If you were a funder, what is the biggest concern you would need answered before supporting IMPALA?
Where does an organisation like GOAL 3 fit in the EA global health landscape? We can think of reasons EA funders might hesitate: our hybrid social enterprise and foundation model, a hardware and health-system intervention rather than a commodity like bed nets, and open questions about who should pay at scale. Which of these, if any, matter most, and are there others we're missing?
Next steps
Let us know what you think
We present what we have for feedback from the community. We are genuinely interested in the questions presented above, and want to hear what you have to say on it. Both us and RP will be available to answer questions in the comments.
Talk to us
We'll be at the organisation fair at EA Global New York, and I'm joining a panel at EAGxBerlin. On Wednesday November 4 at 16:00 CET (10am EST) we will host a webinar together with Rethink Priorities to discuss the report further (sign-up here). You can also email [email protected] to set up a call.
Conflict of Interest
GOAL 3 commissioned and paid for the report; RP had full editorial independence over its conclusions.
This post is published by Niek Versteegde, Founder and CEO of GOAL 3. It is in my interest to promote IMPALA both from a business and impact perspective.
AI Disclaimer
AI (Claude & ChatGPT) were used to help with formatting, spell check and improving general flow. All text was checked before posting.
I want to read more about IMPALA, where can I do that?
Our team has spent the past months making our work, and how we do it, open for anyone to read in depth. You can also hear from us directly: join our next webinar (sign-up here), or email [email protected] and we'll set up a 1:1 call.
Alongside sharing RP's report, we're using this post to officially open theGOAL 3 Public Knowledge Hub on Notion. The Start Here page is the best entry point. The links below take you straight to what you're most likely looking for.
TheEvidence Summary Dashboard goes claim by claim. For each step in our theory of change, it sets out what we claim, which sources support it, how strong we think the evidence is and which questions remain open.
If you want to understand the IMPALA product:
Our product: IMPALA covers the IMPALA Monitor, the IMPALA Platform, the IMPALA Care service model, and how IMPALA differs from other monitoring options.
The costs of IMPALA breaks down the cost of an implementation for a hospital.
If you want to understand the organisation:
Our organization explains where GOAL 3 came from, the problem we work on and our theory of change.
How we work explains our hybrid structure as a social enterprise with a foundation, and our governance.
Absorption capacity sets out how much additional funding we could put to use, financially, operationally and strategically.
Risks, limitations and safeguards are where we set out what could go wrong and how we manage it. If you only have time for one page after this post, we'd suggest this one.
The hub is a living resource, and we update it as new studies come in. If something is missing, unclear or wrong, please tell us in the comments or at [email protected].
Kruk et al. (2018). Mortality due to low-quality health systems in the universal health coverage era: a systematic analysis of amenable deaths in 137 countries. Lancet. https://pubmed.ncbi.nlm.nih.gov/30195398/
Rakers et al. (2024). Cautiously optimistic: paediatric critical care nurses’ perspectives on data-driven algorithms in low-resource settings—a human-centred design study in Malawi. https://link.springer.com/article/10.1186/s44263-024-00108-8
The ninth estimate, from the neonatal unit at Rwamagana Hospital in Rwanda, was essentially flat: a 0.1 percentage-point increase in mortality (95% CI −1.7 to +1.5 pp, p = 0.90), from an uncontrolled before-after comparison. RP interprets this as noise around no effect rather than a sign of harm. Rwamagana also had the lowest baseline neonatal mortality of the neonatal units in our studies, 4.3% before IMPALA, compared with 5.6–27.8% at the other sites, which left little room for a measurable reduction. We think this is the most likely explanation, and it suggests IMPALA's effect may be smaller where outcomes are already relatively good at baseline.
Both discounts are subjective judgements, and RP notes that others could reasonably reach different values. The 60% internal validity discount is in line with how GiveWell treats non-randomised evidence, and given that our evidence is entirely non-randomised, we see it as a reasonable, if cautious, choice. We think the external validity discount is more likely to be too harsh. The studies took place in routine wards across three countries, and IMPALA now runs in more than 60 hospitals in six countries, although RP was only able to assess effects in a small number of them and says it would revisit the discount as evidence from more settings comes in.
Ogero et al. (2018). An observational study of monitoring of vital signs in children admitted to Kenyan hospitals: an insight into the quality of nursing care? Journal of global health. https://doi.org/10.7189/jogh.08.010409
Note: This post was crossposted from the Coefficient Giving Farm Animal Welfare Research Newsletter by the Forum team, with the author's permission. The author may not see or respond to comments on this post.
Subtitle: On vegan advocacy, effective altruism, and FIFA
I turn 40 today. Here are some hot takes.
On factory farming
1. Our biggest challenge is salience. If factory farming led the evening news, it wouldn’t last lo...
Summary
Over the course of a year, the Fieldshaper program develops you into a field-leading grantmaker. You’ll steward $1,000,000 in grant capital, receive a $125,000 annual stipend, and deploy that capital within a sub-field of one of three key areas (farmed animals, global health, or philanthropic infrastructure); building the track record to become a grantmaker who can shape a field.
Role Details
* Title: Grantmak...
Hey this is super cool!
I think this could be a great intervention, but I disagree on the RCT front. I strongly believe that in the medium to long-term, this intervention should, and will only go to scale after high quality RCTs.
Yours and RP's reasons for not doing it aren't compelling to me. I'm almost a bit confused.. This might actually be one of the easier RCTs to do in the development world. If I was a funder, I would be asking why you aren't in the middle of a big RCT at the moment, or at least why you aren't getting started on one. I agree with the funders.
"Funders often tell us to come back with an RCT, which we estimate would need around 72 hospitals across three countries, US$5 million and three to four years." If this is what's needed, then yeah do it. You can get the 5 million - have you asked GiveWell? But I'm skeptical you really need 72 hospitals. Mortality rates are high enough in massive hospitals I would imagine if you are in national referral hospitals in capital cities, or similar, then this number would drastically drop. Maybe you are considering a bunch of smaller hospitals?
From your article
"conventional RCT is hard to run for this intervention. It can't be blinded, because health workers know when monitors are on the ward. It has to randomise whole hospitals, because the central patient overview and the positive impact on workload affects care for every child on the ward, not only those attached to a monitor. Death is a relatively rare outcome, and baseline neonatal mortality across our sites ranges from 4.3% to 27.8%, which makes a trial harder to power. Ministries are also reluctant to keep hospitals in a control arm for years while neighbouring hospitals have monitoring. So is an RCT needed before scaling? If the effect is real, waiting three to four years for trial results means children in hospitals that could have had monitoring go without it. Last but not least it is important to emphasize that continuous monitoring is already considered standard of care in most parts of the world, so perhaps it is less about if it works but rather to what extent it should be prioritized amidst many other priorities and gaps.
From RP
"Note that a traditional randomized controlled trial may be genuinely hard to run in this context: a monitoring system cannot be disguised as a placebo, and it may change how the whole ward operates. For example, staff know when the monitors are present, and the intervention is bundled with training and supervision. The observed effect size therefore reflects both the device plus the attention surrounding it (i.e., the Hawthorne effect[5]),
1. Your current evidence is fairly weak, and low on the evidence ladder. Before and after studies are notorious for being unreliable because there are so many other factors involved. Hospitals are always improving their systems and there are many reasons why mortality could drop. Seeing consistency between studies is encouraging but evidence needs to be better.
2. Blinding is not very important - the staff seeing the monitors is part of the intervention no? I'm not sure why RP and you think this is a problem? If it changes how the whole ward operates that is part of the benefit of the intervention. In fact that might be where you main benefit is! I'm confused about this comment by RP here and not sure what point they are making.... The training and supervision front could be a big deal though, especially when you believe your main impact to be ward spillovers. That's why I think training and supervision can't really last more than a few months, then you need to largely leave the hospital be with the monitors with minimal supervision and see if mortality reductions are sustained over 2-3 years. This isn't that long an RCT!
4. Randomising whole hospitals isn't an ethical issue. Ministries will not refuse to allow hospitals being in a control arm (they are getting free monitors) and the evidence isn't nearly good enough yet to argue an RCT is unethical. If the difference in mortality between hospitals is big after year 2 say, the study will be triggered to stop for ethical reasons anyway.
5. The range of mortality rates doesn't make powering a study harder I don't think? (might be missing something) Its low mortality rates and low sample sizes that might be the issue here. But there are lots of hospitals where multiple kids die every week (I have one 500 meters from my home) If you are in the busiest hospitals where 150+ kids die every year, then I would think you wouldn't need neadly that amount (I might Claude this later). If you need to do a multi-country RCT so be it, that seems likely needed.
6. The urgency argument seems unrealistic. Because of the low level of evidence intervention isn't "obviously" better than other things we could spend money on. I don't think your case is nearly strong enough here to make a "bypass the RCT" argument. I think actually even framing it like this could make people lose confidence in the intervention. It feels more like a sales pitch than a good argument.
7. What is considered "Standard of care" in rich countries is close irrelevant in healthcare interventions in low-income countries. I don't think it should even be a factor really. There are 100 "standard of care" things you could introduce to a pediatric ward in Malawi (Better antibiotics, Dr. nos, nursing ratios, frequency checks, infection control, blood cultures, blood gases etc.). This doesn't mean you don't study their benefit in a low-income setting.
As an unrelated point, I feel like the presentation of the intervention here might be too branded. The intervention is a monitoring device, yet you are at pain to emphasise how your intervention is different from other devices. If your plan is to scale through government, its the monitoring system that matters not the branding. It might be better even to brand machines as "Ministry of Health" machines and think of yourself as a support org as much as an innovation org? This is a tricky one though I know.
I absoutely love this intervention - I'd be keen to get on a call and chat about it even if you might be. Sadly someone ffrom your org was pegged to meet me at EAG New York, but I can't go because of travel restrictions due to ebola :(
Hi Nick, thanks for your message. Happy to jump on a call myself as well to dive deeper feel free to email ([email protected]) or text via the chat function of the forum (if it has that)
as a first quick reflection/response, hopefully I have more time later this week but here are a few reflections, and apologies for not following your exact numbering.
A) we have been trying over the past 4 years to get trials funded (EDCTP, givewell, EU funding, Gates) without success. There are some ways to make it lower cost (lowest we got was about 2,5M, but that also reduces quality and strength of evidence). In between time we have done what we could through different smaller sources of funding to get the best evidence possible. Agree by the way that blinding is not needed and not relevant and that a large effect actually comes from the behavior change due to its clear presence on the ward. Also randomization is not an issue as long as control sites also get the intervention.
B) trials can be great, but if we wait for those results to come out before scaling we would likely not survive as an organization. This is not unique for us, but for all innovations and a major bottleneck for innovative solutions in general. Moreover, 3 years of waiting to scale and finding out that the results are same, would create a huge amount of deaths not averted, happy to provide more precise estimates at a later stage. Or plea is not to skip the trial, our plea is to do it in parallel.
C) While I like trials I actually think real world evidence is better or at least as important. With monitoring and many other interventions it is not about if they work, but how do they work in a very complex environment that is under-resourced and understaffed. In other words, if I would be the minister of health, I would value evidence from 100% of 50 sites in a real world setting a lot higher than the evidence from the intervention under trial circumstances. Just to be clear --> we are doing all this (we are currently running a controlled interrupted time series in multiple countries and multiple other impact evaluations are still ongoing) and still want to do the trial if we can get the funding for it.
D) A trial as the definitive answer to answer the question if something should be scaled is not rational from a perspective of levels of uncertainty. In my view the lower bound of the margin of certainty just needs to be above (or close to) the treshhold because if it is above that level it would not be rational to say we need a trial first before scaling it up, because the statistics already say it is highly unlikely that it is. I sometimes use the picture at the bottom to describe that (NB if evidence changes the direction of the arrow may also change). Of course it is good to discuss if you agree to the estimated impact and the margins of uncertainty to that which Rethink priorities found, because I do think that is important to agree on (or disagree and find ways to close the gap)
E) too salesy --> I think you may underestimate how important the technology, service-model and business model are compared to anything else that is out there. There are dozens if not hundreds of changes we have done and are continuously doing to improve the intervention (at GOAL 3 level) and at the facility (to improve adoption and usage at facility level) all embedded in the business model which allows us to do continue to do this over time. We do not describe this here in this post to sell or brand our solution, but also to make clear that just putting a monitor down will not have the same effect.
Last but not least: it is also a product of passion and enthusiasm that comes from building 8+ years towards this solution. I just can't be fully objective in how I describe our solution. Firstly because it is the answer to the problems I experienced myself when working in the field, secondly I think there are a lot of lessons to learn from our success that can be used by others. I hope you can accept that (and maybe even appreciate it ;-) )
F) thanks again. Really appreciate the open discussion. Would be great to connect
Nice one, some quick replies and lets set up a call. So great you are so passionate about this, and I for one think that your org should be well funded and an RCT should be done.
A) Sorry you can't get the RCT funded. Have you tried the new DIV? You're right randomising at hospital level makes it really tough. I feel like you probably need national referral hospitals for the required power. From Claude if you had hospitals roughly with 1,000 admissions a month at 3% mortality, looking for a 20% reduction in mortality you'd need 4-5 massive (national referral) hospitals, or maybe 2 massive ones and 15-20 smaller? 72 feels like too much but I obviously haven't looked into the details, just asked Claude. As your current results seem to show 30ish percent mortality reduction this seems on the difficult but feesable end?
B) I think you are in a great position continue to grow steadily if not scale without better evidence and survive - your organisation is still so young there's plenty of room to grow and then scale. You are in a pretty good position for philanthropy to scale the setup costs (rather than government) if an RCT shows it is this amazing.
C) This one confuses me I don't really understand what you mean by "Real world evidence". Your RCT should ideally be basically just what you do anyway, so wouldn't that be real world evidence? The problem is that you're looking at these mortality drops as evidence enough, where I would see them more as higher level M&E of your program. Your own M&E can never touch an external study for importance. Orgs like ANSH used mortality drops as a marker for their progrram, but they already know the intervention works from RCTs. I just dont' think it can fly for before/after studies or your own M&E to be the evidence which carries you to scale.
D) I like your chart a lot and I agree with the concept completely! But I don't agree with the levels of uncertainty RP cites, although at least they are pretty big. Given the variation even in your own data, relying on before/after studies, the nature of the effect of these kind of interventions to reduce over time (I think a 7 year horizon is too long for this CEA), and the general trend of hospital mortality improving over time I would put your error bars much wider. Claude told me Kamuzu hospital mortality dropped fom 9% to 3% between 2012 and 2015 of its own accord, so these kind of drops seem not unusual for a bunch of reasons. Also Claude and RP cited a bunch of studies in Western pediatric wards which showed benefits in before/after studies, but then when high quality RCTs were pooled the benefit disappeared. I would have anchored on the best quality research we have on the topic much more heavily than RP did, even though its in a different context. I think if the evidence isn't clear in high income countries, it increases the burden for you somewhat.
https://doaj.org/article/2d05aad6031749878cb5b302f931473f
A) We want it to be representative so only selecting referral hospitals does not do the job as it will not inform a ministry of health what will happen in rural areas. The challenge is that hospitals vary quite a bit, so you have to match them first for certain characteristics and then randomly assign them. The larger the variation the more hospitals you will need, hence we came at 72 through our power calculations. The challenge is most funds only go up to 2 or 3 million USD
B) I think you are right from where we are today, but there have been many moments where it was likely we would not survive. We have our own development team, registrations to maintain all requiring significant investments. If it takes too long, likelihood of dying in the valley of death will increase. There are very few innovators who are actually successful in public health systems for this reason.
C) In any research setting there will be research staff and a lot more data collection, this by itself has a huge effect on how care is provided. Moreover not all sites will be suitable for it, creating a significant selection bias. Therefore the performance in (optimal) trial circumstances often has limited meaning for what it means in real life. Moreover variations across countries, will likely influence success of the intervention. So generalizability is also limited (Why would someone in West Africa, South-East asia or the middle east believe the trial from East Africa) Real world evidence solves a large part of this, whether if it is implementation research or derived from the real world evaluation. We look at each installation for usage, adoption and impact on health workers stress and burnout, where possible (if funding allows) we look at impact on mortality. While the quality of the evidence may be lower, it says a whole lot more about what is actually happening within the specific hospital/country and provides opportunities to intervene (either through us, or the hospital itself) Seeing this over multiple hospitals in one country will help a lot more with decision making then a (potentially outdated) trial from 5000km away (as a figure of speech).
Again, not opposing a trial here just wonder if it is the right instrument for the stage we are at. Just to provide a different perspective. A 5 million implementation program, would allow me to go to 200-300 hospitals in different countries and perform a thorough monitoring and evaluation program in parallel which can look at before-after differences, but also at differences between other hospitals. Those results would be available within 1 year after implementation, so likely 2 years after a grant was awarded.
D) Don't think Claude is accurate. Without a verifiable source on the mortality drop that answer is pretty useless, beyond that it is not wise to assume a conclusion that this is common based on one site. Data quality is notoriously bad in this setting and there is severe under reporting. We have solved this by diving into individual patient files hand by hand (not us, but our research partners)
The study you cite is hardly comparable in my view
- general wards versus ICU/HDU
- High resource settings only - high levels of staffing and no lack of monitoring equipment
- Only adults
- Vast majority is surgical patients
- total number of patients only 1284 over all studies combined
- relative rare occurance of the events making them statistically underpowered to proove significant impact
Moreover if you dive into their discussion you also find that over multiple systematic reviews all findings are consistently positive favoring monitoring, yet not significant. That is relevant to our discussion, because one of the key conclusions is that even at 1-2% mortality reduction IMPALA would still be cost-effective against any bar.