Option A
Current care
The care patients would otherwise receive (for example, medication with routine follow-up).
- Total cost
- €5,000
- Total health
- 1.0 QALY
Psychedelic therapy economics · Part three
This is a plain-English introduction to the terms that appear most often in cost-effectiveness discussions. We will use one simple psychedelic therapy example throughout, then look at what the early research can (and cannot) tell us.
I studied psychology a decade and a half ago. We learned how to ask whether a treatment works, why it might work, and how a person changes. We spent far less time on the question that appears one step later: if a treatment works, is it a good use of limited health-care resources?
At the time, the clinical question felt like the whole question. If people got better, that was the result. Yet a health system has to make another kind of judgement as well. It has a finite number of clinicians, rooms, appointments, and euros (or pounds or dollars), and many treatments competing for them.
That tension is especially visible in psychedelic therapy. A trial can show a meaningful improvement and still leave a health system with a difficult choice. The medicine may be expensive. Sessions can occupy a room and several clinicians for most of a day. Benefits may last for years (or fade within months). Every one of those possibilities changes the economic case.
This is why I wanted to understand the language of health economics, and why I am writing this guide. Cost-effectiveness analysis gives decision-makers a consistent way to put the extra health produced by a treatment next to the extra resources it uses. It does not decide whether a person's recovery has a price. It helps make visible the choices that are already being made whenever limited resources are spent in one place instead of another.
The first step is easy to miss because headlines often leave it out. “Is psychedelic therapy cost-effective?” is not yet a complete question. Compared with what: a daily medicine, a course of psychotherapy, usual care, or no treatment at all?
That alternative is called the comparator. It is the path patients are likely to follow if the new treatment is not introduced. Because current care differs between countries, services, and patient groups, the same therapy can produce a different economic result when the comparator changes.
For the next few sections, imagine a health system choosing between two options. Current care costs €5,000 and produces 1.0 QALY (I explain what a QALY is in the next section). A new psychedelic therapy pathway (including preparation, the treatment day, and follow-up) costs €15,000 and produces 1.5 QALYs. These numbers are invented and deliberately round. They are not an estimate for any real product; they are simply a way to keep the arithmetic visible while the ideas become familiar.
Option A
The care patients would otherwise receive (for example, medication with routine follow-up).
Option B
A psychedelic-assisted treatment with preparation, an administration day, and follow-up.
The new pathway therefore costs €10,000 more and produces 0.5 QALYs more health. Everything that follows is a more careful way of understanding those two differences.
Comparing the cost column is relatively straightforward. The health column is harder. A depression trial may report remission, while a cancer trial reports survival and a pain trial reports mobility. A health system still needs some way to compare the health gained across those very different conditions.
The usual common measure is the quality-adjusted life year, or QALY. One QALY is one year lived in full health. A year lived at half of full health counts as 0.5 QALYs; two years at that level add up to one QALY. The measure therefore combines two things that treatments may change: how well someone lives and how long that change lasts.
Picture health-related quality of life as a line moving through time. If treatment raises that line, keeps it raised for longer, or does both, the area beneath it grows. That extra area is the QALY gain shown in the figure.
Figure 1
When health improves and the improvement lasts, the shaded area grows.
The “quality” part is not simply assigned by a clinician. In England, NICE commonly prefers a short questionnaire called the EQ-5D. It asks people about five parts of daily life: mobility, self-care, usual activities, pain or discomfort, and anxiety or depression. The answers describe a health state, which is then converted into a number using preferences collected from the public.
That process gives researchers a consistent yardstick, but it is still a constructed measure rather than a direct reading from a machine. A value of 0.5 does not mean someone is “half a person.” It is a compact estimate of the health-related quality of life associated with that state, made for comparing outcomes across treatments.
Method note: NICE's current methods guide explains how QALYs and EQ-5D are used in technology appraisal. See the official economic evaluation guidance.
An important boundary
A common yardstick is useful precisely because it reduces many kinds of change to the same scale. That reduction also means some details disappear. Generic questionnaires may be less sensitive to changes that matter greatly within one condition or to one patient.
Psychedelic studies sometimes report changes in wellbeing, connectedness, meaning, relationships, or existential distress. Work, caregiver strain, and participation in family life may matter as well. Some of this will be reflected in a generic health score; some of it may only be partly visible.
The fair conclusion is neither that QALYs capture everything nor that they are useless. They answer a narrower question: how much comparable health was gained, and for how long? A careful assessment reads that answer beside the condition-specific outcomes and patient experience, not instead of them.
Now return to our two columns. The new pathway is both more expensive and more effective. The next question is not its total price, but how much extra cost accompanies each extra unit of health when we choose it over current care.
This is the incremental cost-effectiveness ratio, usually shortened to ICERA note about the ICER acronym in the United States. “Incremental” means extra compared with the alternative. We first subtract the current-care cost and health from the new pathway, then divide the extra cost by the extra QALYs.
Figure 2
Only the differences between the two pathways enter the ratio.
Current care
New psychedelic pathway
Extra cost
€15,000 − €5,000
€10,000
Extra health
1.5 − 1.0 QALYs
0.5 QALYs
ICER
€20,000
per additional QALY
In our example, the new pathway costs €10,000 more and produces 0.5 QALYs. Dividing €10,000 by 0.5 gives an ICER of €20,000 per QALY gained. In ordinary language: the model estimates that each additional QALY costs €20,000 when the new pathway replaces current care.
That number is not the price of a person's life, and it is not a permanent property of the treatment. Change the comparator, price, outcome, time horizon, or health system and the ICER changes too. It belongs to this comparison under these assumptions.
It is tempting to calculate a ratio every time, but first we need to know what kind of trade-off we have. A new treatment can improve or worsen health, and it can raise or lower costs. Put those two dimensions together and four broad possibilities appear.
Figure 3
Our worked example belongs in the better-health, higher-cost corner.
Extra cost
The new treatment is dominated by the alternative.
Is the extra health worth the extra cost?
Are the savings worth the loss in health?
The new treatment is dominant.
Extra health →
Two corners are relatively easy to interpret. If a treatment produces better health and costs less, economists call it dominant: there is no sacrifice to weigh. If it produces worse health and costs more, it is dominated by the alternative. A ratio is not helpful in either case because the direction of the decision is already clear.
The other two corners contain a genuine exchange. A cheaper treatment might produce slightly less health, forcing a decision about whether the saving justifies the loss. More commonly, a new treatment (including our fictional psychedelic pathway) produces better health at a higher cost. The ICER describes that exchange; it does not yet tell us whether the exchange is acceptable.
Our €20,000-per-QALY result only becomes useful when a decision- maker has something to compare it with. That benchmark is often called a cost-effectiveness threshold. Below it, the additional health is more likely to be considered a reasonable use of resources. Above it, the case becomes harder to make.
A threshold is not a market price for a QALY and certainly not a moral price for a life. It is a decision rule used inside a health system with a limited budget. Spending on one treatment means something else cannot be funded, so the benchmark is an attempt to represent the health that may be displaced elsewhere.
Nor does the threshold make the decision automatically. An appraisal committee reaches this point after considering the clinical evidence, the model, uncertainty, severity, unmet need, and other features of the case. A daily tablet and a one-off gene therapy could have the same modelled cost per QALY while creating very different pressures on this year's budget. The ratio informs the judgement; it does not replace it.
This also explains why a result cannot be carried unchanged from one country to another. England uses a relatively explicit range. The Netherlands relates reference values to disease burden and a broader societal assessment. The United States has no single national payer or threshold. The same evidence therefore enters three different decision processes.
England
NICE describes ICERs below £25,000 per QALY as generally acceptable. Between £25,000 and £35,000, committees weigh other factors more closely; above £35,000, a stronger case is needed.
Netherlands
Zorginstituut Nederland evaluates cost-effectiveness alongside effectiveness, necessity, and feasibility. Reference values vary with disease burden and sit within a broader societal assessment.
United States
There is no single national payer threshold. Studies often show results at values such as $100,000 or $150,000 per QALY, while individual payers make their own coverage and budget decisions.
Method note: For the formal decision context, see NICE's committee recommendations and the Dutch cost-effectiveness assessment framework.
Up to this point, the arithmetic has been tidy because our numbers were given to us. Real evaluations start with a messier problem: trials rarely observe every cost and outcome for every year that matters. A six- or twelve-month study cannot directly show what happens five years later.
An economic model bridges that gap. It combines measured evidence with assumptions about what happens after the trial: whether benefits endure, when people relapse, whether they receive another course, and which health-care costs follow. Modelling is not a trick or an automatic weakness. It is a transparent way to make a decision before every uncertainty has been resolved. Its result, however, can only be read alongside the choices used to build it.
The comparator determines the alternative path. The perspective determineswhose costs enter the calculation. A health-care perspective might count medicines, clinicians, and hospital visits; a societal perspective may also include travel, unpaid care, time away from work, or productivity. Neither is the uniquely “real” cost. Each answers a different decision-maker's question.
The time horizon determines how far the model looks. A therapy with a high upfront cost may look poor after six months and attractive after five years if benefits persist and later care is avoided. If the effect fades quickly, a long horizon can exaggerate the value. Durability, retreatment, staffing, price, and the future use of care therefore do much of the work beneath a single headline ICER.
This is where the psychedelic literature becomes especially interesting. We have trial protocols and short-term outcomes for several therapies. We have much less routine-care evidence about what happens after those protocols leave a research setting.
A small but growing group of psychedelic cost-effectiveness studies has investigated some of these questions. Researchers have modelled psilocybin-assisted therapy for depression and MDMA-assisted therapy for PTSD. These are useful first attempts to connect trial results with the decisions health systems may eventually face.
The studies do not begin from nothing. A protocol can specify how many preparation, administration, and follow-up sessions are delivered. Trials measure symptoms, response, remission, adverse events, and sometimes quality of life. Those observations form the closest thing the model has to solid ground.
The next layer is derived. Researchers may translate a symptom score into a health utility, attach costs from other datasets, or project relapse and remission beyond follow-up. Finally come the questions routine psychedelic care has not yet settled: the eventual medicine price, real-world staffing and throughput, durability, retreatment, and longer-term use of health services.
Figure 4
The useful distinction is whether an input was measured, derived from evidence, or remains unresolved for routine care.
Lane 1
Closest to what the trial or protocol directly reports
Lane 2
Built from measured data, other literature, and explicit assumptions
Lane 3
Inputs that routine psychedelic care has not yet settled
Items can move between lanes as better evidence arrives. A utility may be collected directly in one study, mapped from symptoms in another, and borrowed from older literature in a third.
These categories are not fixed forever. A durability estimate that is assumed today can become measured evidence after longer follow-up. A staffing pattern specified by a trial may turn out to be different in routine care. The point is not to dismiss modelled inputs, but to know which kind of knowledge each number represents.
Severe depression · United Kingdom · 6 months
The model found more QALYs than medication or CBT, but cost-effectiveness depended on a lower medicine price and less therapist support.
Still uncertain
Whether effects and downstream savings persist beyond the short model horizon.
Treatment-resistant depression · United States · 12 months
At an assumed $5,000 treatment cost, the model estimated 0.031 additional QALYs and an ICER of $117,517 per QALY.
Still uncertain
Durability, treatment price, and what happens when people need retreatment.
Chronic, severe PTSD · United States · 5 years
Using a $12,000 medicine price per session, the model estimated 0.377 additional QALYs and an ICER of $83,845 per QALY.
Still uncertain
Long-run health-care use, the eventual price, and how trial delivery translates into routine care.
Read together, the three studies suggest a cautious pattern. A favourable result is possible, but it often depends on the price of the course, the amount of therapist support, and benefits that continue beyond the observed trial. The precise ICER belongs to the calculation. It does not make those underlying uncertainties equally precise.
A good analysis therefore does not run the model once and stop. It changes important inputs to see whether the conclusion survives. If changing one assumption overturns the result, that assumption deserves more attention than the reassuring precision of the base-case number.
Economists call the simplest version a one-way sensitivity analysis: change one input at a time while leaving the others alone. A probabilistic analysis goes further by repeatedly drawing many inputs from plausible ranges. Instead of one verdict, it can show how often a treatment is cost-effective at different thresholds.
Figure 5
This schematic shows which assumptions are likely to move a psychedelic model. It is not a ranking of measured effect sizes.
1. Durability of benefit
Longer benefit can accumulate more QALYs and offsets.
2. Course price and staff time
High-touch delivery raises the upfront cost.
3. Clinical effect
Response and remission determine early health states.
4. Relapse and retreatment
A second course reopens both cost and outcome.
5. Comparator costs
The alternative pathway changes the difference.
For psychedelic therapy, we already know more about some delivery inputs than others. A COMP360-style protocol, for example, tells us about preparation, a long administration day, and follow-up. That gives us a basis for estimating staff and room time (which we mapped in our article on the cost of one therapy episode). It does not yet settle the routine price, how services will staff treatment at scale, or how often patients may return.
Durability is often the largest lever because it affects both sides of the model. Longer benefit accumulates more QALYs and may prevent future care; fading benefit reduces those gains and may introduce the cost of retreatment. The most informative result may therefore be the break point: how short would durability need to become, or how high would delivery cost need to rise, before the conclusion changes?
One of the field's longest follow-ups shows why long-term data are valuable (and why they need careful interpretation). In a psilocybin smoking-cessation pilot, 10 of 15 participants were abstinent at 12 months and 9 of 15 at a later follow-up averaging 30 months. That is an unusually durable signal, but it came from a tiny open-label study that combined psilocybin with structured cognitive behavioural therapy. It is not a stable long-term effect estimate that a model can simply copy.
The large COMP360 depression programme is beginning to fill in more of the middle of the timeline. The COMP005 and COMP006 Phase III trials measured their primary depression outcomes at six weeks, followed participants through a blinded 26-week period, and include open-label follow-up through week 52 with protocol-defined retreatment. Those observations are much more informative for a model than a single early endpoint. They still do not establish how long benefit will last in routine care (or how often retreatment will be needed outside a trial).
Ketamine provides a useful counterexample. In depression, its antidepressant effect can arrive quickly but is often transient, which is why researchers have investigated repeated or maintenance ketamine treatment. The available evidence is promising but methodologically limited. Economically, that is a different pathway: if repeated dosing is needed, the model must include repeated medicine, staff, and clinic time rather than treating the first response as a one-off durable gain.
Even a convincing ICER answers only one part of an access decision. It says that the extra health looks like a reasonable use of resources compared with a particular alternative. It does not show that a health system can afford the programme this year, that enough trained clinicians and suitable rooms exist, or that payment reaches providers in a workable form.
The distinction between cost-effectiveness and affordability is easy to miss. A therapy may represent good value for each patient yet create a large immediate budget impact if many eligible people seek it at once. A provider can still face poor service economics if reimbursement does not cover safe staffing and low throughput. A patient can still face travel, time away from work, referral barriers, or co-payments.
Nor does an average QALY result tell us who receives the benefits and who remains excluded. Equity, patient preferences, caregiver effects, and outcomes that generic quality-of-life measures miss belong in the wider decision. Cost-effectiveness is one disciplined view of the problem, not the whole view.
For psychedelic therapy, the broad research agenda is now clearer. Trials can collect utility and resource-use data. Longer follow-up can measure relapse and retreatment. Implementation studies can observe staffing, throughput, and real-world costs. Each result can move an important input from “assumed” towards “observed.”
The next two articles follow those assumptions further: first, why the duration of benefit changes the economics, and then what happens when treatment is needed again.
Start with the comparator. A QALY combines health-related quality of life with time. An ICER divides the extra cost of one option by the extra QALYs it produces. A threshold helps a decision-maker judge that exchange inside a particular health system. The model decides which costs, outcomes, assumptions, and years enter the calculation.
So when someone says a psychedelic therapy is cost-effective, hear a conditional sentence rather than a universal fact. It means the therapy appears to be a reasonable use of resources compared with a specific alternative, using a specific body of evidence and assumptions, at a specific decision threshold. That is less definitive than the headline (and far more useful).