On July 1 of this year, Medicare started paying for weight-loss drugs. The GLP-1 Bridge is a temporary demonstration running through December 2027, under which eligible beneficiaries pay fifty dollars a month and manufacturers accept a negotiated net price of two hundred forty-five. Medicare eats the remaining one hundred ninety-five.
Eight weeks later, Eli Lilly published a study concluding that one hundred ninety-five dollars is a bargain.
The paper appeared in Diabetes, Obesity and Metabolism on August 24. Three of its six authors are Lilly employees and shareholders. The other three work for contract firms Lilly paid to run the analysis. The medical writing was done by an agency Lilly funded. The study was funded by Lilly. Its concluding paragraph names the Bridge program and its exact monthly price.
None of that makes the finding wrong. Industry-funded research is not automatically bad research, and dismissing a paper by its funding line is the lazy version of criticism. The interesting question is narrower. What does this study actually measure, and is it the thing Medicare needs to know?
The methodology is not sloppy. Propensity matching produced good balance on most covariates. The authors ran a pre-index parallel-trends check, which is the right thing to do and is frequently skipped. They used two different censoring approaches and reported both. They flagged the informative-censoring problem themselves and switched their primary analysis in response to it. The limitations section is candid about the follow-up being thin.
The headline numbers are real arithmetic. In the primary analysis, monthly healthcare costs excluding the drug were $145 lower in the treated group at six to twelve months and $319 lower at twelve to eighteen months. Inpatient and emergency visits fell, with an incidence rate ratio of 0.69 in the final window. Outpatient visits went up slightly, which is what you would expect from people who are now seeing a doctor about a prescription.
Notably, the treated group's costs barely moved. It was the control group that got more expensive, rising from $1,031 to $1,244 per month. The claim is not that tirzepatide made people cheaper. It is that untreated people got costlier and treated people held flat. Everything therefore depends on whether the control group's rise is a real trend or an artifact of who was left to measure.
Both cohorts started at 15,843. Here is what was left at each window.
| Window | Tirzepatide | Control | Cost difference |
|---|---|---|---|
| Baseline | 15,843 | 15,843 | — |
| 3 to 6 months | 8,645 | 11,101 | not significant |
| 6 to 12 months | 4,511 | 7,161 | −$145 |
| 12 to 18 months | 1,181 | 3,004 | −$319 |
Under the alternative censoring scheme, the treated arm falls to 793 and the cost difference widens to $607. The effect size and the attrition rate move together, in the same direction, at every step. By the window that carries the policy conclusion, roughly one treated patient in twenty is still being observed.
The first window is worth noting on its own. At three to six months, with the largest surviving sample, there is no effect in either analysis. The primary gives −$37 (p=0.484); the secondary gives positive $7 (p=0.889). The signal only appears after the cohort has already halved.
Why it mattersThe people still on tirzepatide at eighteen months are not a random sample of the people who started it. They tolerated the drug, could afford the refills, navigated prior authorization, and kept showing up. The healthy adherer effect is one of the better-replicated findings in pharmacoepidemiology. A 2006 BMJ meta-analysis of 46,847 patients found that good adherence to placebo was associated with an odds ratio of 0.56 for mortality. Adherence is a marker for the kind of person you are, not only for the drug you take.
The paper is honest that its estimate applies only to persisters. Then it hands the number to Medicare anyway.
Follow-up was censored at death, disenrollment, pregnancy, switching drugs, or discontinuation. Discontinuation is defined as a gap of more than 45 days after the supply runs out. That criterion can only fire in the treated arm. Controls are not taking anything, so they cannot stop.
Why it mattersThis is the structural asymmetry underneath everything else. The treated arm is continuously filtered toward persistence. The control arm is not filtered at all on that dimension, which is why 19 percent of controls survive to the final window against 7.5 percent of the treated. The authors saw this coming and built two corrections for it. Inverse-probability weighting reweights the survivors to stand in for the departed. Pairwise censoring throws away a control whenever their partner leaves.
Both corrections assume censoring is explainable by measured variables. If the reason someone quits tirzepatide is nausea, cost, or a life falling apart, and the claims data cannot see it, no weighting scheme recovers it. The corrections disagree with each other by a factor of two at the final window. That spread is itself the finding: the estimate is not stable under reasonable analytic choices.
Age, sex, race, census region, payer type, index month, Charlson score, comorbidity counts, prior utilization, prior cost, and BMI. That is a serious list, and it addresses the crude version of the affordability objection. These are not rich people compared to poor people. Baseline spending was balanced.
What was not matchedIncome, education, health literacy, food environment, and lifestyle behaviors. The authors say so directly in the limitations. Lab values including HbA1c, lipids, blood pressure, and weight were left out of the propensity model because too many were missing, which means the two arms were matched on diagnosis codes rather than on physiology.
Six covariates still exceeded the balance threshold after matching. The treated group had more Medicare coverage (38.5 versus 31.9 percent), more sleep apnea, more comorbidities, fewer people in the lowest BMI band, and was whiter. Those residual gaps were adjusted for in one analysis and folded into weights in the other.
The deeper problem is selection into treatment itself. To appear in the treated arm you had to obtain a paid pharmacy claim for branded Zepbound during a period when Medicare did not cover it and shortages were common. That is a filter for insurance generosity and administrative persistence, and no covariate in this model captures it. Anyone who bought vials directly from Lilly is invisible in the data and may be sitting in the control arm.
Internal age balance is excellent. Mean 64.5 years treated against 64.7 control, standard deviations of 6.8 and 7.0. Nothing to complain about in the numbers themselves.
How that balance was produced is more interesting, and it is visible only in the attrition figure. Matching was not performed on this cohort. It was performed on the full adult population, 171,925 treated against an eligible control pool of 5,009,184, a ratio of roughly twenty-nine to one. That produced 171,299 matched pairs. The over-55 cohort was then carved out by keeping only pairs in which both members were over 55, which left 15,843 pairs, about nine percent of the matched set.
Age was a covariate in the propensity model, so this is a defensible way to get a balanced subgroup, and the paper calls it pre-specified. But the calipers were tuned to the whole adult distribution, not to older adults, and the both-members rule discards precisely the pairs that straddle the age boundary. The balance is a byproduct rather than a design target.
The external questionThe cohort is defined as over 55 with a mean of 64.5, so roughly half of it is younger than 65. The Bridge program serves beneficiaries 65 and older plus certain disabled enrollees. Only 38.5 percent of the treated arm had Medicare at all. A majority, 55.4 percent, were commercially insured.
The eligibility criteria diverge further. Bridge excludes people with type 2 diabetes, moderate-to-severe sleep apnea, and MASH, on the logic that those conditions already have covered indications. This study also excluded diabetes, which lines up. But 31.5 percent of the treated arm had sleep apnea, a group Bridge largely routes elsewhere. Bridge also covers only the tirzepatide KwikPen and applies tiered BMI and comorbidity requirements.
So the population that generated the $319 is younger than the Bridge population, mostly not on Medicare, and includes a substantial slice that Bridge would not enroll. The paper's own limitation section concedes the findings cannot be extrapolated beyond this age-selected subgroup, then extrapolates them to a program with different age criteria.
Total cost is medical cost plus pharmacy cost. Split the final window into its two parts and both fall apart.
| 12 to 18 months | Medical | Pharmacy | Total |
|---|---|---|---|
| Primary (IPCW) | −$211 p=0.065 |
−$108 p=0.050 |
−$319 p=0.015 |
| Secondary (pairwise) | −$280 p=0.088 |
−$278 p=0.053 |
−$607 p=0.023 |
In the primary analysis the two components sum to exactly $319. Medical does not reach significance. Pharmacy lands precisely on the threshold. Their sum is comfortably significant. In the secondary analysis neither component clears p=0.05 and the total does.
The pharmacy problemNearly half the $607 is pharmacy spending, and pharmacy spending here excludes tirzepatide by design. So the claim embedded in that figure is that starting tirzepatide made the untreated comparison group's other drug costs explode.
In the secondary analysis the control arm's pharmacy cost runs $191, $204, $227, then $477. It more than doubles in the final window, at n=793. The treated arm sits flat at $221. Under the primary analysis the same control-arm series is $207, $215, $235, $297, a rise of 26 percent rather than 110 percent.
Two censoring schemes applied to the same underlying people produce control-arm pharmacy trajectories that differ by a factor of four in their final step. That is the signature of a handful of expensive specialty claims landing in a small denominator, not of a treatment effect.
One more thing the discussion never mentions. At three to six months the treated arm's pharmacy costs were higher by $34 per month, and that one was statistically significant at p=0.046.
Lilly's explanation for the savings is fewer hospitalizations. The primary analysis supports it: inpatient and emergency visit rates give incidence rate ratios of 0.86, 0.76, and 0.69 across the three windows, all significant.
What the secondary analysis showsSame outcome, same data, other censoring scheme. IRRs of 0.87 (p=0.101), 0.83 (p=0.065), and then 0.99 with a confidence interval of 0.60 to 1.64 and p=0.980.
At twelve to eighteen months the two arms are indistinguishable. The control arm's rate falls from 34 to 26 per thousand person-months and meets the treated arm exactly. The difference is two visits per thousand person-months.
Now hold that against the cost result from the same analysis and the same window. The saving is $607, the largest figure in the paper and the one Lilly's announcement leads with. So in the secondary analysis the biggest cost saving coincides with a complete absence of the effect that is supposed to be causing it.
Those two results cannot both be describing tirzepatide keeping older adults out of the hospital. One of them is noise, and the small denominator makes it hard to say which.
This is where the paper gets uncomfortable, because it cites its own refutation.
Reference 29 is a National Bureau of Economic Research working paper by Coady Wing, Sih-Ting Cai, Daniel Sacks, and Kosali Simon, titled Do GLP-1 Medications Pay for Themselves? It analyzes roughly 537,000 GLP-1 initiators using a stacked difference-in-differences design against not-yet-treated controls. Same data type, same broad method, thirty-four times the sample. It finds no reduction in non-GLP-1 spending. Spending was higher at one year and higher at five years. The result held in the subgroup without diabetes, which is this study's exact population.
A 2026 review in the Journal of Managed Care & Specialty Pharmacy summarizing the ICER assessment reached the same place from a different direction. Cost-offset signals show up mainly in patients who have both obesity and diabetes. Obesity-only populations tend to show spending increases. Managed care organizations, it advised, should expect value but not near-term savings.
The Lilly paper disposes of this literature in one sentence. Those studies, it says, looked at older drugs, and they "may have included individuals who were not persistent on the medication."
Read that clause again. The defense against the null result is that the null studies counted the people who stopped taking the drug. Which is true. It is also the entire mechanism by which this study got a different answer.
Lilly announced the study on August 26. The release says costs were up to 15 percent lower "at six months," and that the gap widened "by 12 months" to as much as $607. Then it says savings nearly covered the $195 Bridge price "beginning at six months" and exceeded it "starting at 12 months."
The paper's own conclusion says the drug cost may be partially offset "by 12 months" and may produce savings "by 18 months."
Both are describing identical estimates. The paper labels each window by when it ends. The release labels each window by when it begins. Every statement in the release is defensible on its own. The combined impression is that break-even arrives half a year earlier than the paper says it does. This is paltering in its purest form, and it is done with a labeling convention rather than a false number.
The release also leads its bullets with the secondary analysis, which produces the larger figures, while correctly identifying the primary analysis elsewhere in the text. And it states the study "included 15,843 adults" without mentioning that 1,181 remained when the headline number was computed.
Then there is footnote 4. The release claims lower hospital and emergency visit rates "across every follow-up period," with a footnote conceding the results were "numerically lower, but not statistically significant in the secondary analysis." An incidence rate ratio of 0.99 with p=0.980 is not a near miss on significance. It is a flat null, and it sits in the same window as the $607 the release put in its headline bullet.
The Bridge demonstration runs from July 1, 2026 to December 31, 2027. Eighteen months.
This study finds that cost offsets cross break-even somewhere between month twelve and month eighteen, among patients who stay on treatment. Real-world one-year persistence for weight-loss GLP-1s in commercially insured adults without diabetes ran 33 percent for 2021 initiators and 61 percent for the first half of 2024, according to a Prime Therapeutics analysis the Lilly paper cites as reference 31.
So the program is scheduled to expire at roughly the moment its own sponsor's evidence says the savings would begin, for the minority of enrollees who make it that far. The successor BALANCE model has been indefinitely delayed. Permanent coverage requires Congress.
This study answers a real question well: among older adults who stay on tirzepatide for eighteen months, does other spending fall? Probably yes, by a couple hundred dollars a month, with wide uncertainty.
Medicare is asking a different question. Among everyone we enroll, including the forty to sixty percent who quit inside a year, what happens to total spending? This design cannot answer that, because it removes those people from the denominator by construction.
The number is not fabricated. It is precisely measured on a population that shrinks by ninety-five percent before the measurement is taken, and it is being offered to a program whose enrollees are older, differently insured, and not yet selected for persistence.
Open the figures and it gets thinner still. Neither cost component reaches significance on its own. Half the largest estimate is a doubling of pharmacy spending in the control arm. And the hospitalizations that are supposed to explain the savings show no difference at all in the window where the savings are biggest.
Sources
Primary paper: Upadhyay N, Bonakdar A, Subedi K, Banerjee S, Behrend B, Hankosky ER. "Trends in Cost of Care With Tirzepatide in Adults Aged Over 55 Years With Obesity or Overweight Without Diabetes: A Matched Cohort Analysis." Diabetes, Obesity and Metabolism, 24 August 2026. Open access, CC BY-NC-ND 4.0. doi.org/10.1111/dom.71250
Company announcement: Eli Lilly and Company, "Zepbound linked to lower healthcare costs in adults over age 55 with obesity according to a real-world study," 26 August 2026. investor.lilly.com
Independent counterweight: Wing C, Cai ST, Sacks DW, Simon KI. "Do GLP-1 Medications Pay for Themselves?" NBER Working Paper 34678. nber.org/papers/w34678
NBER summary: "The Health and Healthcare Spending Effects of GLP-1s," NBER Digest, April 2026. nber.org/digest
Healthy adherer effect: Simpson SH, Eurich DT, Majumdar SR, et al. "A meta-analysis of the association between adherence to drug therapy and mortality." BMJ 2006;333:15. Retrieved via PubMed. doi.org/10.1136/bmj.38875.675486.55
Persistence data: Marshall LZ, Gleason PP, Friedlander N, Farley J, Urick BY. "Trends in 1-year persistence and adherence among initiators of high-potency, weight loss-indicated GLP-1 receptor agonists." JMCP 2026;32(3):281-291. doi.org/10.18553/jmcp.2026.32.3.281
ICER context: "ICER report demonstrates both the value and challenges in financing of weight loss medications." JMCP 2026;32(6):753. doi.org/10.18553/jmcp.2026.32.6.753
Associated commentary: Gasoyan H, Rothberg MB. "CMS BALANCE Model for Obesity: Implications for Patients and Clinicians." JAMA 2026;335(16):1383-1384. Cited by the paper as reference 17. doi.org/10.1001/jama.2026.2894
Bridge program mechanics: KFF, "What to Know About the BALANCE Model for GLP-1s in Medicare and Medicaid and the Medicare GLP-1 Bridge." kff.org
Bridge implementation: AJMC, "What You Need To Know Before the Medicare GLP-1 Bridge Goes Live," 20 July 2026. ajmc.com