I used to treat the yearly flu-shot percentage the way most people treat a weather forecast. A number lands in the news, someone repeats it at a family table, and the conversation moves on. Then I sat with one season long enough to notice the number was doing more work than the underlying comparison could support. The shot that year was described as a good match for the strain that dominated. The published estimates still looked tidy. The path that produced them did not.
That is the uncomfortable part. A figure can be calculated carefully and still answer a different question from the one readers think they are hearing. If you have ever nodded along to a headline about flu vaccine effectiveness and then wondered why the office still filled up in January, you are not being difficult. You are noticing a gap between a study design and the claim people carry home.
What Those Seasonal Percentages Actually Compare
Each spring, public-health summaries offer an estimate for the shot used in the season just ended. The figure is usually framed as protection against medically attended illness, sometimes split into urgent visits and hospital stays. It feels precise because it arrives with confidence intervals and age bands. Precision of arithmetic is not the same thing as a clean answer.
Most of these estimates do not come from a trial in which people were assigned, at random, to the shot or to a placebo and then followed through the same weeks. They come from a test-negative design. People who already sought care for an acute respiratory illness are tested. Those who test positive become the cases. Those who test negative become the comparison group. Vaccination status is then compared between the two.
The appeal is real. By keeping the sample inside people who showed up for care, analysts hope to blunt one familiar distortion: people who seek care readily are also more likely to get vaccinated. Restrict the sample to care-seekers, the thinking goes, and that habit is shared on both sides of the ledger. I have found that explanation persuasive on a first read and less persuasive on a second. The restriction cuts one bias and can invite another.
A sample limited to people who already walked through the door can quiet one bias and open a second door at the same time.
That second door has a technical name, colliding bias, and it is less widely discussed outside methods seminars. When you condition on a common effect of two causes, the association you observe between those causes can bend in odd directions. Healthcare-seeking is not a neutral filter. Illness severity, trust in clinics, work schedules, and insurance status all push people toward or away from the door. Once the analysis only sees people who crossed that threshold, vaccination and infection can look related for reasons that have little to do with immune protection.
Perhaps the most interesting aspect is how rarely the published summary admits that the net direction of these forces is unknown. Reduced confounding on one side, added distortion on the other. The remainder is not a rounding error. It is the estimate.
A Season That Should Have Been Favorable
Take a recent season in which the formulation lined up well with the strain that did most of the circulating. A large surveillance network, working with public-health partners, reported effectiveness against emergency and urgent-care encounters and against hospitalization. The headline range sat roughly in the thirties and forties for many adult groups. On paper, that is a modest success. In the footnotes and the timing charts, the story frays.
I am not interested in scoring a single paper as if it were a sporting event. I am interested in whether a reader can trust the percentage as a description of what the shot did in ordinary life. On that narrower question, the season is a useful stress test. Match was not the weak point. Design was.
Why the Background Risk of Infection Matters
Flu does not arrive as a flat line. Risk climbs, peaks, and falls. In the season under discussion, infection pressure was high from October through December and distinctly lower from January through March. Any fair comparison of vaccinated and unvaccinated people has to respect that curve. If the two groups do not spend similar shares of their time under the same pressure, the comparison is already tilted.
Vaccination is tied to the calendar. Clinics push the shot as the wave builds. Most people who are going to get it have done so by the end of December. After that, the share of encounters involving a vaccinated person levels off. That is not a scandal. It is how a seasonal campaign works. It is also a classic setup for confounding by time.
Look at the split in exposure, not just the split in status at the moment of a visit. In the high-risk months, vaccinated people accounted for a smaller share of the relevant encounters than unvaccinated people. In the low-risk months, the pattern flipped. Vaccinated people accumulated more of their observed time when background risk was soft. Unvaccinated people accumulated more of theirs when background risk was hard.
You can feel the bias without a formula. Imagine an extreme version, which I use only as a teaching sketch. Nobody is injected in the first period. At the start of the second period, everyone receives saline. If you then compare “vaccinated” infection rates with “unvaccinated” rates, the saline looks protective, because the labeled group lived mostly in the quiet months and the unlabeled group lived mostly in the loud ones. No biology required. Calendar did the work.
| Window | Background risk | Share of vaccinated exposure | Share of unvaccinated exposure |
| October to December | High | Lower, about 47 percent | Higher, about 61 percent |
| January to March | Low | Higher, about 53 percent | Lower, about 39 percent |
The same pattern has been walked through for other respiratory vaccines during periods of sharp waves. It is less often named in flu work, even though flu rollout routinely chases the rising wave and finishes near the winter peak. In my experience, readers assume the analysts “adjusted for calendar time” and move on. Adjustment can help. It does not automatically erase a structural mismatch in when people were at risk.
There is a cleaner way to ask the question in a cohort. Match an unvaccinated person to a vaccinated person on the vaccination date. Follow both forward. If the unvaccinated person later gets the shot, stop their observation at that point. Both people then share the same subsequent calendar. That design is fussier to run. It is also harder to fool with a wave that rises and falls while the campaign is still filling syringes.
The Quiet Exclusion Called Immortal Time
Timing bias does not stop at the wave. The network also set aside events that fell in a short window after vaccination. The rule, in plain language, excluded encounters in which documented vaccination had occurred fewer than 14 days before the index date. The index date was the earlier of the flu test or the urgent visit or admission.
The intention is understandable. Analysts often treat the first two weeks as a period before protection, if any, would be expected. Dropping those events is meant to avoid diluting an effect. The side effect is famous among people who study observational data, and under-appreciated everywhere else. It is called immortal time bias.
Here is the mechanism without the jargon pile. A person who was vaccinated is, by rule, not allowed to contribute an early bad outcome to the vaccinated column. Those early days are treated as if they did not count against the shot. The unvaccinated comparison has no such grace period. Early events in the vaccinated group vanish from the numerator, or get shunted into an exclusion pile, while time and risk on the other side remain visible. The association bends toward making vaccination look better.
How many early events disappeared? The public flowchart is not generous. Nearly 2,000 outpatient encounters were set aside because vaccination fell 1 to 13 days before the index date, or because vaccination status was unknown. No split between those reasons. “Probably many” is not a satisfying count. It is the count we are given.
I have watched the same exclusion move estimates in other vaccine studies. Putting the early events back in has, in at least one well-described reanalysis, cut apparent effectiveness sharply, sometimes by about half. Apply a similar humility here and a published range of 30 to 40 percent can slide toward 20 percent or lower. That is not a claim that the true number is zero. It is a claim that the published number absorbed a bias with a known direction.
- High-risk weeks were loaded with unvaccinated exposure.
- Low-risk weeks were loaded with vaccinated exposure.
- Events inside 14 days of a shot were removed from the vaccinated side.
- People vaccinated in late December often could not contribute events until January, a softer month, because of that same 14-day rule.
Stack those choices and ask a blunt question. How much of the association in that network is the shot, and how much is the calendar plus the exclusion window? I cannot hand you a single corrected percentage. I can say the share attributable to bias is large enough that “certainly a lot, possibly most” is a fair reading. Elsewhere, a cohort from the same season, re-read against both symptomatic and asymptomatic infection, has been argued down to no detectable benefit. Different data, same season, much less comfort.
When Hospital Estimates Come Out Weaker
Severity is supposed to be a friend of a real effect. If a product reduces infection, or reduces the chance that infection becomes severe, the benefit should look similar or stronger as you move from an urgent-care visit to a hospital bed. Each step is another conditional probability. If the relative risk at each step is at or below 1, the combined relative risk should not mysteriously inflate.
The season’s table did something else. Adjusted estimates against hospitalization were almost always smaller than adjusted estimates against outpatient encounters. Sometimes the gap was modest. Sometimes it was not. That pattern is unusual enough to pause on.
The write-up described effectiveness as similar within the same health systems across ambulatory and inpatient settings. Similar is a generous word for a column that sits consistently lower, and at times substantially lower. Later paragraphs acknowledge that the pattern is unexpected and offer explanations that, to my ear, do not carry the weight. Admission thresholds differ by hospital. Testing habits differ. None of that automatically produces a systematic drop in the adjusted figure.
A reader does not need to accuse anyone of bad faith to find this unsatisfying. Unexpected results happen. They should be labeled as unexpected, not folded into “similar” and waved through. When the more severe outcome shows the weaker association, the first instinct should be to audit bias, not to smooth the sentence.
The Healthy Vaccinee Problem That Models Cannot Finish
Another pattern sat in the same table. In most outpatient analyses, and in every hospitalization analysis, the estimate fell after adjustment. Unadjusted associations looked kinder to the shot. Adjusted ones looked smaller. That direction has a familiar name: the healthy vaccinee bias.
People who get the seasonal shot are, on average, not a random draw. They are more likely to have a regular clinician, to manage chronic illness, to show up for other prevention, and to have the schedule and transport to do it. Some of that advantage is measurable. A lot of it is not. Frailty that never gets coded, family support, diet, smoking history that is half-reported, the simple habit of calling early when a cough starts. Those differences protect against bad outcomes whether or not the syringe did anything.
Regression can nibble at the measured part. It cannot finish the job. The season’s models were unusually elaborate, which is both a credit and a warning. Age, study site, and calendar time were prespecified. Other covariates entered if a standardized difference crossed 0.20. Age and calendar time were fit as natural cubic splines. On top of that, inverse-propensity weights were built with generalized boosted regression trees and then used in logistic models, meant to soak up “additional imbalance.”
Complexity is not transparency. It is unclear how additional imbalance was detected, and which variables actually built the weights. Hand the dataset to a careful outsider and the description still would not let them reproduce the analysis. That is a practical problem, not a philosophical one. If a result depends on a machine-learning weight that nobody else can rebuild from the text, the result is a private number wearing public clothes.
Adjustment can shrink a bias. It rarely deletes the part of health that never made it into a code.
Observational methods, restated plainly
So the adjusted estimates are better than the raw ones, and still biased. How much residual confounding remains is the sort of question papers answer with a sensitivity paragraph and readers answer with a shrug. I would rather see the shrug in the abstract.
The Elderly Result That Does Not Sit Still
Age is where flu immunology usually gets humbler. Immune response softens with age. We expect smaller effectiveness in older adults, not larger. Against outpatient encounters, the season roughly obeyed that expectation: something like 41 percent in older adults versus 45 percent in younger adults. Close enough to call similar, and slightly lower where biology would predict a dip.
Against hospitalization, the pattern reversed. Older adults showed about 41 percent. Younger adults showed about 23 percent. A near doubling, in the outcome where frailty should have made protection harder, not easier. The abstract stated the age split. It did not dwell on how strange the hospitalization half was.
The offered explanation is product mix. Most younger adults received a standard inactivated shot. Most older adults received an enhanced product, high-dose or adjuvanted. That could, in principle, lift the older estimate. It does not explain why the lift appears for hospital stays and not for urgent visits. The authors more or less concede the point for outpatient care: despite enhanced products, effectiveness there looked similar to the standard-dose group. The remarkable benefit, if it is a benefit, is confined to the outcome that already behaved oddly.
Unexplained is not the same as false. It is the same as not yet earned. A trustworthy result should survive the product-mix story without needing a special rule for one column. This one does not.
What the age split roughly showed: Outpatient: older adults slightly lower than younger adults Hospitalization: older adults markedly higher than younger adults Product mix offered as the reason Product mix does not explain the split across outcomes
What Happened After People Were Already in a Bed
Buried past the main effectiveness tables is a secondary look that deserves more air. Among people already hospitalized with confirmed flu, the analysts compared severe in-hospital outcomes by vaccination status, split by age. This is closer to a nested cohort than to the test-negative setup. It asks a different question: once flu has put you in the hospital, did prior vaccination change the course?
They did not adjust this comparison for baseline traits. That limit is stated, and it matters. I will stay with the elderly slice, which held nearly 70 percent of the deaths in view.
Within hospitalized older patients, the vaccinated group was a bit older and a bit sicker. The differences were generally small. Fatality was not. About 4.0 percent in the vaccinated group, about 2.7 percent in the unvaccinated group. The risk ratio is 4.0 divided by 2.7, or roughly 1.5. Fifty percent higher, not lower.
Play the counterfactual honestly. Suppose the shot truly reduced death among elderly people hospitalized with flu, with a risk ratio somewhere between 0.5 and 0.75, meaning 50 to 25 percent effectiveness against death in that already-severe group. To land on an observed ratio of 1.5, confounding would have to shove a true 0.75 up to 1.5, a twofold bend, or a true 0.5 up to 1.5, a threefold bend. That is heavy confounding for differences the paper itself calls generally small. Possible? In frail populations, yes. Comforting? No.
The write-up treats baseline demographics and medical conditions as similar within the age band, then treats severe outcomes, including intensive care, ventilation, and death, as similar across vaccination groups. Similar is doing a lot of lifting for 4.0 versus 2.7. They do not claim a hidden benefit. If a reader accepts their premise of no meaningful confounding, the remaining question is awkward: did vaccination, in this already hospitalized group, line up with higher fatality? The data cannot settle that. They also cannot be summarized as a quiet win.
How a Reader Should Hold the Number
None of this requires a cartoon villain. Surveillance networks do hard logistical work. Testing a fraction of respiratory visits, linking vaccine registries, and publishing before the next season is a real public service. The trouble starts when a biased association is translated into a percentage of protection and then into advice that sounds like a settled effect.
I hold the seasonal figure the way I hold a restaurant rating based only on people who already walked in hungry. Useful as a description of that doorway. Weak as a description of the whole street. Test-negative studies describe an association among care-seekers under a set of exclusions. They do not, by themselves, deliver the effect a healthy adult should expect from a shot taken in October.
- Ask whether vaccinated and unvaccinated exposure sat on the same part of the wave.
- Ask what happened to events in the first two weeks after the shot.
- Ask whether adjustment moved the estimate down, which often signals healthy-user confounding that was only partly removed.
- Ask whether the severe outcome looks stronger than the mild one. If it looks weaker, slow down.
- Ask whether older adults look strangely better only in one column.
Fail two or three of those checks and the headline percentage is a conversation starter, not a decision tool. The season in question fails more than two.
Designs That Would Argue More Cleanly
Randomized trials of the annual reformulated shot, large enough and run through a full wave, are the standard people invoke and then set aside. Strain match changes every year. Ethics boards and manufacturers have little appetite for placebo arms once a product is already recommended. So the trial that would settle a given season is the trial that will not be run. That absence is not a reason to lower the bar for everything else. It is a reason to be choosy about substitutes.
Two substitutes keep coming up in careful re-reads, and neither has so far delivered a glowing endorsement of the annual shot.
The first is a cohort with two outcomes: symptomatic infection and asymptomatic infection. If the shot mainly reduces illness rather than infection, the symptomatic outcome should move more than the silent one. If both sit near zero once timing is handled, the protection story thins out. A reanalysis along those lines for the well-matched season has been reported as flat. Flat is informative. It is also inconvenient for a campaign that prefers a positive integer.
The second is a regression discontinuity design. The idea is to use a rule that sharply changes the chance of vaccination, age cutoffs for enhanced products being the obvious candidate, and to compare people just on either side of the cutoff. They are similar in health and different in product access. If the shot has a sturdy effect, the outcome should jump at the line. So far, that style of check has not produced a promising signal for the annual flu shot either.
I like both ideas because they make fewer bets on statistical adjustment. They do not require a boosted tree to decide who was “similar.” They ask the data to show a break or a paired difference that a reader can see. Until those designs, or a genuine trial, show a clear benefit, the test-negative percentage should stay in the modest-claim drawer.
What This Does and Does Not Say About Personal Choice
A methodological critique is not a personal instruction. Some people have occupational exposure, a fragile household, or a history of bad flu seasons and will still want the shot. Others will look at a biased 30 percent and decide the errand is not worth it. Both can be rational. What is harder to defend is treating the published estimate as if the biases were decorative.
There is also a communication cost. When a season is later revised, or when a well-matched year still leaves hospitals busy, trust leaks. People do not parse splines. They remember being told a number. If the number was partly calendar and partly exclusion rules, the next number inherits the doubt. I would rather public summaries lead with the design limits than with a single percentage in bold type.
Risk is not only infection. It is also the habit of outsourcing judgment to a figure whose construction you have never seen. That habit shows up in markets, in medicine, and in ordinary planning. The fix is not cynicism. The fix is a short list of questions, asked before the percentage is allowed to travel.
A Closer Pass Through the Bias Stack
Let me slow down on the stack, because this is where confident summaries usually skip. Picture two coworkers, same city, same circulating strain. One gets the shot in mid-October. The other never does. Through November and December the wave is high. Both are at risk, and only one is labeled vaccinated. In late December a third coworker gets the shot. Under the 14-day rule, that third person cannot count as a vaccinated case until January, when the wave has eased. Their early January cough, if it arrives inside the window, is excluded. Their later January exposure sits in a quieter period.
Now aggregate thousands of those stories. The vaccinated column is enriched for people whose observed risk-time falls after the peak. The unvaccinated column is enriched for people whose observed risk-time includes the peak. Exclude the awkward early events. Fit a model with splines for calendar time. The spline is a curve drawn through the season. It can absorb average trends. It cannot give the late-vaccinated person the high-risk weeks they were barred from contributing. The coefficient that falls out is then converted to a percentage and called effectiveness.
That conversion is worth a pause. Vaccine effectiveness is usually one minus the odds ratio, expressed as a percent, inside the test-negative sample. It is not a risk reduction in the whole population over a defined follow-up. Readers hear “40 percent effective” and imagine 40 fewer illnesses per 100 people who would otherwise have gotten sick. The study delivered an odds comparison among people who sought care and were tested, after exclusions. The leap from one to the other is where public language gets ahead of the method.
Rough translation people hear: 40 percent fewer illnesses.
What was often computed: 1 minus an odds ratio in a tested, care-seeking sample, after time-based exclusions.
Odds ratios and risk ratios diverge when outcomes are common. Medically attended respiratory illness in peak months is not rare. So even a perfectly unbiased odds ratio would overstate the risk reduction people think they are hearing. Layer bias on top and the public number drifts further from the lived one.
Colliding Bias, in Ordinary Language
I promised the two-edged sword would get a plain telling. Suppose two things both push a person to seek care: feeling quite ill, and being the sort of person who uses clinics. Vaccination is more common in the second group. Infection, if it causes worse symptoms, is more common in the first. The study keeps only people who sought care. Inside that selected room, the two causes can trade off. A vaccinated person may have needed less severe illness to cross the door, because they were already a care-seeker. An unvaccinated person may have needed to feel worse before they came in. The test results inside the room then mix behavior and biology.
Does that make the shot look better or worse? It depends on which force dominates, and on how testing is triggered. That dependence is the point. The design’s main selling point, restriction to care-seekers, is also the design’s open flank. You cannot claim the restriction saved you from confounding and ignore the selection it created. Net bias unknown is the honest summary. It should travel with the percentage.
Why a Good Strain Match Does Not Rescue the Estimate
Match is the variable campaigns like to emphasize, and it matters for mechanism. A shot built against a distant strain has less chance of helping. A shot built against the strain that actually circulated has a fairer test. The season here was that fairer test. If estimates in a well-matched year are still swollen by timing and exclusions, mismatch cannot be the alibi.
This is why I keep returning to one paper rather than to a generic complaint. A generic complaint is easy to dismiss as mood. A single season with a favorable match, a large network, a public table, and a stack of internal inconsistencies is harder to wave off. Lower adjusted than unadjusted estimates. Weaker hospital figures than outpatient figures. A reversed age pattern on the severe outcome. Higher crude fatality among vaccinated elderly patients already in hospital, dismissed as similar. Hidden counts inside the 14-day exclusion. Exposure tilted toward low-risk months for the vaccinated group.
Any one of those might be noise. Together they describe a result that cannot be trusted as a measure of protection. They also describe problems that show up in other network studies of the same product, because the template is shared: test-negative sampling, a post-vaccination buffer, calendar adjustment that does not equal calendar matching, and a healthy-user gap that covariates only partly close.
Reading Tables Without Being Captured by Them
A practical habit has saved me more than once. Read the unadjusted column before the adjusted one. If adjustment shrinks the benefit, ask what the model thinks it removed, and whether that thing is health. Then read the severe column beside the mild one. If severity does not strengthen the association, do not let the prose call them similar without a fight. Then find the exclusion count. If early events are pooled with unknown status, assume the bias-friendly portion is not small until shown otherwise.
Finally, find the epidemic curve and the vaccination curve on the same mental page. If one rises while the other is still being filled, the analysis needed date-matching, not only a spline. Splines are flexible. They are not magic. A flexible curve fit to the average week does not reconstruct the counterfactual week a person would have lived if the campaign had started earlier or later.
I have found that editors like a single number because a single number fits a headline. Scientists sometimes like it because a single number looks like a conclusion. Readers pay for both preferences. The cost is a public that cannot tell a matched cohort from a selected sample, and a debate that swings between cheerleading and blanket rejection. Neither swing is analysis.
What Would Change My Mind
Skepticism should be revisable. A few results would move me. A date-matched cohort from a well-matched season, with the 14-day window reported rather than dropped, showing a stable reduction in both symptomatic and asymptomatic infection. A discontinuity at an age threshold for enhanced products, with a clear jump in outcomes and no jump in unrelated diagnoses. A trial, even a modest one, that follows assigned groups across the same weeks and does not recruit only after the peak. Any of those would deserve more trust than a test-negative percentage with the biases left inside.
Until then, the responsible sentence is narrower than the one that usually ships. In care-seeking samples, after exclusions and adjustment, an association remains in some seasons. The association is not a clean measure of effectiveness. Part of it, and in some seasons possibly all of it, can be produced by when people were vaccinated, which early events were removed, and who tends to get vaccinated in the first place. That sentence will not fit on a poster. It will fit in a mind that wants the poster to be true.
The yearly percentage can stay in the conversation. It should not be allowed to end it. A well-matched season was the easiest year for the estimate to look solid. It still needed the calendar, the buffer, and a very flexible model to look the way it did. I keep that in view the next time a fresh number arrives, tidy, confident, and thinner than it appears.