November 7, 2025
9 min
Kenneth D
July 20, 2026
9 min

You've seen it on the label: "clinically proven," "backed by science," "shown in studies to improve..." Those words are doing real work on you. And once you understand how statistical significance actually functions — and how easy it is to manufacture the appearance of it — you'll read every supplement claim differently.
What's actually true: "Statistically significant" means a result is unlikely to be due to random chance. It does not mean the effect is large enough to matter in real life, that the study was well-designed, or that the product will work for you. The gap between statistical significance and clinical significance is wide, routinely exploited, and almost never explained on a label.
What's misleading or unregulated: Under the Dietary Supplement Health and Education Act (DSHEA) of 1994, supplements require no pre-market clinical proof of efficacy. Manufacturers can cite selectively picked subgroup results, underpowered studies, or relative risk figures that sound dramatic but reflect tiny absolute changes — and none of this triggers regulatory action unless the FTC or FDA pursues enforcement after the fact.
⚕️ LyfeiQ Score: Worth knowing — Most supplement buyers don't know the difference between a real result and a statistically dressed-up one. Learning this distinction is one of the highest-value health literacy moves you can make, and it costs nothing.
Statistical significance is a mathematical statement about probability, not a verdict on whether something works.
When a researcher runs a clinical trial, they're testing a hypothesis: does this supplement do something, compared to a placebo? At the end, they calculate a p-value — a number that tells them how likely it is that the result they observed would have occurred by chance alone, assuming the product had no effect.
The conventional threshold is p < 0.05. Cross that line and the result is labeled "statistically significant" — meaning there's less than a 5% probability the finding was a fluke. That sounds rigorous. And it is, as far as it goes.
But that threshold tells you nothing about three things that actually matter to a consumer.
1. Effect size. How big is the difference? A supplement that reduces joint pain scores from 62 to 61 on a 100-point scale might reach statistical significance in a large enough trial. That 1-point difference is real. It's also meaningless.
2. Absolute vs. relative risk. A result reported as "50% more effective" may mean the response rate went from 2% to 3%. The relative gain is 50%. The absolute gain is 1 percentage point. Supplement marketing almost always leads with the relative figure — because it sounds more impressive. The American Academy of Family Physicians has documented that even primary care physicians can be misled by relative risk framing when absolute risk isn't provided alongside it.
3. Whether the study was adequately powered. A small trial might miss a real effect — too few participants to detect it reliably — or, more dangerously for consumers, report a significant-looking result in one subgroup by chance. Run 10 subgroup analyses and probability alone predicts that roughly one will cross the p < 0.05 threshold even if the product does nothing.
This third problem has a name: p-hacking. It refers to the practice — sometimes deliberate, often unconscious — of running multiple analyses, selecting favorable subgroups, or stopping data collection at a convenient moment, then reporting only the result that "worked." As described in a widely cited 2011 paper by Simmons et al. in the journal Psychological Science, undisclosed flexibility in data collection and analysis allows presenting almost anything as significant. The problem is pervasive in nutrition and supplement research, and it rarely makes the label.
Before you buy based on a study citation, five questions can tell you whether that "clinically proven" claim actually holds up.
Ask: what was the effect size, not just the direction? A study can show a supplement "significantly" reduces something without showing it reduces it meaningfully. Look for the actual numbers — mean difference, change in score, percentage points — not just the statistical notation.
Ask: is this a relative or absolute risk figure? "Reduces risk by 40%" is a relative figure. "Reduces risk from 5% to 3%" is absolute. The first sounds impressive. The second shows you're looking at a 2-percentage-point real-world change. If the marketing doesn't show the baseline, that's information.
Ask: how many people were in the study? A trial of 20 people can generate statistical significance if the effect looks large enough — but a 20-person study has very low power and very wide error bars. Any single study with under 100 participants should be treated as preliminary, not proof.
Ask: was this the primary outcome, or a subgroup? If a study's main result was null but a subgroup — women over 50, people with a specific deficiency, participants in one geographic region — showed a significant finding, and that subgroup result is what the marketing cites, you're looking at a post-hoc subgroup analysis. Not the same as a preregistered, powered trial.
Ask: has this been replicated? One statistically significant study is a starting point. The scientific community begins to trust an effect when multiple independent teams find it. A supplement citing "a clinical study" — singular — is citing a starting point.
The FDA does not approve dietary supplements for efficacy before they reach shelves.
Under the Dietary Supplement Health and Education Act (DSHEA) of 1994, the burden of proof is essentially reversed compared to pharmaceuticals. With a drug, a company must demonstrate safety and efficacy before FDA approval. With a supplement, the FDA must prove the product is unsafe or mislabeled to take action — after it's already on shelves and in consumers' hands. As the FDA states in its own guidance, manufacturers generally do not have to provide the FDA with the evidence they rely on to substantiate safety before or after marketing their products.
The FTC governs advertising claims. In December 2022, the FTC published a 40-page Health Products Compliance Guidance document — incorporating lessons from over 200 enforcement cases — clarifying that claiming a product is "clinically proven" without adequate substantiation is an unlawful act. In April 2023, the FTC sent notices to approximately 670 companies marketing supplements, OTC drugs, and functional foods, warning them they could face civil penalties up to $50,120 per violation for unsubstantiated efficacy claims.
The structural problem the FTC guidance makes explicit: a study showing a "modest but statistically significant" result in a short-duration trial may not substantiate a broad effectiveness claim — especially if longer or larger trials show no effect. Cherry-picking the favorable study while ignoring contradicting evidence is, according to FTC guidance, grounds for a claim being "unsubstantiated."
The supplement industry is a $50-billion-a-year market that runs largely on structure/function claims — language carefully engineered to imply effectiveness without technically triggering FDA drug-claim standards.
Labels say "supports joint health" rather than "treats arthritis." They say "promotes cognitive function" rather than "improves memory." This framing is not accidental — it's the DSHEA architecture in practice. As long as the label includes a disclaimer that the FDA has not evaluated the claim, manufacturers can make structure/function claims with a level of evidence that would never pass FDA drug-approval review.
This creates an environment where a single small-sample study with a statistically significant subgroup result can fuel years of advertising. The GAIT trial — a $12.5 million, 1,583-patient NIH-funded study published in the New England Journal of Medicine in 2006 — found that glucosamine and chondroitin, separately or in combination, produced no statistically significant improvement over placebo in the overall study population with knee osteoarthritis. One subgroup of 354 patients with moderate-to-severe pain showed a significant result for the combination. That subgroup — roughly 22% of participants, in a single trial — became the basis for decades of marketing claims. Subsequent GAIT follow-up data at 24 months found no statistically significant effect on joint-space narrowing across any treatment group.
Similarly, ginkgo biloba has been marketed for memory support for decades. A 2026 systematic review from Georgetown University School of Medicine analyzed 82 randomized controlled trials involving 10,613 participants and concluded that ginkgo offers little to no benefit for individuals with subjective memory complaints or mild cognitive impairment. An earlier NCBI systematic review was more direct: there is no convincing evidence that ginkgo biloba extracts had a positive effect on any aspect of cognitive performance in healthy people under the age of 60 years. The product still markets prominently to exactly that demographic.
From a consumer's perspective, the phrase "studies show" functions as social proof — not as a claim that can be directly interrogated.
Most buyers don't have access to the underlying study, don't know how to read a statistical methods section, and have no way of knowing whether the cited research was the primary outcome of a preregistered trial or a post-hoc subgroup finding in a study of 47 people.
Sellers commonly advertise that a product is "backed by clinical research" when the underlying research is a single industry-funded study, conducted on a small sample, measuring a surrogate endpoint — like a biomarker change — rather than a clinical outcome like whether pain actually decreased. A biomarker improvement can be statistically significant and clinically irrelevant simultaneously.
For echinacea — one of the most widely purchased immune support supplements — a 2014 Cochrane systematic review of 24 double-blind trials involving 4,631 participants found that none of the 12 prevention comparisons showed a statistically significant reduction in cold incidence. Some treatment trials suggested shorter cold duration with certain preparations, but the Cochrane authors noted that the overall evidence for clinically relevant effects is weak, and that positive results may partly reflect reporting bias — the tendency for statistically significant results to get published while null results don't.
One responsible counterexample: some manufacturers have moved toward publishing full study data, registering trials in advance on ClinicalTrials.gov, and citing effect sizes with confidence intervals in their marketing rather than just statistical significance flags. These practices, while not yet industry-standard, represent what transparent evidence communication looks like.
The honest answer is: usually somewhere between the study's results section and the label.
Statistical significance is a mathematical tool for distinguishing signal from noise. It was never designed to be a consumer-facing quality seal — and it functions poorly as one. A p-value below 0.05 tells you the effect probably didn't happen by chance. It tells you nothing about whether you'll feel a difference, whether the effect size matters to your life, or whether the study that found it was adequately designed and powered to find it.
The three gaps that marketing routinely exploits are now clear: relative risk numbers that convert small absolute differences into large-sounding percentages; subgroup results from underpowered or exploratory analyses presented as primary findings; and single statistically significant studies cited while conflicting evidence goes unmentioned. Each of these is technically defensible. Each is designed to make a product look more effective than the totality of evidence supports.
The regulatory framework doesn't close these gaps. DSHEA's reversed burden of proof means the system depends on the FTC and FDA to catch bad actors after the fact — and with hundreds of thousands of products on the market, enforcement is necessarily selective. The FTC's 2023 notices to 670 companies named a principle; they didn't change the underlying architecture.
A few developments could narrow the gap between "statistically significant" and "actually works." The FDA's proposed guidance on new dietary ingredient notifications, still in development, could eventually require pre-market evidence of a higher quality for novel ingredients. Some researchers and journals are actively pushing for effect size and confidence interval reporting to replace or accompany p-value significance flags — a shift the American Statistical Association endorsed in guidance published in 2016. And the growing use of ClinicalTrials.gov pre-registration for supplement research would make it harder for companies to selectively report favorable subgroup results from studies that didn't find an overall effect. None of these changes are imminent or comprehensive. But they represent the direction the evidence-quality conversation is moving.
Why It Matters scale: Low stakes · Worth knowing · Matters a lot · Act on this now
Evidence Strength: Well-established — the gap between statistical and clinical significance is thoroughly documented in peer-reviewed literature and regulatory guidance
- Perception Gap: High — most consumers treat "clinically proven" as equivalent to "works reliably and meaningfully"
- Real-World Risk: Moderate — primarily financial, but can include delayed use of more effective treatments, or overconfidence in supplement-only management of conditions
- How Often It's Exploited: High — standard marketing practice in the supplement industry, documented across hundreds of FTC enforcement cases
- Risk-Benefit Ratio: Favorable — for a reader who understands the distinction; the risk is information asymmetry, which this article closes
- Medical/Regulatory Consensus: FDA requires no pre-market efficacy proof for supplements; FTC requires substantiation for "clinically proven" claims but enforcement is after-the-fact
👉 Who should care most: Anyone who buys supplements based on study citations, label claims, or influencer endorsements — particularly for joint support, memory, immune function, and weight management products, where the statistical/clinical gap is most often exploited.
👉 Who can safely ignore this: People who don't use supplements and aren't considering them; those already working with a clinician who evaluates the full evidence base on their behalf.
⚕️ LyfeiQ Score: Worth knowing — Understanding this distinction doesn't require a statistics degree. It requires three questions: how big was the effect, was this the primary result or a subgroup finding, and has it been replicated? Ask those before the next purchase.
Related: Compounded Semaglutide and Industrial Salts: What the Label Doesn't Tell You
Disclaimer: This content includes personal opinions and interpretations based on available sources and should not replace medical advice. This content includes interpretation of available research and should not replace medical advice. Although the data found in this blog and infographic has been produced and processed from sources believed to be reliable, no warranty expressed or implied can be made regarding the accuracy, completeness, legality or reliability of any such information. This disclaimer applies to any uses of the information whether isolated or aggregate uses thereof.