“Clinically proven.” “Backed by science.” “Supported by published research.” These phrases appear on supplement labels and marketing pages with reassuring frequency. But what do they actually mean? And how can a dog owner — without a biostatistics degree — evaluate whether the science behind a product is robust or rhetorical?
This guide provides a practical framework for reading and interpreting clinical trial evidence. No PhD required. Just a willingness to look past the marketing summary and ask a few pointed questions.
The Hierarchy of Evidence
Not all research carries equal weight. Evidence exists on a hierarchy:
- Systematic reviews and meta-analyses — pooling data from multiple trials. Highest level.
- Randomized controlled trials (RCTs) — the gold standard for intervention studies.
- Cohort and case-control studies — observational; show association, not causation.
- Case series and case reports — individual patient outcomes; hypothesis-generating only.
- In vitro studies — test-tube experiments; necessary but far from sufficient.
- Expert opinion and anecdotes — lowest level; no controlled comparison.
When a supplement company cites “published research,” determine which level they are referencing. An in vitro study showing that a compound kills bacteria in a petri dish is not evidence that the product works in a living dog. A case report of one dog improving is not evidence of efficacy. An RCT in the target species, with appropriate controls and adequate sample size, is the minimum standard for an efficacy claim.
Anatomy of a Clinical Trial
Study Design
The strongest design for supplement evaluation is the randomized, double-blind, placebo-controlled trial:
- Randomized: Dogs are assigned to treatment or control groups by chance, not by owner or investigator choice. This distributes confounders (age, breed, baseline health) evenly.
- Double-blind: Neither the owner nor the assessing veterinarian knows which group each dog is in. This eliminates observer bias in subjective outcomes (fecal scoring, behavior assessment).
- Placebo-controlled: The control group receives an identical-looking inactive product. This accounts for the placebo effect (which exists in veterinary medicine through owner-report bias) and natural disease fluctuation.
If a study lacks any of these elements, note which ones are missing and how that limits interpretation.
Population
Ask:
- How many dogs were enrolled? (Sample size)
- What breeds, ages, and health statuses were included?
- Were dogs with concurrent medications or diseases excluded?
- Does the study population resemble your dog?
A trial in 15 healthy adult Beagles may not predict outcomes in your 11-year-old Golden Retriever with early kidney disease and concurrent NSAID use.
Intervention
- What exact product/strain/dose was tested?
- How long was the intervention period?
- Was compliance monitored (did owners actually administer the product as directed)?
Outcomes
- Primary endpoint: The main outcome the study was designed to measure. This is what the sample size calculation is based on.
- Secondary endpoints: Additional outcomes measured but not powered for definitive conclusions.
- Objective vs. subjective: Fecal calprotectin concentration (objective, lab-measured) carries more weight than “owner-perceived improvement in energy” (subjective, unblinded).
Understanding P-Values (Without the Math)
The p-value is the most cited and most misunderstood statistic in biomedical research.
What It Is
A p-value answers one specific question: “If there were truly no difference between treatment and control, how likely would we be to see a result this extreme (or more extreme) just by random chance?”
- p = 0.03 → There is a 3% probability that random variation alone would produce this result.
- p = 0.001 → There is a 0.1% probability. Stronger evidence against the null.
- p = 0.08 → There is an 8% probability. Conventionally “not significant” but not proof of no effect.
What It Is NOT
- It is NOT the probability that the treatment works.
- It is NOT the probability that the null hypothesis is true.
- It does NOT measure the size or importance of the effect.
- It does NOT tell you whether the study was well-designed.
The 0.05 Threshold Is Arbitrary
The convention of p < 0.05 as “significant” was proposed by Ronald Fisher in 1925 as a convenient cutoff, not a natural law. A p-value of 0.051 is not meaningfully different from 0.049. Context, effect size, and study quality matter more than whether a number crosses an arbitrary line.
Confidence Intervals: The More Informative Statistic
A confidence interval (CI) tells you the range of plausible values for the true treatment effect.
Example: “The probiotic group showed 1.2 days faster diarrhea resolution (95% CI: 0.3 to 2.1 days, p = 0.01).”
- The point estimate is 1.2 days.
- The 95% CI means: if we repeated this study 100 times, 95 of those repetitions would produce a result between 0.3 and 2.1 days.
- The CI does not cross zero → the effect is statistically significant.
- The CI is relatively narrow → reasonable precision.
Compare: “The probiotic group showed 1.2 days faster resolution (95% CI: -0.8 to 3.2 days, p = 0.22).”
- Same point estimate, but the CI crosses zero and is very wide.
- This study is inconclusive. The true effect could be a 3.2-day benefit OR a 0.8-day harm.
- The sample size was likely too small to draw firm conclusions.
Practical rule: Always look at the confidence interval, not just the p-value. A narrow CI that excludes zero is strong evidence. A wide CI that includes zero is weak evidence regardless of the point estimate.
Sample Size and Statistical Power
Why Veterinary Studies Are Often Small
Enrolling client-owned dogs in clinical trials is expensive and logistically challenging. Owners must commit to follow-up visits, compliance monitoring, and potential placebo assignment. As a result, many veterinary supplement trials enroll 10-30 dogs per group.
The Power Problem
Statistical power is the probability of detecting a real effect if one exists. Power depends on:
- Sample size (more dogs = more power)
- Effect size (larger effects are easier to detect)
- Outcome variability (less variable outcomes need fewer subjects)
- Significance threshold (lower alpha = less power)
A study with 12 dogs per group has approximately 80% power to detect a very large effect (Cohen’s d > 1.2). It has less than 50% power to detect a moderate effect (d = 0.6). This means: a “negative” result in a small study does not prove the treatment is ineffective. It may simply prove the study was too small to detect the effect.
What to Look For
- Did the authors report a sample size calculation before the study? (Prospective power analysis)
- Is the sample size justified for the primary endpoint?
- For negative results: is the CI narrow enough to exclude a clinically meaningful effect?
Conflicts of Interest: Following the Money
Why It Matters
A 2023 meta-epidemiological analysis in PLOS ONE found that industry-funded nutrition studies were 2.4 times more likely to report conclusions favorable to the sponsor’s product compared to independently funded studies of the same interventions. This does not mean industry-funded research is fraudulent. It means that design choices (population selection, comparator choice, endpoint selection, statistical methods) can be made — consciously or not — to favor a desired outcome.
What to Check
- Funding disclosure: Who paid for the study? Is it stated clearly?
- Author affiliations: Are authors employees of the manufacturer? Do they hold patents or equity?
- Comparator choice: Was the product compared to placebo (easy to beat) or to an established effective treatment (harder)?
- Endpoint selection: Are primary endpoints clinically meaningful, or are they surrogate markers chosen because they are likely to show a difference?
- Publication venue: Is the journal peer-reviewed and indexed in PubMed? Or is it a low-impact, pay-to-publish outlet?
- Replication: Has any independent group (unaffiliated with the manufacturer) replicated the findings?
Industry Funding Is Not Disqualifying
Most supplement research is necessarily industry-funded, because government agencies rarely fund product-specific trials. An industry-funded study can be rigorous, transparent, and reproducible. The key is whether the design and analysis are sound regardless of who paid. Look for pre-registered protocols, independent statistical analysis, and full data transparency.
Red Flags in Supplement Research
- No control group: “We gave 20 dogs our product and owners reported improvement.” Without a placebo group, you cannot distinguish treatment effect from natural fluctuation, owner expectation, or regression to the mean.
- Open-label design: Owners know their dog is receiving the product. Owner-reported outcomes (energy, coat quality, behavior) are highly susceptible to expectation bias.
- Multiple endpoints without correction: Testing 20 outcomes and reporting the 2 that reach p < 0.05 is not evidence. This is “p-hacking” or the “multiple comparisons problem.”
- Post-hoc subgroup analysis: “The product didn’t work overall, but in the subgroup of dogs aged 5-7 with brown coats, it was significant.” Subgroup findings are hypothesis-generating, not confirmatory.
- Surrogate endpoints only: “Increased fecal Lactobacillus counts” does not equal “improved health.” Microbiome changes without clinical outcome data are mechanistically interesting but clinically incomplete.
- Abstract-only publication: Findings presented at a conference but never published as a full peer-reviewed paper have not undergone complete scrutiny.
A Practical Checklist for Dog Owners
When a product claims clinical evidence, ask these questions:
- Is there a published, peer-reviewed RCT in dogs (not just in vitro, not just in humans)?
- Was the study randomized, blinded, and placebo-controlled?
- What was the sample size? Is it adequate for the claimed effect?
- What were the primary endpoints? Are they clinically meaningful?
- What was the effect size and confidence interval?
- Who funded the study? Are there declared conflicts of interest?
- Has the finding been replicated by an independent group?
- Does the study population match my dog (breed, age, health status)?
If a product cannot satisfactorily answer questions 1-4, treat its “clinically proven” claim with appropriate skepticism.
The Bottom Line
You do not need a statistics degree to evaluate supplement evidence. You need a framework: study design, sample size, endpoints, effect size, and funding source. These five elements tell you more than any p-value or marketing summary.
The supplement industry benefits from information asymmetry — from owners who see “published research” and stop asking questions. The antidote is not cynicism. It is literacy. The science, when it is good, withstands scrutiny. When it is not, a few pointed questions reveal the gap.
Frequently Asked Questions
What does a p-value actually mean?
A p-value represents the probability of observing results as extreme as those in the study IF the null hypothesis (no real difference) were true. A p-value of 0.03 means there is a 3% chance the observed difference is due to random variation alone. It does NOT mean the treatment has a 97% chance of working, nor does it measure effect size.
Why does sample size matter in veterinary studies?
Small sample sizes (common in veterinary research, often n=10-30 per group) produce wide confidence intervals and low statistical power. A study with 12 dogs per group may miss a real 20% improvement or overstate a trivial effect. Larger samples produce more reliable estimates of true treatment effects.
How do I spot conflicts of interest in supplement research?
Check the funding source and author affiliations. Industry-funded studies show statistically higher rates of favorable conclusions. Look for: manufacturer employees as authors, funding from the product’s maker, no independent replication, and studies published only in low-impact journals.
What is the difference between statistical significance and clinical relevance?
Statistical significance (p < 0.05) means the result is unlikely due to chance. Clinical relevance means the effect is large enough to matter in practice. A study can find a statistically significant 2% improvement that has no meaningful impact on the dog’s quality of life. Always evaluate effect size, not just p-values.