How to Read a Clinical Trial: A Dog Owner’s Guide
Our Fact-Checking Team —
On this page
“Clinically proven.” “Backed by science.” “Supported by published research.” These phrases show up on supplement labels with reassuring frequency, and they’re designed to end the conversation right there. But what do they actually mean? And how is a dog owner without a biostatistics degree supposed to tell whether the science behind a product is robust — or merely rhetorical?
I wrote this guide because I think most owners underestimate how much they can figure out on their own. You don’t need a PhD. You need a framework and a willingness to look past the marketing summary and ask a few pointed questions. The rest of this piece is that framework.
The Hierarchy of Evidence
| Level | Study Type | Relative Strength | What It Can Tell You |
|---|---|---|---|
| 1 | Systematic review / meta-analysis | Highest | Pooled conclusion across many trials |
| 2 | Randomized controlled trial (RCT) | High | Causal effect of one intervention |
| 3 | Cohort / case-control | Moderate | Associations, not causation |
| 4 | In-vitro / animal model | Low (for efficacy) | Mechanism and hypothesis generation |
| 5 | Expert opinion / marketing | Lowest | No evidentiary weight on its own |
Not all research carries equal weight, and this is the first thing worth internalizing. Evidence sits on a hierarchy:
- Systematic reviews and meta-analyses — pooling data from multiple trials. Highest level.
- Randomized controlled trials (RCTs) — the gold standard for intervention studies.
- Cohort and case-control studies — observational; show association, not causation.
- Case series and case reports — individual patient outcomes; hypothesis-generating only.
- In vitro studies — test-tube experiments; necessary but far from sufficient.
- Expert opinion and anecdotes — lowest level; no controlled comparison.
So when a supplement company cites “published research,” your first move is to find out which level they mean. An in vitro study showing that a compound kills bacteria in a petri dish isn’t evidence that the product works in a living dog. A case report of one dog improving isn’t evidence of efficacy. I say this gently, because these studies get waved around constantly. An RCT in the target species, with appropriate controls and an adequate sample size, is the minimum standard for an efficacy claim — anything less is a suggestion, not a demonstration.


Anatomy of a Clinical Trial
Study Design
The strongest design for supplement evaluation is the randomized, double-blind, placebo-controlled trial:
- Randomized: Dogs are assigned to treatment or control groups by chance, not by owner or investigator choice. This distributes confounders (age, breed, baseline health) evenly.
- Double-blind: Neither the owner nor the assessing veterinarian knows which group each dog is in. This eliminates observer bias in subjective outcomes (fecal scoring, behavior assessment).
- Placebo-controlled: The control group receives an identical-looking inactive product. This accounts for the placebo effect (which exists in veterinary medicine through owner-report bias) and natural disease fluctuation.
If a study is missing any of these elements, make a note of which ones and how that limits what you can conclude. An open-label trial isn’t worthless — it’s just worth less, and you should discount it accordingly.
Population
Ask:
- How many dogs were enrolled? (Sample size)
- What breeds, ages, and health statuses were included?
- Were dogs with concurrent medications or diseases excluded?
- Does the study population resemble your dog?
This matters more than it sounds. A trial in 15 healthy adult Beagles may not predict a thing about your 11-year-old Golden Retriever with early kidney disease and concurrent NSAID use. I always check who was actually in the study before I trust what it claims to show.
Intervention
- What exact product/strain/dose was tested?
- How long was the intervention period?
- Was compliance monitored (did owners actually administer the product as directed)?
Outcomes
- Primary endpoint: The main outcome the study was designed to measure. This is what the sample size calculation is based on.
- Secondary endpoints: Additional outcomes measured but not powered for definitive conclusions.
- Objective vs. subjective: Fecal calprotectin concentration (objective, lab-measured) carries more weight than “owner-perceived improvement in energy” (subjective, unblinded).
Understanding P-Values (Without the Math)
The p-value is the most cited — and the most misunderstood — statistic in all of biomedical research. I’ve watched reviewers and marketers alike treat it as a verdict, which it isn’t.
What It Is
A p-value answers one specific question: “If there were truly no difference between treatment and control, how likely would we be to see a result this extreme (or more extreme) just by random chance?”
- p = 0.03 → There’s a 3% probability that random variation alone would produce this result.
- p = 0.001 → There’s a 0.1% probability. Stronger evidence against the null.
- p = 0.08 → There’s an 8% probability. Conventionally “not significant” but not proof of no effect.
What It Is NOT
- It’s NOT the probability that the treatment works.
- It’s NOT the probability that the null hypothesis is true.
- It does NOT measure the size or importance of the effect.
- It does NOT tell you whether the study was well-designed.
The 0.05 Threshold Is Arbitrary
The convention of p < 0.05 as “significant” was proposed by Ronald Fisher in 1925 as a convenient cutoff, not a natural law. He essentially picked it because it was handy. A p-value of 0.051 isn’t meaningfully different from 0.049, yet one gets called a failure and the other a finding. Context, effect size, and study quality all matter more than whether a number happens to cross an arbitrary line.
Confidence Intervals: The More Informative Statistic
A confidence interval (CI) tells you the range of plausible values for the true treatment effect.
Example: “The probiotic group showed 1.2 days faster diarrhea resolution (95% CI: 0.3 to 2.1 days, p = 0.01).”
- The point estimate is 1.2 days.
- The 95% CI means: if we repeated this study 100 times, 95 of those repetitions would produce a result between 0.3 and 2.1 days.
- The CI doesn’t cross zero → the effect is statistically significant.
- The CI is relatively narrow → reasonable precision.
Compare: “The probiotic group showed 1.2 days faster resolution (95% CI: -0.8 to 3.2 days, p = 0.22).”
- Same point estimate, but the CI crosses zero and is very wide.
- This study is inconclusive. The true effect could be a 3.2-day benefit OR a 0.8-day harm.
- The sample size was likely too small to draw firm conclusions.
Practical rule: look at the confidence interval, not just the p-value. A narrow CI that excludes zero is strong evidence. A wide CI that includes zero is weak evidence, full stop — no matter how promising the point estimate looks.
Sample Size and Statistical Power
Why Veterinary Studies Are Often Small
Enrolling client-owned dogs in clinical trials is expensive and logistically challenging. Owners must commit to follow-up visits, compliance monitoring, and potential placebo assignment. As a result, many veterinary supplement trials enroll 10-30 dogs per group.
The Power Problem
Statistical power is the probability of detecting a real effect if one exists. Power depends on:
- Sample size (more dogs = more power)
- Effect size (larger effects are easier to detect)
- Outcome variability (less variable outcomes need fewer subjects)
- Significance threshold (lower alpha = less power)
Put some numbers on it. A study with 12 dogs per group has roughly 80% power to detect a very large effect (Cohen’s d > 1.2). It has less than 50% power to detect a moderate effect (d = 0.6). The practical consequence matters: a “negative” result in a small study doesn’t prove the treatment is ineffective. Half the time it only proves the study was too small to see the effect that was there. An n=12 is exactly the kind of sample I refuse to generalize from.
What to Look For
- Did the authors report a sample size calculation before the study? (Prospective power analysis)
- Is the sample size justified for the primary endpoint?
- For negative results: is the CI narrow enough to exclude a clinically meaningful effect?
Conflicts of Interest: Following the Money
Why It Matters
A 2023 meta-epidemiological analysis in PLOS ONE found that industry-funded nutrition studies were 2.4 times more likely to report conclusions favorable to the sponsor’s product than independently funded studies of the same interventions. Let me be clear what that does and doesn’t mean. It doesn’t mean industry-funded research is fraudulent. It means that design choices — population selection, comparator choice, endpoint selection, statistical methods — can be made, consciously or not, to steer toward a desired outcome. The bias is usually in the framing, not the data.
What to Check
- Funding disclosure: Who paid for the study? Is it stated clearly?
- Author affiliations: Are authors employees of the manufacturer? Do they hold patents or equity?
- Comparator choice: Was the product compared to placebo (easy to beat) or to an established effective treatment (harder)?
- Endpoint selection: Are primary endpoints clinically meaningful, or are they surrogate markers chosen because they’re likely to show a difference?
- Publication venue: Is the journal peer-reviewed and indexed in PubMed? Or is it a low-impact, pay-to-publish outlet?
- Replication: Has any independent group (unaffiliated with the manufacturer) replicated the findings?
Industry Funding Is Not Disqualifying
Most supplement research is necessarily industry-funded, because government agencies rarely fund product-specific trials, and I don’t hold that against a study on its own. An industry-funded trial can be rigorous, transparent, and reproducible. The question is whether the design and analysis are sound regardless of who signed the checks. Pre-registered protocols, independent statistical analysis, and full data transparency are what I look for.
Red Flags in Supplement Research
- No control group: “We gave 20 dogs our product and owners reported improvement.” Without a placebo group, you can’t distinguish treatment effect from natural fluctuation, owner expectation, or regression to the mean.
- Open-label design: Owners know their dog is receiving the product. Owner-reported outcomes (energy, coat quality, behavior) are highly susceptible to expectation bias.
- Multiple endpoints without correction: Testing 20 outcomes and reporting the 2 that reach p < 0.05 isn’t evidence. This is “p-hacking” or the “multiple comparisons problem.”
- Post-hoc subgroup analysis: “The product didn’t work overall, but in the subgroup of dogs aged 5-7 with brown coats, it was significant.” Subgroup findings are hypothesis-generating, not confirmatory.
- Surrogate endpoints only: “Increased fecal Lactobacillus counts” doesn’t equal “improved health.” Microbiome changes without clinical outcome data are mechanistically interesting but clinically incomplete.
- Abstract-only publication: Findings presented at a conference but never published as a full peer-reviewed paper haven’t undergone complete scrutiny.
A Practical Checklist for Dog Owners
When a product claims clinical evidence, ask these questions:
- Is there a published, peer-reviewed RCT in dogs (not just in vitro, not just in humans)?
- Was the study randomized, blinded, and placebo-controlled?
- What was the sample size? Is it adequate for the claimed effect?
- What were the primary endpoints? Are they clinically meaningful?
- What was the effect size and confidence interval?
- Who funded the study? Are there declared conflicts of interest?
- Has the finding been replicated by an independent group?
- Does the study population match my dog (breed, age, health status)?
If a product can’t satisfactorily answer questions 1 through 4, I treat its “clinically proven” claim as marketing until proven otherwise. That isn’t cynicism; it’s just the bar.
The Five Things That Actually Tell You Something
You don’t need a statistics degree to evaluate supplement evidence. You need a framework: study design, sample size, endpoints, effect size, and funding source. Those five elements will tell you more than any p-value or marketing summary ever will.
The supplement industry profits from information asymmetry — from owners who see “published research” and stop asking questions. The antidote isn’t cynicism. It’s literacy. Good science withstands scrutiny; I’ve never seen a solid trial fall apart under a few honest questions. When the science is weak, those same questions reveal the gap almost immediately.
Frequently Asked Questions
What does a p-value actually mean?
A p-value represents the probability of observing results as extreme as those in the study IF the null hypothesis (no real difference) were true. A p-value of 0.03 means there is a 3% chance the observed difference is due to random variation alone. It does NOT mean the treatment has a 97% chance of working, nor does it measure effect size.
Why does sample size matter in veterinary studies?
Small sample sizes (common in veterinary research, often n=10-30 per group) produce wide confidence intervals and low statistical power. A study with 12 dogs per group may miss a real 20% improvement or overstate a trivial effect. Larger samples produce more reliable estimates of true treatment effects.
How do I spot conflicts of interest in supplement research?
Check the funding source and author affiliations. Industry-funded studies show statistically higher rates of favorable conclusions. Look for: manufacturer employees as authors, funding from the product’s maker, no independent replication, and studies published only in low-impact journals.
What is the difference between statistical significance and clinical relevance?
Statistical significance (p < 0.05) means the result is unlikely due to chance. Clinical relevance means the effect is large enough to matter in practice. A study can find a statistically significant 2% improvement that has no meaningful impact on the dog’s quality of life. Always evaluate effect size, not just p-values.
References
- Shmalberg J, Montalbano C, Morelli G, Buckley GJ. A Randomized Double Blinded Placebo-Controlled Clinical Trial of a Probiotic or Metronidazole for Acute Canine Diarrhea. Frontiers in Veterinary Science. 2019;6:163. DOI: 10.3389/fvets.2019.00163
- Bonel-Ayuso DP, Roca M, Llopis M, et al. Effects of Postbiotic Administration on Canine Health: A Systematic Review and Meta-Analysis. Microorganisms. 2025;13(7):1572. PMID: 40732081
- Weese JS, Martin H. Assessment of commercial probiotic bacterial contents and label accuracy. Canadian Veterinary Journal. 2011;52(1):43-46. PMC3003573
- Manson-Smith DF, Stewart CJ, Bhatt A, et al. Longitudinal Survey of Fecal Microbiota in Healthy Dogs Administered a Commercial Probiotic. Frontiers in Veterinary Science. 2021;8:664318. DOI: 10.3389/fvets.2021.664318
- Sordillo A, Casella L, Turcotte R, Sheth RU. A Novel Postbiotic Reduces Canine Halitosis. Animals (Basel). 2025;15(11):1596. PMID: 40509062
