A client asks about the risk of joint infection after a steroid injection. You remember a JAVMA paper on it, something reassuring, a few hundred injections, one bad outcome. You quote a rate off the top of your head. That rate is real. It is also built on one dog.
That is the trap. Every JAVMA study reports a headline sample size, the number in the title or the first line of the abstract. But the specific claim you're about to repeat to a client almost never rests on that headline number. It rests on whatever smaller n sits inside the table the claim actually came from. Checking which n you're standing on takes fifteen seconds and it is the difference between citing a finding and citing a coincidence.
Twelve Cases, Eight Dogs
Take a paper vet cardiologists still cite from memory: Fascetti et al. "Taurine Deficiency in Dogs with Dilated Cardiomyopathy: 12 Cases (1997-2001)," JAVMA, 2003. The title says 12. That is a case series, not a trial, and 12 is already a small n to build a clinical impression on.
But the number that usually gets repeated is narrower than 12. Repeat echocardiograms were obtained in 9 of the 12 dogs, and the specific measurement most often quoted, E-point to septal separation, was reported on 8 dogs. That is the n behind the statistic, not the n in the title. An 8-dog before-and-after comparison in a retrospective case series is a real data point. It is not the same claim as "a study of 12 dogs found."
The paper's own design is a limitation the authors would not dispute: no control group, no randomization, referral-population dogs that may not represent the general patient pool. We are not second-guessing the finding. We are pointing out that the n changes depending on which sentence you're citing, and most people citing it from memory use the biggest one.
One Event Out of 505
The joint-injection paper above is a cleaner example of the same problem, because it has three different sample sizes stacked on top of each other. Miller, Carney, Markmann, and Frye published a retrospective review in JAVMA (2023;261:397-402) built from Cornell's hospital records: 505 joint injections, across 283 patient visits, in 178 client-owned dogs, treated for osteoarthritis and related joint disease between 2010 and 2022.
The finding people quote is the safety one: one case of septic arthritis out of 505 injections. Said out loud, that becomes "less than one in five hundred." That framing is not wrong, but it is doing more work than the data can support. It is a single event. The authors themselves flagged that it was unclear whether the infection came from the injection at all, versus an unrelated hematogenous source. A rate built on one occurrence swings enormously with the next case that comes through the door. It could double at the next data pull and the paper's math would not have been wrong, because n=1 events do not produce stable rates. Minor complications, by contrast, were common enough to trust more: 70 of 283 visits, mostly transient soreness. That is a real subgroup with a real n behind it. The septic arthritis number is not in the same category, and the paper's own denominator, 505 injections versus 178 dogs, versus 283 visits, is a reminder that "out of how many" needs a specific answer before you repeat a rate.
What the P Value Is Hiding
There is a second, quieter version of this problem, and it does not need a headline-grabbing case count to show up. Weng and Messam laid it out in the Journal of Veterinary Internal Medicine (2025;39:e17258, "Reporting and Interpreting Statistical Results in Veterinary Medicine: Calling for Change"). They built a hypothetical example, not real patient data, of four therapeutic renal diets tested against a control in cats with chronic kidney disease. One diet showed an 11.9 percentage-point difference in body weight change with a 95% confidence interval spanning 6.2 to 17.6, P=.036. Another diet showed a 1.6 percentage-point difference with a much tighter interval, 1.1 to 2.1, P=.019.
The second diet has the smaller P value. It also describes a smaller, more clinically trivial effect, measured with more precision because of the underlying sample size and variability, not because the effect itself is more important. A P value tells you whether a result cleared a threshold. It does not tell you how big the effect is or how much you should trust its size. The confidence interval width is doing that job, and the interval width is a direct readout of how much data is actually behind the number.
JAVMA's own editorial page has been saying something close to this out loud. Constance White, writing in JAVMA (2025;263:645, "The Perilous P Value"), pointed out that low-powered studies, the kind common in clinical veterinary research where 20 or 40 cases is a normal sample, are prone to missing real differences and returning a non-significant P value even when a true effect exists. A "no significant difference" finding in an underpowered study is not evidence of no difference. It is often just evidence of too small an n to detect one.
The Check
None of this means retrospective case series or small trials are useless. Most of veterinary clinical evidence is built on exactly this kind of study, because multi-center RCTs with hundreds of patients are rare in a profession this size. The fix is not skepticism about small studies. It is precision about which number came from where.
Before repeating a study finding to a client or in a case discussion, trace the specific number back to its table, not its abstract. Ask what n produced that specific statistic, whether it is a rate built on a single-digit event count, and whether the paper's own limitations section already flagged the gap. The abstract's case count is the marketing copy. The table's denominator is the data.