Paper Check
Public validation disclosure was scarce across 71 veterinary AI products.
A 2026 systematic audit found very little validation, safety or data-provenance information in public vendor materials. It measured disclosure—not product performance—and cannot show that validation never happened.
The signal
For most products, a buyer could not inspect the basics before entering the sales process.
Across the 71 included tools, the mean unweighted transparency score was 6.4%. The paper reports that 45 of the 71 audited entries—about 63%—did not publicly disclose a single item from its 25-point transparency framework.
Why it matters
Veterinarians need more than an accuracy claim.
Before relying on a clinical AI tool, a buyer should be able to ask what population it was trained on, how it was tested, what its uncertainty looks like and where it fails. If those details are not public, clinicians cannot independently judge whether a result is likely to transfer to their species mix, equipment and case population.
This is a procurement and evidence problem. It is not, by itself, proof that a specific tool is inaccurate or unsafe.
What the study did
A seven-day market search, followed by a structured audit of archived vendor materials.
The author searched for commercially available clinical veterinary AI products from 17 to 24 November 2025. The search used four channels:
- exhibitor lists from VMX 2025, WVC 2025 and the AVMA Convention 2025;
- Crunchbase and LinkedIn searches for veterinary AI and machine-learning companies;
- the Apple App Store and Google Play Store; and
- a structured Google search, screening the first 100 results per query.
Products had to be commercially accessible through a purchase or demo route, explicitly claim AI or machine learning, and perform a clinical task. Purely administrative software, direct-to-consumer apps, academic prototypes and “coming soon” pages were excluded.
The search produced 1,353 initial records and 71 final inclusions. On 28 November 2025, the included vendor websites were archived as time-stamped offline copies. The audit then used those fixed snapshots.
How transparency was scored
The author developed a 25-item Veterinary AI Transparency Index (VATI), mapped to FDA Good Machine Learning Practice, CONSORT-AI and Coalition for Health AI model-card concepts. It covered four domains: data provenance and composition; performance and validation; safety and risk management; and usability and documentation.
Each item received binary credit only for a specific, quantitative or verifiable disclosure. Broad marketing language did not count. An isolated statistic also received no credit unless it came with context such as a sample size, confidence interval or validation-study citation.
The paper reports both an unweighted percentage and a risk-weighted score. Safety items received 3 points, validation items 2 and descriptive or administrative items 1. A blinded repeat audit of 14 entries after seven days produced intra-rater agreement above 0.90 by Cohen's kappa.
What it found
Low disclosure overall, with imaging ahead of generative and ambient tools.
Market composition
- 47 generative and ambient tools (66.2%)
- 19 diagnostic imaging tools (26.8%)
- 5 specialized tools (7.0%)
Disclosure differences
The mean risk-weighted transparency score was 13.1% for diagnostic imaging and 1.8% for generative and ambient tools (Mann–Whitney U, p = 0.003). The overall difference among product domains was also statistically significant (Kruskal–Wallis, p < 0.001).
| Audited disclosure | Imaging (n = 19) | Generative / ambient (n = 47) | p-value |
|---|---|---|---|
| Peer-reviewed evidence / accuracy metrics | 36.8% | 2.1% | < 0.001 |
| Independent test set defined | 26.3% | 2.1% | 0.006 |
| Confidence intervals reported | 15.8% | 0.0% | 0.021 |
| Risk mitigation / guardrails | 5.3% | 0.0% | 1.0 |
| Known failure modes | 5.3% | 0.0% | 0.50 |
| Bias / fairness evaluation | 5.3% | 0.0% | 0.29 |
| Signalment demographics | 0.0% | 2.1% | 1.0 |
Percentages describe what was disclosed in public documentation. They are not estimates of how many products were validated, accurate or safe.
What it does not prove
This was a transparency audit, not a product validation study.
It shows that the audited public materials did not contain any of the framework's qualifying disclosures.
No product outputs were tested against a reference standard in this study.
A study supporting a narrow task cannot automatically be extended to every marketed capability.
The findings belong to a fixed November 2025 snapshot; vendor documentation may have changed.
Limitations
Useful baseline, constrained view.
- Public material only. Technical evidence shared during sales, under contract or behind an NDA was outside the audit.
- One primary rater. The repeat audit showed strong intra-rater consistency, but the full dataset did not receive an independent second review.
- Rapidly changing market. A narrow collection window improved consistency but makes the result time-sensitive.
- North American focus. US-based search engines and North American conference lists shaped product discovery.
- New instrument. The VATI was mapped to established human-health frameworks but had not undergone formal content validation by an independent expert panel.
- Small specialized subgroup. The five Specialized/Other tools were excluded from pairwise statistical comparisons.
- Disclosure, not evidence quality. Binary VATI scoring recorded whether qualifying information was public, not whether an underlying study was methodologically strong.
Funding and disclosures
The paper reports no financial support for the work or its publication and no commercial or financial relationships that could be construed as a potential conflict of interest.
The author also disclosed using Google Gemini (Models 3 Pro) as a research-structuring aid, for scripting and statistical workflow support, troubleshooting and formatting, and for outlining, summarization and copy-editing. The paper attributes all study roles to David Brundage.
VetAI Trust take
The strongest conclusion is about inspectability.
This paper provides a valuable baseline for what a veterinary AI buyer could independently inspect in late November 2025. Its central finding is serious: public documentation often lacked the information needed to assess model fit, validation design, uncertainty and failure modes.
But evidence discipline cuts both ways. The study supports saying that public validation disclosure was often absent. It does not support calling every low-scoring product unvalidated, inaccurate or unsafe. Those questions require product-specific evidence and, where appropriate, independent testing.
For buyers, the practical response is straightforward: request the intended-use population, external test-set design, per-task performance with uncertainty, subgroup results, version information and known failure modes before clinical adoption.
Source / citation
Brundage D. A systematic audit of transparency and validation disclosure in commercial veterinary artificial intelligence. Frontiers in Veterinary Science. 2026;13:1761038.
DOI: 10.3389/fvets.2026.1761038
Read the full open-access paper
Paper Check by VetAI Trust. This review distinguishes the paper's data from the author's interpretation and from VetAI Trust's editorial assessment. See the review methodology or return to all Evidence.