Paper Check

Public validation disclosure was scarce across 71 veterinary AI products.

A 2026 systematic audit found very little validation, safety or data-provenance information in public vendor materials. It measured disclosure—not product performance—and cannot show that validation never happened.

PaperA systematic audit of transparency and validation disclosure in commercial veterinary artificial intelligence
AuthorDavid Brundage
Published05 March 2026
JournalFrontiers in Veterinary Science, volume 13

The signal

For most products, a buyer could not inspect the basics before entering the sales process.

Across the 71 included tools, the mean unweighted transparency score was 6.4%. The paper reports that 45 of the 71 audited entries—about 63%—did not publicly disclose a single item from its 25-point transparency framework.

Why it matters

Veterinarians need more than an accuracy claim.

Before relying on a clinical AI tool, a buyer should be able to ask what population it was trained on, how it was tested, what its uncertainty looks like and where it fails. If those details are not public, clinicians cannot independently judge whether a result is likely to transfer to their species mix, equipment and case population.

This is a procurement and evidence problem. It is not, by itself, proof that a specific tool is inaccurate or unsafe.

What the study did

A seven-day market search, followed by a structured audit of archived vendor materials.

The author searched for commercially available clinical veterinary AI products from 17 to 24 November 2025. The search used four channels:

  • exhibitor lists from VMX 2025, WVC 2025 and the AVMA Convention 2025;
  • Crunchbase and LinkedIn searches for veterinary AI and machine-learning companies;
  • the Apple App Store and Google Play Store; and
  • a structured Google search, screening the first 100 results per query.

Products had to be commercially accessible through a purchase or demo route, explicitly claim AI or machine learning, and perform a clinical task. Purely administrative software, direct-to-consumer apps, academic prototypes and “coming soon” pages were excluded.

The search produced 1,353 initial records and 71 final inclusions. On 28 November 2025, the included vendor websites were archived as time-stamped offline copies. The audit then used those fixed snapshots.

How transparency was scored

The author developed a 25-item Veterinary AI Transparency Index (VATI), mapped to FDA Good Machine Learning Practice, CONSORT-AI and Coalition for Health AI model-card concepts. It covered four domains: data provenance and composition; performance and validation; safety and risk management; and usability and documentation.

Each item received binary credit only for a specific, quantitative or verifiable disclosure. Broad marketing language did not count. An isolated statistic also received no credit unless it came with context such as a sample size, confidence interval or validation-study citation.

The paper reports both an unweighted percentage and a risk-weighted score. Safety items received 3 points, validation items 2 and descriptive or administrative items 1. A blinded repeat audit of 14 entries after seven days produced intra-rater agreement above 0.90 by Cohen's kappa.

What it found

Low disclosure overall, with imaging ahead of generative and ambient tools.

71commercial tools included
6.4%mean unweighted VATI score
45audited entries with no qualifying public item
1vendor disclosing signalment distribution and subgroup performance

Market composition

  • 47 generative and ambient tools (66.2%)
  • 19 diagnostic imaging tools (26.8%)
  • 5 specialized tools (7.0%)

Disclosure differences

The mean risk-weighted transparency score was 13.1% for diagnostic imaging and 1.8% for generative and ambient tools (Mann–Whitney U, p = 0.003). The overall difference among product domains was also statistically significant (Kruskal–Wallis, p < 0.001).

Public disclosure rates reported in the paper
Audited disclosureImaging
(n = 19)
Generative / ambient
(n = 47)
p-value
Peer-reviewed evidence / accuracy metrics36.8%2.1%< 0.001
Independent test set defined26.3%2.1%0.006
Confidence intervals reported15.8%0.0%0.021
Risk mitigation / guardrails5.3%0.0%1.0
Known failure modes5.3%0.0%0.50
Bias / fairness evaluation5.3%0.0%0.29
Signalment demographics0.0%2.1%1.0

Percentages describe what was disclosed in public documentation. They are not estimates of how many products were validated, accurate or safe.

What it does not prove

This was a transparency audit, not a product validation study.

It does not prove that 63.3% of products were never validated.

It shows that the audited public materials did not contain any of the framework's qualifying disclosures.

It does not measure diagnostic accuracy, clinical benefit or patient harm.

No product outputs were tested against a reference standard in this study.

It does not establish that evidence for one feature covers an entire platform.

A study supporting a narrow task cannot automatically be extended to every marketed capability.

It does not describe today's market.

The findings belong to a fixed November 2025 snapshot; vendor documentation may have changed.

Limitations

Useful baseline, constrained view.

  • Public material only. Technical evidence shared during sales, under contract or behind an NDA was outside the audit.
  • One primary rater. The repeat audit showed strong intra-rater consistency, but the full dataset did not receive an independent second review.
  • Rapidly changing market. A narrow collection window improved consistency but makes the result time-sensitive.
  • North American focus. US-based search engines and North American conference lists shaped product discovery.
  • New instrument. The VATI was mapped to established human-health frameworks but had not undergone formal content validation by an independent expert panel.
  • Small specialized subgroup. The five Specialized/Other tools were excluded from pairwise statistical comparisons.
  • Disclosure, not evidence quality. Binary VATI scoring recorded whether qualifying information was public, not whether an underlying study was methodologically strong.

Funding and disclosures

The paper reports no financial support for the work or its publication and no commercial or financial relationships that could be construed as a potential conflict of interest.

The author also disclosed using Google Gemini (Models 3 Pro) as a research-structuring aid, for scripting and statistical workflow support, troubleshooting and formatting, and for outlining, summarization and copy-editing. The paper attributes all study roles to David Brundage.

VetAI Trust take

The strongest conclusion is about inspectability.

This paper provides a valuable baseline for what a veterinary AI buyer could independently inspect in late November 2025. Its central finding is serious: public documentation often lacked the information needed to assess model fit, validation design, uncertainty and failure modes.

But evidence discipline cuts both ways. The study supports saying that public validation disclosure was often absent. It does not support calling every low-scoring product unvalidated, inaccurate or unsafe. Those questions require product-specific evidence and, where appropriate, independent testing.

For buyers, the practical response is straightforward: request the intended-use population, external test-set design, per-task performance with uncertainty, subgroup results, version information and known failure modes before clinical adoption.

Source / citation

Brundage D. A systematic audit of transparency and validation disclosure in commercial veterinary artificial intelligence. Frontiers in Veterinary Science. 2026;13:1761038.

DOI: 10.3389/fvets.2026.1761038

Read the full open-access paper

Paper Check by VetAI Trust. This review distinguishes the paper's data from the author's interpretation and from VetAI Trust's editorial assessment. See the review methodology or return to all Evidence.