Independent tests

Benchmarks only when they answer a real question.

Testing is used when evidence review cannot settle a practical claim. Protocols should name the task, baseline, dataset limits, scoring method and decision relevance.

Test in progressLaunch placeholder

No public benchmark result yet

Future test pages will disclose methodology before presenting any result. No vendor score or ranking is invented for launch.

Model watchBaseline

Specialized model versus frontier model

When relevant, veterinary-specific AI should be compared with an appropriate frontier-model baseline, not only with weaker or undefined alternatives.

Test designGuardrail

Free submission, no guaranteed outcome

Free testing lowers barriers to scrutiny. It does not guarantee publication, favorable framing or inclusion in a comparison.

Benchmark principles

What every test must make clear.

A benchmark that cannot explain what it measures can become another marketing claim. VetAI Trust treats test design as part of the evidence.

Task definitionThe clinical or operational question the system is expected to answer.
Population and limitsSpecies, case types, exclusions and known sources of distribution shift.
Comparison setRelevant human workflow, software, veterinary AI and frontier-model baselines.
MetricsRecall, false positives, calibration, time saved or other endpoints matched to the claim.
Publication thresholdWhether results can be reported clearly without overstating what was tested.