Claim Check
Did SignalPET's AI “match the best radiologist”?
The reported accuracy point estimates were close: 90.2% for SignalRAY and 91.0% for the best-accuracy radiologist. But failure to detect a statistically significant difference does not itself demonstrate equivalence.
The claim
“Equivalent Accuracy: SignalPET’s AI matched the best radiologist in overall accuracy.”
SignalPET published this wording on its article “New Research Validates SignalPet’s AI in Radiographic Interpretation”. The page was live when reviewed on 18 August 2026. It identifies SignalPET as the author and shows a publication date of 24 March 2025.
Our status
Partially supported. Equivalence was not demonstrated.
The narrower observation is supported: in this dataset and under this reference-standard method, the reported accuracy point estimates were close—90.2% for SignalRAY and 91.0% for the best-accuracy radiologist. The study tested for a difference and reported p = 0.0561.
The stronger wording “Equivalent Accuracy” is not demonstrated by the study design. Failure to detect a significant difference does not itself establish equivalence, non-inferiority, or equal clinical diagnostic utility.
The evidence
A small reader study of coded descriptive findings.
The supporting paper was published in Frontiers in Veterinary Science on 21 February 2025. It compared SignalRAY with 11 board-certified veterinary radiologists holding ECVDI or ACVR credentials.
- 50 radiographic studies: 40 canine and 10 feline, retrospectively drawn from one institutional PACS.
- Reader coverage: each radiologist reported 24 or 25 studies; the AI processed all 50. Each study received a mean of 5.46 radiologist reports.
- AI task: images only, producing coded normal/abnormal descriptive findings. The AI did not receive the clinical context supplied to radiologists.
- Human task: full written descriptive and diagnostic reports, later manually reduced to codes matching the AI output.
- Differential diagnoses: excluded from evaluation because the tested AI did not generate them.
- Software timing: SignalRAY was accessed in July 2022. The paper says it was continuously updated and had no version numbers.
How “correct” was defined
The study did not use pathology, surgery, follow-up outcomes, or a separate adjudication panel as an independent reference standard. Instead, the normal/abnormal state of each finding was set by the majority of the observers assessing it.
The AI was itself counted as an observer in that consensus. If an observer did not mention a finding raised by another observer, the omission was treated as a normal judgment. This reference standard measures agreement with the constructed consensus, not independently confirmed clinical truth.
What the data show
Close overall accuracy, different error profiles.
| Observer | Accuracy | Sensitivity | Specificity |
|---|---|---|---|
| SignalRAY | 90.2% | 68.8% | 94.4% |
| Best-accuracy radiologist | 91.0% | 77.9% | 93.6% |
| Median-accuracy radiologist | 87.5% | 86.0% | 87.8% |
“Best” and “median” depend on the metric. The same radiologist was not necessarily best for accuracy, sensitivity and specificity.
The headline comparison
SignalRAY's accuracy estimate was 0.8 percentage points below the best-accuracy radiologist: 90.2% versus 91.0%. The paper's Bonferroni-corrected proportion z-test reported p = 0.0561 for that difference. Against the median-accuracy radiologist, SignalRAY was 2.7 points higher; the paper reported p = 0.0022.
The error profile matters
SignalRAY was less sensitive to abnormal findings than nine radiologists, with the paper reporting p < 0.0001 for those comparisons. It was more specific for normal findings than ten radiologists, also reported at p < 0.0001. In plain language: this system was stronger at agreeing that a finding was normal than at finding abnormalities.
| Subset | Accuracy | Sensitivity | Specificity |
|---|---|---|---|
| Low ambiguity | 92.7% | 57.8% | 96.2% |
| High ambiguity | 70.1% | 44.4% | 78.5% |
The paper did not report confidence intervals for the headline performance estimates or the AI–best-radiologist difference.
The statistical catch
“Not different” and “equivalent” answer different questions.
Its z-test produced p = 0.0561 for the two accuracy estimates. At a 0.05 threshold, the study did not reject the null hypothesis of zero difference.
That requires a clinically justified margin chosen in advance and a confidence interval contained within that margin.
A non-significant p-value can occur because the true difference is small. It can also occur because the study is too small or too noisy to identify a difference. Here, no prespecified equivalence or non-inferiority margin and no confidence interval evaluated against such a margin were reported. The study therefore did not demonstrate equivalence or non-inferiority.
The peer-reviewed commentary raised a second statistical concern: the z-test treated repeated findings as independent even though observations were correlated by reader and case. A multireader, multicase analysis should account for those correlations. That issue makes the reported p-values less secure; it does not establish which observer is better.
Claim escalation analysis
Where the wording becomes stronger.
- 1 · Data90.2% versus 91.0% overall accuracy.
SignalRAY also had lower sensitivity and higher specificity than the best-accuracy radiologist.
- 2 · Statistical resultNo significant accuracy difference detected by the study's z-test.
The reported p-value was 0.0561. This is a failure to detect a difference, not a positive equivalence result.
- 3 · Authors' interpretation“Almost as well as the best” for descriptive findings.
The authors also wrote that the AI was “no less accurate,” while acknowledging lower abnormality sensitivity and no differential diagnoses.
- 4 · Vendor interpretationThe study “validates” expert-level interpretation.
SignalPET's research article presents the result as confirmation of reliable, specialist-comparable performance.
- 5 · Marketing claim“Equivalent Accuracy” and “matched the best radiologist.”
This wording converts an inconclusive difference test into an affirmative equivalence message.
Where the claim holds
The point estimates were close in this study.
- SignalRAY's 90.2% overall accuracy was close to the best radiologist's 91.0% within the study's constructed consensus framework.
- The AI's overall accuracy exceeded the median-accuracy radiologist's estimate under the reported analysis.
- The AI had high specificity for normal findings and performed particularly well in the study's low-ambiguity group.
- The authors' narrower formulation—almost as well for coded descriptive findings—is more consistent with the data than a broad claim of equivalent radiologist-level interpretation.
Where it stretches
“Matched” carries more certainty and scope than the evidence.
The analysis looked for a difference. It did not show that any plausible difference was small enough to be clinically unimportant.
Radiologists produced full reports with diagnostic reasoning and differential diagnoses. Evaluation reduced their work to coded normal/abnormal findings because SignalRAY did not generate differential diagnoses.
The reference standard included the AI under evaluation and did not independently resolve disagreements using outcomes, surgery or pathology.
The tested software was accessed in July 2022 and had no fixed version number. SignalPET's current Immediate Report advertises 63+ pathologies, while its current Complete Report adds clinical context, conclusions, differentials and recommendations—capabilities this study did not evaluate.
The experiment did not test changes in treatment, missed diagnoses in practice, clinician-AI interaction or clinical benefit.
Limitations
Promising comparison, narrow evidentiary base.
- Small, single-source dataset. Fifty studies from one institutional PACS cannot represent the full case mix, acquisition quality and disease spectrum of general practice.
- Class imbalance. Consensus labeled 83.9% of 16,434 findings normal and 16.1% abnormal. Calling every finding normal would yield about 84% accuracy, making sensitivity essential context.
- Non-independent reference standard. Observer consensus included SignalRAY, the system being evaluated.
- Coding assumptions. An unmentioned finding was treated as normal, and insignificant abnormalities in radiologist reports were coded abnormal. These decisions may affect human–AI comparisons.
- Reader and case clustering. The reported proportion z-tests did not model repeated findings within cases and repeated readings by radiologists.
- No equivalence framework. There was no prespecified margin, equivalence confidence interval or power calculation for proving a clinically acceptable difference.
- Evaluation-set independence unclear. The paper describes train, validation and test splits used in model development but does not establish whether these 50 evaluation studies were absent from every development dataset.
- Different exposure and tasks. The AI assessed all 50 studies using images alone; each radiologist assessed roughly half and received signalment, history and clinical signs.
- Version traceability. The continuously updated July 2022 system had no version number, limiting reproducibility and relevance to later releases.
Vendor involvement
SignalPET funded the study and supplied operational support.
The paper states that SignalPET supported radiologist reports and publication costs; provided PACS access; processed the cases with SignalRAY; tabulated coded findings; transferred radiologist findings into diagnostic codes matched to the software; and supplied technical information about SignalRAY.
The authors state that SignalPET had no role in study design, data collection, analysis, data interpretation, manuscript writing or the decision to publish. They declared no commercial or financial relationships that could be construed as a conflict of interest. Axel Ockenfels and Yero S. Ndiaye also acknowledged research funding through the German Research Foundation's Excellence Strategy.
Funding does not invalidate a study. It does make clear disclosure, reproducible methods and independent validation especially important when the result is later used in marketing.
The later peer-reviewed commentary reported no funding and no commercial or financial conflicts.
VetAI Trust assessment
The study supports proximity—not proven equivalence.
The narrower observation is supported: SignalRAY's reported accuracy point estimate was close to the best participating radiologist's—90.2% versus 91.0%—within this small, consensus-based evaluation.
The commercial phrase “Equivalent Accuracy” goes further and is not demonstrated by the study design. The study tested for a difference and reported p = 0.0561; it did not report a prespecified equivalence or non-inferiority margin or a confidence interval evaluated against such a margin. The peer-reviewed commentary additionally questions the z-test because observations were correlated by reader and case.
Our status is therefore partially supported — equivalence not demonstrated. This does not mean the claim is false, nor does it establish that the current product performs worse than a radiologist. It means the published evidence does not resolve that stronger question.
See how VetAI Trust maps claims to evidence. Material factual or methodological challenges are handled under the Corrections policy.
Right to respond
SignalPET is invited to provide additional evidence.
VetAI Trust welcomes factual corrections, a prespecified equivalence or non-inferiority analysis, independent external validation, version-specific performance data, or other evidence that materially changes this assessment.
Sources
- Commercial claim. SignalPET. New Research Validates SignalPet’s AI in Radiographic Interpretation. Published 24 March 2025; reviewed 18 August 2026.
- Current product context. SignalPET. Immediate Report — Veterinary AI X-Ray Screening. Reviewed 18 August 2026.
- Primary study. Ndiaye YS, Cramton P, Chernev C, Ockenfels A, Schwarz T. Comparison of radiological interpretation made by veterinary radiologists and state-of-the-art commercial AI software for canine and feline radiographic studies. Frontiers in Veterinary Science. 2025;12:1502790. doi:10.3389/fvets.2025.1502790. Full PDF.
- Peer-reviewed commentary. Joslyn SK, Faulkner J, Ma D, Appleby R. Commentary: Comparison of radiological interpretation made by veterinary radiologists and state-of-the-art commercial AI software for canine and feline radiographic studies. Frontiers in Veterinary Science. 2025;12:1615947. doi:10.3389/fvets.2025.1615947. Full PDF.
- General methodological background. Piaggio G, Elbourne DR, Pocock SJ, Evans SJW, Altman DG; CONSORT Group. Reporting of Noninferiority and Equivalence Randomized Trials: Extension of the CONSORT 2010 Statement. JAMA. 2012;308(24):2594–2604. doi:10.1001/jama.2012.87802. This randomized-trial reporting guidance is cited only as general background for the logic of formal equivalence and non-inferiority reasoning; it is not presented as a standard governing this veterinary multireader diagnostic-accuracy study.
Claim Check by VetAI Trust. Vendor statements, study findings, published commentary and VetAI Trust's editorial assessment are presented as distinct layers. No author reply or correction was linked from the Frontiers study record when reviewed on 18 August 2026.