How trustworthy is AI for clinical data summary: A case study
23rd July 2026 by Sharaz Luke
Introduction
According to an American Medical Association survey, over 80% of clinicians use AI as part of their everyday practice (1).
Within this group, 39% specifically use AI to summarise medical research and standards of care, and its use could help to reduce work-related burnout.
This raises a question on how reliable are AI tools to carry tasks that could shape everyday medical practices?
A change in healthcare data consumption
Clinicians increasingly turn to a browser or AI assistant when a new trial reads out a major congress, often before the full publication is even available. This shift represents a great concern, because it can change where the message first lands, and whether the information clinicians act on matches what the sponsor intended to say.
To investigate the accuracy of AI summaries, we tested this against a real, recent Phase 3 trial example, presented at a major congress. The details in this article have been anonymised, therefore, the drug will be referred as “SummariZol” and the trial as “SNAIL-Z”, but the findings below reflect what was actually observed.
1. Implementing a process for comparisons
To see how these tools handled the data, we used the sponsor’s official press release and an independent oncology article as “control” references, setting out the facts an accurate summary should contain, and roughly the order they should appear in.
The same query, the results of the SNAIL-Z, was then submitted across:
- Edge & Chrome browsers,
- Widely used AI assistants (ChatGPT, Grok, and Claude).
Each response was assessed against nine facts a clinician would reasonably need, each weighted by clinical importance.
Five were rated high importance:
- The headline efficacy reading,
- The study design,
- Median event-free survival,
- 24-month event-free survival,
- Whether the trial met its primary endpoint.
Three were rated medium importance:
- The patient population,
- The use of independent or peer-reviewed sourcing,
- Explicit reference to results being "statistically significant" or "clinically meaningful."
One was rated low importance:
- The mention of brand name.
2. Results
The outputs from the five AI platforms were summarised into a single Red, Amber, Green (RAG) dashboard with an overall assessment selected by accuracy on weighted responses.
According to our criteria, and in reference to our controls, 2/5 of the AI platforms passed - ChatGPT and Grok - though ChatGPT was a borderline pass.
Our detailed report goes into the details for each criteria, however in this summary case study we have only outlined some of the key observations from this event.
2.1 The most concerning answer came from Claude
Two of the five AI assistants passed, but this was not without caveats. Edge summary omitted the risk-reduction figure and hazard ratio entirely.
- However, the finding that stood out most was how Claude opened its response to the question of median event-free survival (EFS). Although it included the data, its opening sentence was: "the median EFS was not reached."
- While this phrasing is technically accurate and the actual 24-month EFS was positive, the failure with the median EFS can be misinterpreted.
A hurried healthcare professional scanning the summary, or a non-native English speaker, could easily reached this to mean the trial failed to meet its primary endpoint rather than recognising it as a favourable result.
Therefore, this statement is labelled as high-level of concerns as it carries the risk of miscommunication to a global audience.
- ChatGPT didn't mention the 24-month EFS finding at all,
- Chrome did, but in different terminology to the reference control (2-year EFS mentioned which technically should be 24-month EFS).
2.2 AI browsers struggle with trial fundamentals
Although Chat GPT and Grok stated the trial met its primary endpoint, Edge and Chrome did not clearly mention the trial met its primary endpoint.
- Our control references both open with a positive result stated plainly,
- Edge, by contrast, saying the treatment “significantly improved EFS”,
- Chrome says it “significantly improved outcomes”.
Both statements go on to suggest the data supports a new standard of care. Although the primary endpoint being met could be implied from these statements, it is not factually confirmed, which led to a fail.
2.3 The patient population wasn't explicitly clear with Edge and ChatGPT
There were also nuances in the patient population.
- Edge and ChatGPT describing it as stage IB-IIIA,
- However the trial’s primary analysis population was actually Stage II-IIIA,
- ChatGPT was the only tool to explain both, i.e. Stage IB patients were part of the trial but excluded from the primary analysis. This could be due to its citations leaned more heavily on the sponsor's own published material than on peer-reviewed or independent coverage.
Being able to get the right population and the trial nuances matter in practice, as it determines which real patients the data is relevant to.
3. Conclusions
Overall, across five widely used AI tools, accuracy on this trial readout was inconsistent throughout the queries.
These findings indicate AI-generated summaries cannot yet be treated as a reliable source for primary trial data.
As reliance on these tools grows, even small inaccuracies in framing, terminology or completeness carry meaningful downstream risk for clinical understanding.
AI answers are not static, meaning repeating one browser query two hours later resulted in a different and more accurate response. It was now showing the brand name and correct hazard ratio, alongside a jump of 1,300 queries on that same overview.
The case study highlighted the potential for miscommunication with experts, even in a therapy area widely trialled and at a major oncology congress.
Our recommendations is that pharma teams with new data need to monitor and react to any miscommunication errors.
A clear dashboard view of the information shaping customer conversations enables medical and commercial teams to refine future content and optimise messaging for generative engines.
For more information about our AI-generated summary service, TRUST, email support@meddigital.com