Large language models are entering nuclear cardiology at a pace that makes the usual question, „Can AI support physicians?”, increasingly outdated. The more relevant question is now operational: how should these systems be governed, validated and integrated into clinical imaging workflows without compromising medical accountability?
In the first edition of Axcellant’s „3 Questions with…” series, Yana Arlouskaya, Head of Operations & Regulatory at Axcellant, spoke with Prof. Piotr Słomka, Professor at Cedars-Sinai Medical Center in Los Angeles, a recognised expert in artificial intelligence, quantitative imaging and nuclear cardiology, and a member of Axcellant’s Advisory Board.
Over the past two decades, Prof. Słomka has contributed to the development of advanced software and automated imaging analysis methods for cardiac diagnostics and clinical research.
The conversation focused on three areas that are becoming increasingly important for sponsors, imaging teams and clinical research organisations: the rapid evolution of LLM capabilities, their role in diagnostic imaging, and the strategic choice between broad general models and narrow specialised systems.
Yana Arlouskaya: How rapidly are large language models advancing in terms of their application in nuclear cardiology? A study you co-authored, first posted as a preprint in July 2024 and published in 2025, shows that the accuracy of decisions made by the latest models at the time, such as GPT-4, GPT-4o, and Gemini, ranged from about 40% to over 60%. Does each new generation (we are now at GPT-5.5 and Claude Opus 4.7 at the time of your study, and newer ones since) also represent a qualitative leap in the field of nuclear cardiology?
Prof. Piotr Słomka: Absolutely, and the timing of this question is very relevant. In our recent work, we tested the two leading large language models available at the time of the study, GPT-5.5 and Claude Opus 4.7, on the 168-question ASNC Board Preparation Exam. The results were genuinely surprising. Claude Opus 4.7 scored 86.3% and GPT-5.5 scored 86.7%, compared with 78% for the average fellow in training.
That is a significant shift. Compared with our 2024 results, that is a 23-percentage-point improvement in less than two years. Both models also answered the image-based questions using retrieval-augmented generation to draw on relevant textbooks and guidelines, although image interpretation remains their weakest area: 73% for Claude and 77% for GPT-5.5, still below the 82% achieved by the fellows.
Newer models have already been released since we ran these experiments, which is itself the point: any specific benchmark figure has a short shelf life, and the trend line matters more than the individual number.
From my perspective, this moves the discussion beyond whether LLMs can be useful in nuclear cardiology. They clearly can. The more important question is how we integrate them responsibly into clinical, educational and research workflows.
There is also a major productivity aspect. This work has now been accepted for publication in the Journal of Nuclear Cardiology. We used Claude to produce the first draft, including figures and reference formatting, as a deliberate test of AI-assisted scientific writing — which is disclosed in the paper itself. The research team reviewed and approved every result before submission, but the first draft came from the model, under human direction at every step.
The efficiency was remarkable. We completed five full iterations of the analysis, with all figures updated, in a single working session. But there is a critical caveat. These models are extremely confident, and if something is missing, they may fill the gap themselves. In one case, a supplement with the fellows’ performance data was missing from the input, and the model filled the gap with invented numbers — which our verification step caught before submission.
That is why human supervision is non-negotiable. LLMs can accelerate work dramatically, but acceleration without expert oversight is not progress. It is risk.
Yana Arlouskaya: In your opinion, will LLMs become agents that independently perform steps in diagnostic imaging in the near future, rather than just suggesting solutions? How much will they accelerate and improve the diagnostic process for the benefit of the patients (e.g., avoiding serious complications as a result of timely treatment)? Please explain your point.
Prof. Piotr Słomka: They are getting there technically, but I think independence is the wrong frame. Recently published work has demonstrated systems in which LLMs generate full interactive diagnostic reports from scans, clinical history and laboratory data. The capability is real. We are no longer talking only about summarising text or simplifying reports for patients, although those applications are already being used in clinical settings.
The step towards full diagnostic reporting has been technically demonstrated. What remains is a regulatory, clinical and operational question.
However, the final decision about a patient must remain with the physician. That will not change. These systems should support expert decision-making, not replace it.
Where I see particularly strong potential is in incidental findings. A cardiologist looking at a scan is naturally focused on the heart. That is appropriate, but it also means that something outside the main clinical question, such as an enlarged aorta visible on the same scan, may be missed.
A model has the time, breadth and consistency to flag that type of finding. The physician then decides what to do with the information.
This is not only about speed. In many cases, the real value will be accuracy, completeness and the ability to expand the physician’s field of attention.
One practical issue we observed in our study was also very interesting. GPT-5.5 refused to answer an average of 12 of the 168 questions per run, around 7%, because its safety filters flagged legitimate exam content on radiopharmaceutical handling, radiation safety, patient management and stress pharmacology. Claude Opus 4.7 answered every question across all five runs. In a medical context, that difference matters enormously.
It shows that performance is not the only factor. Model behaviour, restrictions, reliability and context-awareness are just as important when we consider clinical use.
Yana Arlouskaya: SLMs narrowly and locally trained on specific guidelines or a universal LLM with a vast database – which do you think has greater potential for use in nuclear cardiology?
Prof. Piotr Słomka: My personal view is clear: nuclear cardiology should bet on large general models. We have spent years building narrow, single-task systems. For example, a model trained to detect coronary disease from one specific type of scan may perform well within that limited task. But show it anything outside that task, and its performance degrades quickly.
That is the core limitation of specialised models. They often do not generalise. A model fine-tuned at one hospital may fail at another. A system trained for one workflow may not work in a different clinical environment. Some narrow models barely understand broader clinical context.
Large language models bring something fundamentally different. They combine broad knowledge, strong general reasoning, natural language capabilities and the ability to handle unexpected inputs. They can work with clinical context, imaging-related information, guidelines, reports and language at the same time.
Of course, they are not perfect. They hallucinate. They are more expensive. They need constraints. But the right strategy is not to avoid them. The right strategy is to apply strong guardrails that keep them focused, traceable and accountable.
For nuclear cardiology and imaging-driven clinical research, this is especially important. These environments are complex, variable and often multicentre by design. A narrow model may work in one setting but fail when the input changes. A large model, properly constrained, has a better chance of adapting to that complexity.
So the direction is clear to me: use large models, but constrain them carefully.
That is where nuclear cardiology is heading. We have already presented this work at SNMMI in Los Angeles in June, including an abstract on an LLM-powered interactive reporting system.
What this means for clinical research
The evolution of LLMs in nuclear cardiology points towards a broader shift. AI is becoming less of a standalone innovation topic and more of a potential infrastructure layer for clinical research. For sponsors, biotech companies and CROs operating in complex imaging environments, the implications are significant. LLMs may support literature review, research documentation, imaging workflow design, report structuring and data interpretation. They may also help identify missing information, surface inconsistencies and improve the completeness of imaging review.
None of the systems discussed here are validated or regulatory-cleared for clinical decision-making; they are research and educational tools at this stage. Their introduction into clinical research must therefore be deliberate. AI-enabled workflows require clear governance: who reviews outputs, how sources are controlled, what data enters the model, how hallucinations are detected, and which decisions remain exclusively human.
This is particularly important in nuclear medicine, radiopharmaceutical development and multicentre imaging studies, where scientific complexity, operational variability and regulatory scrutiny already intersect. The opportunity is substantial, but the standard must be high. In medical research, the goal is not to use AI because it is impressive. The goal is to use it where it improves quality, speed, consistency or decision support without weakening accountability.
At Axcellant, this is where the conversation becomes especially relevant. Advanced imaging, AI-supported analysis and robust clinical operations are increasingly interconnected. The challenge is not only to follow innovation, but to translate it into reliable, compliant and operationally useful research practice.
The future of LLMs in nuclear cardiology will not be defined by autonomy. It will be defined by controlled intelligence: powerful models, expert oversight and clinical accountability working together.
Reference
About the authors:

An experienced manager of complex, multi-country studies, Yana has delivered 30+ Phase I–IV trials across APAC, the US, and Europe—covering IMPs, medical devices, and ATMPs. A specialist in global regulatory oversight (FDA, EMA, PMDA), she drives end-to-end operations—from protocol design and CTIS submissions to risk mitigation—ensuring top-tier compliance (ICH GCP, ISO 14155) and moving every project toward successful approval.

As an Advisory Board Member, Professor Piotr J. Słomka supports Axcellant in the strategic design and optimisation of imaging-based clinical studies, with a particular focus on quantitative analysis and artificial intelligence applications. Leveraging his combined expertise in AI-driven image analysis, nuclear imaging and cardiology he advises on the development and validation of automated, reproducible imaging endpoints suitable for multicentre and regulatory-grade trials.
How Axcellant helped a diagnostic study overcome isotope supply, manufacturing, and logistics constraints In radiopharmaceutical development, we often say that…
Europe’s clinical trials market is entering a more practical phase in its approach to competitiveness. With the Clinical Trials Regulation…
The Axcellant team recently attended the International Conference on Nuclear Cardiology (ICNC) in Berlin – one of the most important…
Copyright @ 2026 Axcellant