Cognita Imaging, a subsidiary of Mosaic Clinical Technologies, has been awarded a $1.2 million contract by the U.S. Food and Drug Administration (FDA) to test a new approach for evaluating AI-generated radiology reports.
The project will see the use of multiple large language models (LLMs) to “work as a jury” in assessing AI-generated reports across around a million patient exams, according to Mosaic. Cognita’s role will be the development and validation of a framework incorporating several LLMs to check human and AI-generated radiology reports, comparing the judgement of the models to produce a more consistent evaluation.
Once the framework has been tested, it will be used against the data from patient exams, with researchers looking at performance across different patient groups and care settings. Radiologists will then review “clinically important” disagreements to find out whether the issue stemmed from the AI-generated report, the LLMs, or the original radiologist report, Mosaic notes. These findings will reportedly go toward revising the framework and exploring the role of automated review in premarket validation and post-market monitoring.
Akshay Chaudhari, Cognita co-founder and principal investigator on the contract, said: “Once an AI system starts writing the entire report, the evaluation problem changes. A few hundred cases may tell you whether a model works in a narrow setting. They do not show every way it can fail in practice. We want to test whether a group of language models, with radiologists reviewing the difficult cases, can give us a clearer and more repeatable picture at real-world scale.”
Part of Cognita’s role will be supporting the FDA with software code and guidance for building LLM juries, Mosaic continues. Taking to LinkedIn, the company shared: “As AI takes on a greater role in producing radiology, evaluating these systems for safety and performance is increasingly important. Through this research, we will test how LLMs combined with radiologist expertise could help advance the way generative AI clinical devices are evaluated, aiming to ensure these systems can be safely reviewed at scale.”
Wider trend: Health AI
The National Commission into the Regulation of AI has put forward a series of recommendations for a future regulatory framework, finding strong support for the use of AI in healthcare being “conditional, rather than automatic”. The central conclusion, the commission goes on, is that future regulation needs to be more proportionate, lifecycle-based, and system-wide. “Current approaches were largely designed for products that are more static and easier to reliably assess at a single point in time,” it states. “AI-enabled products may iterate rapidly, perform differently in different settings and depend on the data, workflows, people and organisations around them.”
The government has launched the first procurement competitions for its £100 million sovereign AI R&D procurement scheme, aiming to support British AI companies to “start here, scale here and win globally”. According to the government, the scheme is designed to act as a venture capital fund, focusing on demonstrator-stage solutions ready to be tested in real-world environments. Applications will be assessed by the sovereign AI team, participating government departments, and technical experts, with the first competitions designed to test the delivery model for future phases.
An AI spin-out from the University of Birmingham has launched, with an aim to speed up the process of extracting data from EHRs to support medical researchers and reduce the time for drug safety and effectiveness studies. Dexter AI is based on a software developed by a team from the Birmingham Department of Applied Health Sciences, Professor Krishnarajah Nirantharakumar, Dr Krishna Gokhale, and Professor Joht Singh Chandan. Its website outlines the potential for use in the design of epidemiological studies, cohort identification, emulating clinical trials with real world data, automating clinical audits, and enabling population health management.


