Cognita Imaging recently received a $1.29 million research contract from the Food and Drug Administration to test a new method of evaluating artificial intelligence-generated radiology reports. The goal is to assess whether a “jury” of large language models can evaluate another AI model’s report, and which cases need radiologist oversight.
The contract, which started on June 22 and will run for 18 months, is intended to address a core problem with generative AI. Currently, most tools are evaluated against a panel of radiologist experts, but for models that can evaluate a wide variety of cases and generate different outputs, this becomes more challenging and time-consuming.
Today, most of the AI tools authorized by the FDA are in radiology. Many of them are one-off solutions, such as to detect certain blood clots in a head CT scan, or a collapsed lung in a chest X-ray. Evaluating these types of models is straightforward, said Akshay Chaudhari, Cognita co-founder and an associate professor of radiology and biomedical data science at Stanford University. However, tools that only identify one or two things are not very helpful in radiologist practices.
Newer tools being developed by Cognita and its peers look to generate full text radiology reports.
“But it’s more open-ended in nature, so there could be hundreds of findings that could potentially be represented in this output,” Chaudhari said. “So, how do you meaningfully evaluate these hundreds of findings?”
Cognita is working with the Center for Devices and Radiological Health’s regulatory science team to test the approach of using multiple LLMs to review the output of AI radiology models, with the goal of being comfortable with their understanding of the benefits and risks of those tools. The company will build and test the framework, and then apply it to 1 million patient exams from a large U.S. cohort. Researchers will examine the model’s performance across different patient groups, care settings, imaging equipment and diseases, including rare findings.
When there are disagreements that have clinical significance, radiologists will be brought in to see if the problem came from the AI-generated report, the LLM jury or the original radiologist’s report.
Two concerns with generative AI are the potential for hallucinations, or information that seems plausible but is incorrect, and omissions, when a model does not include something that should be there. Chaudhari said he intends to be able to quantify the rate of hallucinations and omissions and whether they were clinically meaningful through this approach.
Cognita will provide the FDA with a report on the project’s findings and limitations, as well as the software code, guidance for building LLM juries, analyses comparing large and small validation cohorts, and radiologist-reviewed discrepancy studies.
Chaudhari said he wants the learnings from the test to be generalizable, so that the FDA may use them in regulations and guidance.
Founded in 2024, Cognita was sold a year later to Radiology Partners, the largest radiology practice in the U.S. Chaudhari said having feedback from those radiologists about what challenges they’re facing, across diverse practice areas from an academic medical center to an outpatient community facility, has been helpful.
Currently, the U.S. faces a shortfall of radiologists, resulting in longer work hours and larger case backlogs.
Breakthrough device designation
Cognita is working with the FDA on a separate initiative to develop a vision-language model that can interpret chest X-rays and generate reports, which are then reviewed by a radiologist. The company received the FDA’s breakthrough designation earlier this year for the device, which has not been authorized by the agency yet.
Chaudhari said the breakthrough designation is separate from the company’s research contract. Cognita is working with the FDA to determine what type of clinical study the company could use to evaluate the accuracy, benefits and risks of its vision-language model.
“There’s still work to be done in developing the exact study plan and actually implementing the study, but we’re confident that through these very nice partnerships that the FDA fosters that there is a path forward, because there is a real clinical need for these tools,” Chaudhari said.
Between the breakthrough designations, the research contracts and a discussion paper shared by the agency last month, the FDA is taking a closer look at how to regulate devices that use generative AI.

