Radiologists routinely compare current and prior chest X-rays to track disease progression, producing follow-up reports that describe multiple findings, each localised to an anatomical region and annotated with a temporal change status. Existing automated methods either generate reports from a single image without modelling temporal context, or incorporate temporal information but do not ground their outputs spatially. The few approaches that combine temporal reasoning with spatial grounding are restricted to single-finding descriptions, leaving multi-finding reports with mixed change directions unaddressed. We present GRCD, a framework for grounded report generation from chest X-ray pairs in the multi-finding setting. We first construct a rigorously cleaned dataset of temporal chest X-ray pairs by identifying and correcting two systematic labelling errors in the source annotations. We then introduce a Region-Guided Change Token module that encodes per-region temporal change across anatomical structures and injects this signal into a language model through a dual-pathway strategy combining prepended spatial tokens with gated cross-attention. On a multi-finding test set, GRCD outperforms existing baselines on text generation and clinical accuracy metrics, with gains in change detection.
GRCD addresses the clinically realistic multi-finding setting that prior work leaves unaddressed.
First model to generate temporal reports with multiple co-occurring findings, each with its own change label and bounding box grounding.
Three-tier pipeline that identifies and corrects two systematic errors in Chest ImaGenome: same-study AP retakes and silent NLP absence. Yields a 40,250-pair benchmark.
Region-Guided Change Tokens encode per-region temporal change and reach the LLM via both prepended tokens and gated cross-attention for complementary spatial signals.
GRCD combines a frozen BioViL-T temporal encoder with a Region-Guided Change Token module and an autoregressive LLM decoder.
The RGCT module crops 25 anatomical regions from both feature maps, compares them across time, and produces 25 change tokens. Six key tokens are prepended to the LLM input (Pathway A), while all 25 are injected via gated cross-attention (Pathway B).
We identify and fix two systematic errors in standard Chest ImaGenome temporal pair construction.
14% of pairs come from the same exam (duplicate AP exposures). Both images share one report, so no temporal information can be learned. Removed by checking study ID equality.
77% of NEW/RESOLVED transitions lack textual evidence. NLP-derived labels fabricate change when findings are simply omitted. Three-tier confidence system retains only verified labels.
Evaluated on 7,967 multi-finding test pairs. GRCD values are mean ± std over 3 seeds.
| Metric | BioViL-T Clf | Retrieval | TRACE | GRCD (Ours) |
|---|---|---|---|---|
| Natural Language Generation | ||||
| BLEU-4 | — | 0.273 | 0.078 | 0.398 ± 0.003 |
| METEOR | — | 0.374 | 0.213 | 0.494 ± 0.003 |
| ROUGE-L | — | 0.376 | 0.176 | 0.493 ± 0.003 |
| Clinical Accuracy | ||||
| CheXbert F1 | — | 0.368 | 0.175 | 0.567 ± 0.009 |
| RadGraph F1 | — | 0.353 | 0.222 | 0.473 ± 0.002 |
| Grounding | ||||
| Mean IoU | — | — | 0.545 | 0.540 ± 0.005 |
| IoU > 0.5 | — | — | 64.0% | 61.4 ± 0.5% |
| Change Detection | ||||
| Report-level | 37.6% | — | 39.8% | 53.3 ± 0.2% |
| Balanced Acc. | 46.6% | — | 35.7% | 50.0 ± 0.3% |
Qualitative examples. Each row: the prior CXR, the current CXR with ground-truth boxes, and with GRCD's predicted boxes. Box colours mark change type (worsening / improving / stable); green and red text highlight matched and mismatched report sentences.
Each component contributes to the final performance. RGCT is essential; dual pathway outperforms either strategy alone.
| Configuration | BLEU-4 | METEOR | CheXbert F1 | IoU > 0.5 | Change Det. |
|---|---|---|---|---|---|
| A: No RGCT | 0.007 | 0.067 | 0.062 | — | 37.6% |
| B: 6-region, prepend only | 0.362 | 0.467 | 0.570 | 64.2% | 51.6% |
| C: 25-region, cross-attention only | 0.394 | 0.476 | 0.512 | 57.9% | 54.4% |
| D: 25-region, dual pathway (GRCD) | 0.402 | 0.498 | 0.576 | 61.8% | 53.6% |