A Disease-Memory Guided Framework for Radiology Report Generation from Multi-View Chest X-rays
Abstract
Chest X-rays are widely used, low-cost diagnostic tools. Writing chest X-ray reports manually is a time-consuming and tedious task. The increasing number of X-rays burdens radiologists, creating a need for a clinically accurate, fast and automated report generation system. To solve this challenge, we propose a new disease-memory-guided vision-language framework for radiology report generation. Frontal and lateral chest X-rays (multi-view) are encoded using a Swin-Transformer. These encoded visual features are then passed through a multi-label classification head that predicts disease probabilities, from which the top three labels are converted into learnable disease-memory embeddings. These are concatenated with projected visual tokens and fixed task instruction embeddings to form a multimodal prompt which is fed into a frozen decoder. A LLaMA-2-7B model adapted to the radiology domain using LoRA fine-tuning on the train-only MIMIC-CXR subset is used as an autoregressive decoder. The disease-memory embeddings serve as disease guidance during report generation. An extensive experimental analysis of different experimental variants is conducted on the benchmark Indiana University Chest X-ray dataset using three aspects: 1) statistical n-gram overlap, 2) semantic similarity, and 3) clinical label-based metrics. Experimental results of the proposed system show improvements in semantic evaluation metrics, with BERTScore precision, recall, and F1 values of 0.898, 0.891 and 0.894, respectively, and ROUGE-L and METEOR scores of 0.327 and 0.281, indicating improved contextual consistency between generated and reference reports. However, performance on statistical metrics (e.g., BLEU4=0.068) and clinical evaluation metrics remains limited, with Micro-F1, Macro-F1, Precision and Recall values of 0.4211, 0.0640, 0.4272, and 0.4151, respectively, indicating only partial abnormality description. This framework improves abnormality grounding while maintaining computational efficiency and can be used as a supportive tool for preliminary report drafting in research or resource-constrained settings.
Downloads
References
Z. He, A. N. N. Wong, and J. S. Yoo, ‘Radiology report generation using automatic keyword adaptation, frequency-based multi-label classification and text-to-text large language models’, Computers in Biology and Medicine, vol. 196, p. 110625, Sep. 2025, doi: 10.1016/j.compbiomed.2025.110625.
F. Dong, S. Nie, M. Chen, F. Xu, and Q. Li, ‘Keyword-based AI assistance in the generation of radiology reports: A pilot study’, npj Digit. Med., vol. 8, no. 1, p. 490, Aug. 2025, doi: 10.1038/s41746-025-01889-4.
X. Wang, Y. Zhang, Z. Guo, and J. Li, ‘TMRGM: A Template-Based Multi-Attention Model for X-Ray Imaging Report Generation’, JoAIMS, vol. 2, no. 1–2, pp. 21–32, 2021, doi: 10.2991/jaims.d.210428.002.
F. Yu et al., ‘Evaluating progress in automatic chest X-ray radiology report generation’, Patterns, vol. 4, no. 9, p. 100802, Sep. 2023, doi: 10.1016/j.patter.2023.100802.
K. Kale, P. Bhattacharyya, and K. Jadhav, ‘Replace and Report: NLP Assisted Radiology Report Generation’, in Findings of the Association for Computational Linguistics: ACL 2023, Toronto, Canada: Association for Computational Linguistics, 2023, pp. 10731–10742. doi: 10.18653/v1/2023.findings-acl.683.
O. Alfarghaly, R. Khaled, A. Elkorany, M. Helal, and A. Fahmy, ‘Automated radiology report generation using conditioned transformers’, Informatics in Medicine Unlocked, vol. 24, p. 100557, 2021, doi: 10.1016/j.imu.2021.100557.
D. Gao et al., ‘Simulating doctors’ thinking logic for chest X-ray report generation via Transformer-based Semantic Query learning’, Medical Image Analysis, vol. 91, p. 102982, Jan. 2024, doi: 10.1016/j.media.2023.102982.
Z. Liu et al., ‘Swin Transformer: Hierarchical Vision Transformer using Shifted Windows’, in 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada: IEEE, Oct. 2021, pp. 9992–10002. doi: 10.1109/ICCV48922.2021.00986.
Y. Tang, Y. Yuan, F. Tao, and M. Tang, ‘Cross-Modal Augmented Transformer for Automated Medical Report Generation’, IEEE J. Transl. Eng. Health Med., vol. 13, pp. 33–48, 2025, doi: 10.1109/JTEHM.2025.3536441.
S. Yang, X. Wu, S. Ge, Z. Zheng, S. K. Zhou, and L. Xiao, ‘Radiology report generation with a learned knowledge base and multi-modal alignment’, Medical Image Analysis, vol. 86, p. 102798, May 2023, doi: 10.1016/j.media.2023.102798.
Z. Chen, Y. Shen, Y. Song, and X. Wan, ‘Cross-modal Memory Networks for Radiology Report Generation’, in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), Online: Association for Computational Linguistics, 2021, pp. 5904–5914. doi: 10.18653/v1/2021.acl-long.459.
Z. Chen, Y. Song, T.-H. Chang, and X. Wan, ‘Generating Radiology Reports via Memory-driven Transformer’, in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Online: Association for Computational Linguistics, 2020, pp. 1439–1449. doi: 10.18653/v1/2020.emnlp-main.112.
M. Y. Ouis and M. A. Akhloufi, ‘ChestBioX-Gen: contextual biomedical report generation from chest X-ray images using BioGPT and co-attention mechanism’, Front. Imaging, vol. 3, p. 1373420, Apr. 2024, doi: 10.3389/fimag.2024.1373420.
J. Ji, Y. Hou, X. Chen, Y. Pan, and Y. Xiang, ‘Vision-Language Model for Generating Textual Descriptions From Clinical Images: Model Development and Validation Study’, JMIR Form Res, vol. 8, p. e32690, Feb. 2024, doi: 10.2196/32690.
A. Selivanov, O. Y. Rogov, D. Chesakov, A. Shelmanov, I. Fedulova, and D. V. Dylov, ‘Medical image captioning via generative pretrained transformers’, Sci Rep, vol. 13, no. 1, p. 4171, Mar. 2023, doi: 10.1038/s41598-023-31223-5.
Z. Wang, L. Liu, L. Wang, and L. Zhou, ‘R2GenGPT: Radiology Report Generation with frozen LLMs’, Meta-Radiology, vol. 1, no. 3, p. 100033, Nov. 2023, doi: 10.1016/j.metrad.2023.100033.
M. Kapadnis, S. Patnaik, A. Nandy, S. Ray, P. Goyal, and D. Sheet, ‘SERPENT-VLM : Self-Refining Radiology Report Generation Using Vision Language Models’, in Proceedings of the 6th Clinical Natural Language Processing Workshop, Mexico City, Mexico: Association for Computational Linguistics, 2024, pp. 283–291. doi: 10.18653/v1/2024.clinicalnlp-1.24.
C.-Y. Lin, ‘ROUGE: A Package for Automatic Evaluation of Summaries’, in Text Summarization Branches Out, Barcelona, Spain: Association for Computational Linguistics, Jul. 2004, pp. 74–81. Accessed: Sep. 01, 2026. [Online]. Available: https://aclanthology.org/W04-1013/
R. Vedantam, C. L. Zitnick, and D. Parikh, ‘CIDEr: Consensus-based image description evaluation’, in 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA: IEEE, Jun. 2015, pp. 4566–4575. doi: 10.1109/CVPR.2015.7299087.
P. Singh and S. Singh, ‘ChestX-Transcribe: a multimodal transformer for automated radiology report generation from chest x-rays’, Front. Digit. Health, vol. 7, p. 1535168, Jan. 2025, doi: 10.3389/fdgth.2025.1535168.
A. Nicolson, J. Dowling, and B. Koopman, ‘Improving chest X-ray report generation by leveraging warm starting’, Artificial Intelligence in Medicine, vol. 144, p. 102633, Oct. 2023, doi: 10.1016/j.artmed.2023.102633.
Z. Babar, T. Van Laarhoven, and E. Marchiori, ‘Encoder-decoder models for chest X-ray report generation perform no better than unconditioned baselines’, PLoS ONE, vol. 16, no. 11, p. e0259639, Nov. 2021, doi: 10.1371/journal.pone.0259639.
Y. Pan, L.-J. Liu, X.-B. Yang, W. Peng, and Q.-S. Huang, ‘Chest radiology report generation based on cross-modal multi-scale feature fusion’, Journal of Radiation Research and Applied Sciences, vol. 17, no. 1, p. 100823, Mar. 2024, doi: 10.1016/j.jrras.2024.100823.
S. Zhang, C. Zhou, L. Chen, Z. Li, Y. Gao, and Y. Chen, ‘Visual prior-based cross-modal alignment network for radiology report generation’, Computers in Biology and Medicine, vol. 166, p. 107522, Nov. 2023, doi: 10.1016/j.compbiomed.2023.107522.
J. Zhao et al., ‘Automated Chest X-Ray Diagnosis Report Generation with Cross-Attention Mechanism’, Applied Sciences, vol. 15, no. 1, p. 343, Jan. 2025, doi: 10.3390/app15010343.
H. Yin, S. Zhou, P. Wang, Z. Wu, and Y. Hao, ‘KIA: Knowledge-Guided Implicit Vision-Language Alignment for Chest X-Ray Report Generation’, in Proceedings of the 31st International Conference on Computational Linguistics, O. Rambow, L. Wanner, M. Apidianaki, H. Al-Khalifa, B. D. Eugenio, and S. Schockaert, Eds, Abu Dhabi, UAE: Association for Computational Linguistics, Jan. 2025, pp. 4096–4108. Accessed: Sep. 01, 2026. [Online]. Available: https://aclanthology.org/2025.coling-main.276/
D. Gu, Y. Gao, Y. Zhou, M. Zhou, and D. Metaxas, ‘RadAlign: Advancing Radiology Report Generation with Vision-Language Concept Alignment’, in Medical Image Computing and Computer Assisted Intervention – MICCAI 2025, vol. 15966, J. C. Gee, D. C. Alexander, J. Hong, J. E. Iglesias, C. H. Sudre, A. Venkataraman, P. Golland, J. H. Kim, and J. Park, Eds, Cham: Springer Nature Switzerland, 2026, pp. 484–494. doi: 10.1007/978-3-032-04981-0_46.
T. Gu, D. Liu, Z. Li, and W. Cai, ‘Complex Organ Mask Guided Radiology Report Generation’, in 2024 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA: IEEE, Jan. 2024, pp. 7980–7989. doi: 10.1109/WACV57701.2024.00781.
T. Tanida, P. Müller, G. Kaissis, and D. Rueckert, ‘Interactive and Explainable Region-guided Radiology Report Generation’, in 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada: IEEE, Jun. 2023, pp. 7433–7442. doi: 10.1109/CVPR52729.2023.00718.
H. Sharma et al., ‘MAIRA-Seg: Enhancing Radiology Report Generation with Segmentation-Aware Multimodal Large Language Models’, in Proceedings of the 4th Machine Learning for Health Symposium, PMLR, Feb. 2025, pp. 941–960. Accessed: Sep. 01, 2026. [Online]. Available: https://proceedings.mlr.press/v259/sharma25a.html
S. Lee, J. Youn, H. Kim, M. Kim, and S. H. Yoon, ‘CXR-LLaVA: a multimodal large language model for interpreting chest X-ray images’, Eur Radiol, vol. 35, no. 7, pp. 4374–4386, Jan. 2025, doi: 10.1007/s00330-024-11339-6.
J. S. Ryu, H. Kang, Y. Chu, and S. Yang, ‘Vision-language foundation models for medical imaging: a review of current practices and innovations’, Biomed. Eng. Lett., vol. 15, no. 5, pp. 809–830, Sep. 2025, doi: 10.1007/s13534-025-00484-6.
I. Hartsock and G. Rasool, ‘Vision-language models for medical report generation and visual question answering: a review’, Front. Artif. Intell., vol. 7, p. 1430984, Nov. 2024, doi: 10.3389/frai.2024.1430984.
R. Tanno et al., ‘Collaboration between clinicians and vision–language models in radiology report generation’, Nat Med, vol. 31, no. 2, pp. 599–608, Feb. 2025, doi: 10.1038/s41591-024-03302-1.
J. Xu, ‘Uncertainty Estimation in Large Vision Language Models for Automated Radiology Report Generation’, in Proceedings of the 4th Machine Learning for Health Symposium, PMLR, Feb. 2025, pp. 1039–1052. Accessed: Sep. 01, 2026. [Online]. Available: https://proceedings.mlr.press/v259/xu25a.html
Y. Zhang, H. Jiang, Y. Miura, C. D. Manning, and C. P. Langlotz, ‘Contrastive Learning of Medical Visual Representations from Paired Images and Text’, 2020, Proc. Mach. Learn. Res., vol. 182, pp. 2–25. doi: 10.48550/ARXIV.2010.00747.
D. Demner-Fushman et al., ‘Preparing a collection of radiology examinations for distribution and retrieval’, Journal of the American Medical Informatics Association, vol. 23, no. 2, pp. 304–310, Mar. 2016, doi: 10.1093/jamia/ocv080.
A. E. W. Johnson et al., ‘MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports’, Sci Data, vol. 6, no. 1, p. 317, Dec. 2019, doi: 10.1038/s41597-019-0322-0.
W. Li, X. Wang, W. Li, and B. Jin, ‘A Survey of Automatic Prompt Engineering: An Optimization Perspective’, 2025, arXiv. doi: 10.48550/ARXIV.2502.11560.
T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi, ‘BERTScore: Evaluating Text Generation with BERT’, in Proc. Int. Conf. Learn. Represent. (ICLR), 2020. doi: 10.48550/ARXIV.1904.09675.
Copyright (c) 2026 Prajakta Dhamanskar, Chintan Thacker, Saurabh Shah

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution-ShareAlikel 4.0 International (CC BY-SA 4.0) that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).


.png)
.png)
.png)
.png)
.png)