Structured Nursing Handover Report Generation from Clinical Speech using Fine-Tuned XLSR-53 and T5: A Benchmarking Study

Keywords: Automatic speech recognition; XLSR-53; clinical NLP; nursing handover; T5 transformer; partial layer freezing; medical speech recognition; patient safety

Abstract

Accurate nursing handovers are critical for patient safety, as miscommunication during shift transitions leads to irreversible clinical errors. This work proposed an end-to-end pipeline that converts unstructured clinical nursing speech into standardized handover reports using a fine-tuned XLSR-53 acoustic model and T5-base text-to-text transformer. An Australian English clinical corpus of 200 synthetic nursing handover recordings from the CSIRO data access portal was utilised in this work. This benchmarking study was conducted within the CSIRO synthetic Australian English nursing handover corpus and does not represent a broad cross-domain clinical ASR benchmark. A domain-specific benchmarking study across seven state-of-the-art ASR architectures (Whisper Tiny/Base/Small, Wav2Vec2 Base/Large, HuBERT Large, XLSR-53) was conducted using this corpus. The experimental results further revealed XLSR-53 as the optimal architecture for clinical nursing speech recognition. A partial layer-freeze strategy was adopted in XLSR-53 by freezing the first 12 of 24 encoder layers, empirically validated through an ablation study with five freeze configurations (L=0, 6, 12, 18, 24). XLSR-53 preserves cross-lingual phonetic representations while enabling clinical vocabulary adaptation. A clinically motivated evaluation framework using curated medical vocabulary terms computes Medical Precision, Recall, and F1-Score along with standard WER, CER, and PER to assess reliability in clinical term recognition. Benchmarking against Google Health AI's MedASR zero-shot revealed that the proposed system XLSR-53 (L=12) achieved 17.15% WER against MedASR's 28.87% (p<0.001) with a Medical F1-Score of 0.98 and ROUGE-L of 0.92. Although results were obtained on synthetic Australian English speech, performance under real clinical conditions with background noise, overlapping speakers, and spontaneous interruptions requires further validation.

Downloads

Download data is not yet available.

Author Biographies

Sasikala D, Department of Computer Science, School of Engineering and Technology, Pondicherry University, Puducherry, India.

Sasikala D,  completed  her  B.E from  Government  College  of Technology  and  M.E  from  Sri Venkateswara    College    of Engineering. She worked as an assistant   professor   in   the Department    of Computer Science and Engineering at Sri Venkateswara    College    of Engineering  nearly  12  years. Currently she is pursing Ph.D(Part-time) in Pondicherry University, India. Also she is working as an Assistant Professor in the Department of Computer Science and Engineering at Amrita Vishwa Vidyapeetham,Chennai. Her  research  interests include Speech  Processing,  Natural  Language Processing,  Machine  Learning,  and  Deep  Learning. She has authored and co-authored 12 scopus indexed publications,  which  include a book  chapter, one journal publication and 10 international  conferences  in  the  field  of  Computer Science

Siva Sathya S, Department of Computer Science, School of Engineering and Technology, Pondicherry University, Puducherry, India

Dr. Siva Sathya S, is a Professor and Dean of the Department of Computer Science and School of Engineering & Technology at Pondicherry University, India. With over 25 years of academic experience, her research spans Artificial Intelligence, Natural Language Processing, Evolutionary Computing, and Spatio-Temporal Data Mining. She currently works in the field of medical informatics and AI-driven clinical support systems, serving as a technical lead for national AI consortia. A recipient of the prestigious National Teachers’ Award (2025) and the Nari Shakti Puraskar (2017) conferred by the President of India, her work focuses on leveraging deep learning architectures for social and healthcare innovations. She publishes her research works in SCIE and Scopus-indexed journals and serves as an expert on several national academic committees.

References

M. I. Alhosani, F. R. Ahmed, N. Al-Yateem, H. S. Mobarak, and M. E. AbuRuz, "Assessment of nursing workload and adverse events reporting among critical care nurses in the United Arab Emirates," The Open Nursing Journal, vol. 17, e18744346281511. ISSN: 1874-4346, 2023. DOI: 10.2174/0118744346281511231120054125.

Z. Banda, M. Simbota, and C. Mula, "Nurses' perceptions on the effects of high nursing workload on patient care in an intensive care unit of a referral hospital in Malawi: A qualitative study," BMC Nursing, vol. 21(1), 136, 2022. DOI: 10.1186/s12912-022-00918-x.

M. Shohani and H. Tavan, "Factors affecting medication errors from the perspective of nursing staff," Journal of Clinical and Diagnostic Research, vol.12(3), pp. IC01–IC04, 2018. DOI: 10.7860/JCDR/2018/28447.11336.

E. Gesner, P. C. Dykes, L. Zhang, and P. Gazarian, "Documentation burden in nursing and its role in clinician burnout syndrome," Applied Clinical Informatics, vol. 13(5), pp. 983–990, 2022. DOI: 10.1055/s-0042-1757157.

R. A. Atinga, M. N. Gmaligan, A. Ayawine, and J. K. Yambah, "It's the patient that suffers from poor communication: Analyzing communication gaps and associated consequences in handover events from nurses' experiences," SSM - Qualitative Research in Health, vol. 6, Art. no. 100482, 2024. DOI: 10.1016/j.ssmqr.2024.100482.

J. Y. Uhm, E. Y. Lim, and J. Hyeong, "The impact of a standardized inter-department handover on nurses' perceptions and performance in Republic of Korea," Journal of Nursing Management, vol. 26, pp. 933–944, 2018. DOI: 10.1111/jonm.12608.

S. Ghosh, L. Ramamoorthy, and B. Pottakat, "Impact of structured clinical handover protocol on communication and patient satisfaction," Journal of Patient Experience, vol. 8, Art. no. 2374373521997733, 2021. DOI: 10.1177/2374373521997733.

A. L. Cooper, J. A. Brown, S. P. Eccles, N. Cooper, and M. A. Albrecht, "Is nursing and midwifery clinical documentation a burden? An empirical study of perception versus reality," Journal of Clinical Nursing, vol. 30, pp. 1645–1652, 2021. DOI: 10.1111/jocn.15718.

I. Köse Tosunöz and A. Aydınlı, "Nurses' perceptions of the effectiveness of handover practices, influencing factors and perceived barriers: A descriptive cross-sectional study of medical and surgical nurses," Collegian, vol. 32, pp. 242–249, 2025. DOI: 10.1016/j.colegn.2025.06.002.

J. A. Brown, A. L. Cooper, and M. A. Albrecht, "Development and content validation of the Burden of Documentation for Nurses and Midwives (BurDoNsaM) survey," Journal of Advanced Nursing, vol. 76, pp. 1273–1281, 2020. DOI: 10.1111/jan.14320.

X. Luo, L. Zhou, K. Adelgais, et al., "Assessing the effectiveness of automatic speech recognition technology in emergency medicine settings: A comparative study of four AI-powered engines," Journal of Healthcare Informatics Research, vol. 9, pp. 494–512, 2025. DOI: 10.1007/s41666-025-00193-w.

K. Denecke, L. Meier, J. G. Bauer, M. Bender, and C. Lueg, "Information capturing in pre-hospital emergency medical settings (EMS)," Studies in Health Technology and Informatics, vol. 270, pp. 613–617, 2020. DOI: 10.3233/shti200233.

C. C. Chiu, A. Tripathi, K. Chou, C. Co, N. Jaitly, D. Jaunzeikare, A. Kannan, P. Nguyen, H. Sak, A. Sankar, et al., "Speech recognition for medical conversations," in Proc. Interspeech 2018, pp.2972–2976, 2018. DOI: 10.21437/Interspeech.2018-40.

E. Edwards, W. Salloum, G. P. Finley, J. Fone, G. Cardiff, M. Miller, and D. Suendermann-Oeft, "Medical speech recognition: Reaching parity with humans," in Speech and Computer, A. Karpov, R. Potapova, and I. Mporas, Eds. Cham, Switzerland, pp.512–524, 2017. DOI: 10.1007/978-3-319-66429-3_51.

T. Hodgson, F. Magrabi, and E. Coiera, "Efficiency and safety of speech recognition for documentation in the electronic health record," Journal of the American Medical Informatics Association, vol. 24, pp. 1127–1133, 2017. DOI: 10.1093/jamia/ocx073.

Angel, Maricel; Suominen, Hanna; Zhou, Liyuan; & Hanlen, Leif (2014): Synthetic nursing handover training and development data set - audio files. v1. CSIRO. Data Collection. DOI: 10.4225/08/58d0977ab4888.

Sasikala D, S. Siva Sathya, S. Baghavathi Priya, D. Niranjan Kumar, and S. Vignesh, "A review of automatic speech recognition approaches for bridging communication gaps in clinical handover," in Proc. 2nd IEEE International Conference for Women in Computing (InCoWoCo), pp.1–7, 2025. DOI: 10.1109/InCoWoCo68239.2025.11407097.

A. J. Nashwan, A. Abujaber, and S. K. Ahmed, "Charting the future: The role of AI in transforming nursing documentation," Cureus, vol. 16, no. 3, 2024. DOI: 10.7759/cureus.57304.

S. Yadav, "Embracing artificial intelligence: Revolutionizing nursing documentation for a better future," Cureus, vol. 16, no. 4, 2024. DOI: 10.7759/cureus.57725.

J. Kodish-Wachs, E. Agassi, P. I. Kenny, and J. M. Overhage, "A systematic comparison of contemporary automatic speech recognition engines for conversational clinical speech," AMIA Annual Symposium Proceedings, vol. 2018, pp. 683–689, 2018. PMCID: PMC6371385.

A. Femi-Abodunde, K. Olinger, L. M. B. Burke, T. Benefield, E. R. Lee, K. McGinty, and B. M. Mervak, "Radiology dictation errors with COVID-19 protective equipment: Does wearing a surgical mask increase the dictation error rate?", Journal of Digital Imaging, vol. 34, pp. 1294–1301, 2021. DOI: 10.1007/s10278-021-00502-w.

K. Le-Duc, "VietMed: A dataset and benchmark for automatic speech recognition of Vietnamese in the medical domain",arXiv, arXiv:2404.05659, 2024. DOI: 10.48550/arXiv.2404.05659.

S. Banerjee, A. Agarwal, and P. Ghosh, "High-precision medical speech recognition through synthetic data and semantic correction: UNITED-MEDASR," arXiv preprint, arXiv:2412.00055, 2024. DOI: 10.48550/arXiv.2412.00055.

T. Afonja, T. Olatunji, S. Ogun, N. A. Etori, A. Owodunni, and M. Yekini, "Performant ASR models for medical entities in accented speech," in Proc. Interspeech 2024, pp. 2315–2319, 2024. DOI: 10.21437/Interspeech.2024-2261.

A. Prasad and P. Jyothi, "How accents confound: Probing for accent information in end-to-end speech recognition systems," in Proc. 58th Annual Meeting of the Association for Computational Linguistics (ACL), 2020, pp. 3739–3753. DOI: 10.18653/v1/2020.acl-main.345.

M. Najafian, A. DeMarco, S. Cox, and M. Russell, "Unsupervised model selection for recognition of regional accented speech," in Proc. Interspeech 2014, pp.2967-2971, 2014. DOI: 10.21437/Interspeech.2014-495.

C. T. Do, S. Imai, R. S. Doddipatla, and T. Hain, "Improving accented speech recognition using data augmentation based on unsupervised text-to-speech synthesis," in Proceedings of the 2024 32nd European Signal Processing Conference (EUSIPCO), Lyon, France, 2024, pp. 136–140, 2024. DOI: 10.23919/EUSIPCO63174.2024.10715166.

A. Conneau, A. Baevski, R. Collobert, A. Mohamed, and M. Auli, "Unsupervised cross-lingual representation learning for speech recognition," in Proc. Interspeech 2021, Brno, Czech Republic, pp. 2426–2430, 2021. DOI: 10.21437/Interspeech.2021-329.

X. Zhang and L. He, "End-to-end cross-lingual spoken language understanding model with multilingual pretraining," in Proc. Interspeech 2021, pp.4728–4732, 2021. DOI: 10.21437/Interspeech.2021-818.

T. Pekarek Rosin and S. Wermter, "Replay to remember: Continual layer-specific fine-tuning for German speech recognition," in Proc. 32nd International Conference on Artificial Neural Networks (ICANN 2023), pp.489–500, 2023. DOI: 10.1007/978-3-031-44195-0_40.

Y. Li, K. Harrigian, A. Zirikly, and M. Dredze, "Are clinical T5 models better for clinical text?" ,in Proc. 4th Machine Learning for Health Symposium, Proceedings of Machine Learning Research, vol. 259, pp.636-667, 2025. DOI: 10.48550/arXiv.2412.05845.

P. Manakul, Y. Fathullah, A. Liusie, V. Raina, and M. Gales, "CUED at ProbSum 2023: Hierarchical ensemble of summarization models," in Proc. 22nd Workshop on Biomedical Natural Language Processing and BioNLP Shared Tasks, Toronto, Canada, pp. 516–523, 2023. DOI: 10.18653/v1/2023.bionlp-1.51.

C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, "Exploring the limits of transfer learning with a unified text-to-text transformer," arXiv preprint, arXiv:1910.10683, 2020. DOI: 10.48550/arXiv.1910.10683.

J. Giorgi et al., "Clinical note generation from doctor-patient conversations using large language models," in Proc. ClinicalNLP Workshop, pp.323-334, 2023. DOI: 10.18653/v1/2023.clinicalnlp-1.36.

A. Toma, P. R. Lawler, J. Ba, R. G. Krishnan, B. B. Rubin, and B. Wang, "Clinical Camel: An open expert-level medical language model with dialogue-based knowledge encoding," arXiv preprint, arXiv:2305.12031, 2023. DOI: 10.48550/arXiv.2305.12031.

Y. Labrak, A. Bazoge, E. Morin, P.-A. Gourraud, M. Rouvier, and R. Dufour, "BioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical Domains," in Findings of the Association for Computational Linguistics: ACL 2024, Bangkok, Thailand, pp. 5848–5864, 2024. DOI: 10.18653/v1/2024.findings-acl.348.

P. C. Nair, D. Gupta, and B. I. Devi, "Extracting clinical relationships from discharge summaries of suprasellar lesion patients using Gemini LLM", Procedia Computer Science, vol. 258, pp.2391-2404, 2025. DOI: 10.1016/j.procs.2025.04.502.

D. Sasikala, R. Sudarshan, and S. Sivasathya, "Harnessing LLMs for medical insights: NER extraction from summarized medical text," in Proc. 15th International Conference on Computing Communication and Networking Technologies(ICCCNT), pp.1–6, 2024. DOI: 10.1109/ICCCNT61001.2024.10724860.

A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, "Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks," in Proc. 23rd International Conference on Machine Learning (ICML), pp. 369–376, 2006. DOI: 10.1145/1143844.1143891.

D. Klakow and J. Peters, "Testing the correlation of word error rate and perplexity," Speech Communication, vol. 38, no. 1–2, pp. 19–28, 2002. DOI: 10.1016/S0167-6393(01)00041-3.

C. J. van Rijsbergen, Information Retrieval, 2nd ed. London, UK: Butterworths, 1979. Available: https://dl.acm.org/doi/book/10.5555/539927.

C.-Y. Lin, "ROUGE: A package for automatic evaluation of summaries," in Proc. ACL Workshop on Text Summarization Branches Out, Barcelona, Spain, pp. 74–81, 2004. ACL Anthology ID: W04-1013 . Available: https://aclanthology.org/W04-1013/.

A. V. Oppenheim and R. W. Schafer, Discrete-Time Signal Processing, 3rd ed. Upper Saddle River, NJ: Prentice Hall, 2009. ISBN: 978-0-13-198842-2.

A. Radford et al., "Robust speech recognition via large-scale weak supervision," in Proc. International Conference on Machine Learning (ICML), vol. 202, pp. 28492–28518, 2023. DOI: 10.48550/arXiv.2212.04356.

W. N. Hsu, B. Bolte, Y. H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, "HuBERT: Self-supervised speech representation learning by masked prediction of hidden units," IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 3451–3460, 2021. DOI: 10.1109/TASLP.2021.3122291.

A. Baevski, H. Zhou, A. Mohamed, and M. Auli, "wav2vec 2.0: A framework for self-supervised learning of speech representations," in Proc. Advances in Neural Information Processing Systems (NeurIPS), 2020, pp. 12449–12460. DOI: 10.48550/arXiv.2006.11477.

K. Wu, E. Variani, T. Bagby, S. Reddy, and R. Pilgrim, "MedASR: An Open-Source Model for High-Accuracy Medical Dictation,"

arXiv, arXiv:2605.16555, 2026. DOI: 10.48550/arXiv.2605.16555.

S. S. Shapiro and M. B. Wilk, "An analysis of variance test for normality (complete samples)," Biometrika, vol. 52, no. 3/4, pp. 591–611, 1965. DOI: 10.2307/2333709.

F. Wilcoxon, "Individual comparisons by ranking methods," Biometrics Bulletin, vol. 1, no. 6, pp. 80–83, 1945. DOI: 10.2307/3001968.

W. H. Kruskal and W. A. Wallis, "Use of ranks in one-criterion variance analysis," Journal of the American Statistical Association, vol. 47, no. 260, pp. 583–621, 1952. DOI: 10.1080/01621459.1952.10483441.

J. Cohen, Statistical Power Analysis for the Behavioral Sciences, 2nd ed. Hillsdale, NJ: Lawrence Erlbaum Associates, 1988. DOI: 10.4324/9780203771587.

B. Efron and R. J. Tibshirani, An Introduction to the Bootstrap. New York, NY: Chapman and all/CRC, 1994. DOI: 10.1201/9780429246593.

Suominen H, Zhou L, Hanlen L, Ferraro G, "Benchmarking clinical speech recognition and information extraction: New data, methods, and evaluations," JMIR Medical Informatics, vol. 3, no. 2, Art. no. e19, 2015.

DOI: 10.2196/medinform.4321.

Published
2026-08-15
How to Cite
[1]
S. D, S. S. S, N. K. D, and V. S, “Structured Nursing Handover Report Generation from Clinical Speech using Fine-Tuned XLSR-53 and T5: A Benchmarking Study”, j.electron.electromedical.eng.med.inform, vol. 8, no. 4, pp. 1371-1389, Aug. 2026.
Section
Medical Informatics