Journal of Electronics, Electromedical Engineering, and Medical Informatics
https://jeeemi.org/index.php/jeeemi
<p>The Journal of Electronics, Electromedical Engineering, and Medical Informatics, (JEEEMI), is a peer-reviewed periodical scientific journal aimed at publishing research results of the Journal focus areas. The Journal is published by the Department of Electromedical Engineering, Health Polytechnic of Surabaya, Ministry of Health, Indonesia. The role of the Journal is to facilitate contacts between research centers and the industry. The aspiration of the Editors is to publish high-quality scientific professional papers presenting works of significant scientific teams, experienced and well-established authors as well as postgraduate students and beginning researchers. All articles are subject to anonymous review processes by at least two independent expert reviewers prior to publishing on the International Journal of Electronics, Electromedical Engineering, and Medical Informatics website.</p>Department of Electromedical Engineering, POLTEKKES KEMENKES SURABAYAen-US Journal of Electronics, Electromedical Engineering, and Medical Informatics2656-8632<p><strong>Authors who publish with this journal agree to the following terms:</strong></p> <ol> <li class="show">Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution-ShareAlikel 4.0 International <a title="CC BY SA" href="https://creativecommons.org/licenses/by-sa/4.0/" target="_blank" rel="noopener">(CC BY-SA 4.0)</a> that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.</li> <li class="show">Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.</li> <li class="show">Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See <a href="http://opcit.eprints.org/oacitation-biblio.html" target="_new">The Effect of Open Access</a>).</li> </ol>STFT-Based Multiclass Heart Sound Classification Using BiLSTM and CNN-BiLSTM Models
https://jeeemi.org/index.php/jeeemi/article/view/1724
<p>Cardiovascular diseases require early and reliable screening because manual auscultation may be affected by noise, subjective interpretation, and inter-observer variability. This study aimed to develop an STFT-based deep learning framework for multiclass phonocardiogram (PCG) classification. The proposed framework was designed to provide a reproducible evaluation procedure by combining standardized preprocessing, time–frequency feature extraction, and deep learning-based classification under the same experimental conditions. Unlike approaches that may evaluate segmented signals without clearly preserving recording-level separation, this study emphasizes a leakage-free splitting strategy to reduce the risk of overestimated performance and to provide a more reliable assessment of model generalization. The Yaseen PCG dataset, consisting of 1000 recordings from five classes (AS, MR, MS, MVP, and Normal), was divided using a leakage-free recording-level split before segmentation and spectrogram generation. After preprocessing, 2-second PCG segments with 50% overlap were converted into 128 × 128 STFT spectrograms and classified using BiLSTM and CNN-BiLSTM models. Both models were trained and tested using the same dataset split, preprocessing pipeline, and evaluation metrics, including accuracy, precision, recall, F1-score, specificity, and confusion matrices. The BiLSTM model achieved 92.36% accuracy in the final independent test run, while the CNN-BiLSTM model achieved 95.83%. Across three repeated runs, BiLSTM achieved 93.85% ± 1.47%, whereas CNN-BiLSTM achieved 95.94% ± 0.71%. These results show that CNN-BiLSTM provides higher and more stable classification performance for five-class PCG classification, while BiLSTM remains a simpler alternative for lightweight implementation. Overall, the proposed STFT-based framework provides a reliable approach for automated heart sound classification and may support future computer-aided cardiac screening applications.</p>Noor S.Ehab Abdulrazzaq HusseinLaith Ali Abdul-Rahaim
Copyright (c) 2026 Noor S., Ehab Abdulrazzaq Hussein, Laith Ali Abdul-Rahaim
https://creativecommons.org/licenses/by-sa/4.0
2026-08-032026-08-03841296131310.35882/jeeemi.v8i4.1724Electrooculography-Based Voluntary Eye-Blink Detection for Arabic Assistive Communication
https://jeeemi.org/index.php/jeeemi/article/view/1848
<p class="Abstract" style="margin: 0cm -1.15pt 8.0pt 0cm;"><span style="font-family: 'Arial',sans-serif;">People with profound neuromotor disabilities commonly experience significant speech disability and lack of voluntary control over their limbs, retaining, however, the ability to perform deliberate eyelid blinks. Such retained capability presents an opportunity to be used as assistive communication. Nevertheless, reliable detection of voluntary blinks and decoding of communication patterns based on them is challenged by variation in the waveforms, amplitudes, durations, and inter-blink intervals of the voluntary eye blinks. In this work, a cost-efficient vertical electrooculography (EOG) approach for detecting voluntary eye blinks and decoding predefined assistive messages in Arabic was presented and assessed. The pipeline consists of signal conditioning, root-mean-square envelope calculation, adaptive hysteresis-based thresholding, temporal segmentation, and rule-based interpretation of single- and double-blink patterns. Quiet-background threshold updating and locking it during the communication window period were used to increase the robustness of the pattern detection process, while autocorrelation analysis and temporal constraints were utilized to interpret patterns. A single blink is associated with binary 0, and a double blink with binary 1; four consecutive pattern positions form a 4-bit frame that represents one of sixteen predefined Arabic messages. Unclear patterns are excluded from processing and not mapped into any valid message. The proposed approach was evaluated offline on 4,000 word-level recordings from 25 healthy participants, with verification using videos. The F1-score of 95.69%, <span dir="RTL" lang="AR-SA">macro-F1-score</span> of 98.30%, and message-decoding accuracy of 92.98% were obtained. These findings demonstrate the possibility of implementing an interpretable and cost-efficient vertical-EOG approach for message-level Arabic assistive communication without morse-code-type encoding and letter-by-letter spelling. Direct mapping of patterns into messages can help reduce communication effort and time. Nevertheless, since the present evaluation was conducted offline and included only healthy subjects, further validation with target users and real-time implementation should be conducted before deployment.</span></p>Afrah ThamerMussab Alaziz
Copyright (c) 2026 Afrah Thamer, Mussab Alaziz
https://creativecommons.org/licenses/by-sa/4.0
2026-08-112026-08-11841331135310.35882/jeeemi.v8i4.1848A Dual-Stage Deep Lesion Segmentation Framework with Gradient and Boundary Optimization for Diabetic Retinopathy Retinal Images
https://jeeemi.org/index.php/jeeemi/article/view/1734
<p>Diabetic Retinopathy (DR) is a leading global cause of preventable blindness, and the early detection and precise segmentation of pathologies play a key role in clinical decisions in this disease. Conventional deep learning models, however, suffer from the problem of weak lesion boundaries, gradient inconsistency, and structural distortion in heterogeneous datasets. To overcome these drawbacks, we present a sophisticated systemthat combines GEONet (Gradient and Edge-Optimized Network) to achieve accurate gradient- and edge-aware segmentation, and BESNet (Boundary-Enhanced Segmentation Network) to introduce boundary refinement and structural preservation. This two-step strategy can guarantee the high-fidelity capture of subtle retinal lesions with anatomical consistency. Comparative analysis was carried out against state-of-the-art baselines. Experimental results on the EyePACS dataset demonstrate the effectiveness of the proposed framework. The integrated GEONet+BESNet architecture achieved a Dice Coefficient of 89.0%, Intersection-over-Union (IoU) of 87.0%, Boundary Accuracy of 89.0%, and a Hausdorff Distance of 5.1, outperforming all competing methods in terms of segmentation fidelity and boundary preservation. Structural consistency was further enhanced, attaining a Dice Similarity Coefficient (DSC) of 90.5% and a Structural Similarity Index Measure (SSIM) of 90.3%. From a clinical screening perspective, the proposed framework achieved a Precision of 92.8% and Specificity of 94.6%, indicating reliable lesion localization with a reduced rate of false-positive detections. These findings confirm that the synergistic integration of gradient-aware segmentation through GEONet and boundary-enhanced refinement through BESNet effectively preserves lesion morphology and improves segmentation robustness across heterogeneous retinal imaging conditions.</p>Malathi PDhinakaran DJeyalakshmi SKavitha PPalpandi SPrabaharan S
Copyright (c) 2026 Malathi P, Dhinakaran D, Jeyalakshmi S, Kavitha P, Palpandi S, Prabaharan S
https://creativecommons.org/licenses/by-sa/4.0
2026-08-112026-08-11841314133010.35882/jeeemi.v8i4.1734A Multimodal Graph Neural Network for Multiclass ADHD and ASD Classification with Leakage-Aware Evaluation
https://jeeemi.org/index.php/jeeemi/article/view/1801
<p>Neurodevelopmental disorders such as attention deficit hyperactivity disorder (ADHD) and autism spectrum disorder (ASD) share overlapping clinical symptoms, complicating diagnosis and motivating objective, data-driven approaches using neuroimaging and machine learning. Graph neural networks (GNNs) have shown strong performance in this domain, yet many existing studies rely on single-modality data, transductive learning, and feature selection procedures that may introduce information leakage and inflate reported accuracy. This study proposes a multimodal graph learning framework that integrates resting-state fMRI (rs-fMRI), structural MRI (sMRI), and demographic data for classifying ADHD, ASD, and healthy controls (HC). The framework integrates temporal stability-based functional connectivity, hybrid feature selection, and adaptive multi-graph learning to exploit complementary information across these modalities. Using the ADHD-200 and ABIDE datasets, the framework is evaluated under three protocols that progressively tighten control over information leakage: transductive learning with global feature selection, inductive learning with global feature selection, and inductive learning with fold-wise feature selection. Results show that classification performance is highest under the transductive, globally-selected setting (85.5% accuracy for HC vs ADHD vs ASD, 92.5% for HC vs ASD, and 90.4% for HC vs ADHD), but decreases under the strictest leakage-aware protocol (70.9%, 81.7%, and 79.3%, respectively). This performance gap indicates that conventional evaluation protocols can substantially overestimate real-world generalization. Importantly, the proposed framework still achieves reasonable accuracy under the strictest setting, suggesting genuine discriminative capability beyond evaluation artifacts. These findings emphasize that leakage-aware evaluation, although yielding lower numbers, provides a more realistic and trustworthy estimate of model performance, highlighting its importance for developing reliable neuroimaging-based GNN models</p>Chofifatul HidayahWiharto WihartoEsti Suryani
Copyright (c) 2026 Chofifatul Hidayah, Wiharto Wiharto, Esti Suryani
https://creativecommons.org/licenses/by-sa/4.0
2026-08-112026-08-11841354137010.35882/jeeemi.v8i4.1801Adaptive Sparse Cross-Scale Transformer for Computationally Efficient Brain Tumor MRI Segmentation
https://jeeemi.org/index.php/jeeemi/article/view/1742
<p>Brain tumor segmentation from Magnetic Resonance Imaging (MRI) is a critical task in medical image analysis, owing to the irregular morphology of tumors, heterogeneous intensity distributions across MRI modalities, and the presence of multiple overlapping sub-regions. Accurate delineation of these regions is essential for clinical diagnosis, treatment planning, and longitudinal disease monitoring. Although Vision Transformer (ViT)-based architectures have demonstrated promising performance in medical image segmentation by capturing long-range dependencies and global contextual information, conventional multi-scale transformer models remain constrained by static grouped attention mechanisms that introduce redundant computations and exhibit quadratic complexity with respect to input size. To address these limitations, this paper proposes the Adaptive Sparse Cross-Scale Transformer (ASCT), a computationally efficient framework for brain tumor MRI segmentation. ASCT incorporates three key innovations: (i) a Dynamic Scale Routing (DSR) module that adaptively weights multi-scale features using learned routing coefficients, replacing fixed channel grouping; (ii) a Sparse Token Attention (STA) mechanism that restricts attention computation to the most informative token pairs, reducing complexity from quadratic O(N²) to near-linear O(Nk); and (iii) a linearized attention approximation that significantly reduces GPU memory consumption during training. Additionally, cross-scale feature fusion is performed prior to the attention operation to suppress redundant computations and enhance inter-scale feature interaction. The proposed ASCT model is evaluated on the BraTS 2021 dataset for multi-region segmentation, targeting Whole Tumor (WT), Tumor Core (TC), and Enhancing Tumor (ET) sub-regions. Experimental results demonstrate that ASCT achieves Dice scores of 89.2%, 84.1%, and 83.5% for WT, TC, and ET, respectively, yielding an average Dice score of 85.6%. Compared to the baseline Vision Transformer, the proposed model reduces computational complexity from 94.6 GFLOPs to 58.4 GFLOPs and GPU memory usage from 11.2 GB to 7.3 GB, confirming its efficiency and practical viability for real-world clinical applications.</p>Ravichandra BandiSelvanayaki SAmitha I. C.Sudha SubramaniamSuganthi RAnand Rajendran
Copyright (c) 2026 Ravichandra Bandi, Selvanayaki S, Amitha Ida Chandran, Sudha Subramaniam, Suganthi R, Anand Rajendran
https://creativecommons.org/licenses/by-sa/4.0
2026-08-132026-08-13841389140510.35882/jeeemi.v8i4.1742Structured Nursing Handover Report Generation from Clinical Speech using Fine-Tuned XLSR-53 and T5: A Benchmarking Study
https://jeeemi.org/index.php/jeeemi/article/view/1778
<p>Accurate nursing handovers are critical for patient safety, as miscommunication during shift transitions leads to irreversible clinical errors. This work proposed an end-to-end pipeline that converts unstructured clinical nursing speech into standardized handover reports using a fine-tuned XLSR-53 acoustic model and T5-base text-to-text transformer. An Australian English clinical corpus of 200 synthetic nursing handover recordings from the CSIRO data access portal was utilised in this work. This benchmarking study was conducted within the CSIRO synthetic Australian English nursing handover corpus and does not represent a broad cross-domain clinical ASR benchmark. A domain-specific benchmarking study across seven state-of-the-art ASR architectures (Whisper Tiny/Base/Small, Wav2Vec2 Base/Large, HuBERT Large, XLSR-53) was conducted using this corpus. The experimental results further revealed XLSR-53 as the optimal architecture for clinical nursing speech recognition. A partial layer-freeze strategy was adopted in XLSR-53 by freezing the first 12 of 24 encoder layers, empirically validated through an ablation study with five freeze configurations (L=0, 6, 12, 18, 24). XLSR-53 preserves cross-lingual phonetic representations while enabling clinical vocabulary adaptation. A clinically motivated evaluation framework using curated medical vocabulary terms computes Medical Precision, Recall, and F1-Score along with standard WER, CER, and PER to assess reliability in clinical term recognition. Benchmarking against Google Health AI's MedASR zero-shot revealed that the proposed system XLSR-53 (L=12) achieved 17.15% WER against MedASR's 28.87% (p<0.001) with a Medical F1-Score of 0.98 and ROUGE-L of 0.92. Although results were obtained on synthetic Australian English speech, performance under real clinical conditions with background noise, overlapping speakers, and spontaneous interruptions requires further validation<strong>.</strong></p>Sasikala DSiva Sathya SNiranjan Kumar DVignesh S
Copyright (c) 2026 Sasikala D, Siva Sathya S, Niranjan Kumar D, Vignesh S
https://creativecommons.org/licenses/by-sa/4.0
2026-08-152026-08-15841371138910.35882/jeeemi.v8i4.1778