QT
Q. Tao
info
Please Note
<p>This page displays the records of the person named above and is not linked to a unique person identifier. This record may need to be merged to a profile.</p>
3 records found
1
AI-Assisted Point-of-Care Ultrasound for Breast Lesion Detection
Feasibility of a Standardized Blind-Sweep Workflow and Deep Learning-Based Lesion Detection
Introduction: Breast cancer remains one of the most common cancers among women worldwide, and early detection is essential for improving clinical outcomes. Patients presenting with breast complaints in primary care are often referred to hospital-based breast imaging because clinical assessment alone cannot reliably exclude malignancy. However, many referred patients are ultimately found to have benign findings or no focal abnormality. Point-of-care ultrasound (POCUS), particularly when combined with artificial intelligence (AI), may provide an accessible first-line assessment tool to support referral decisions before specialist breast imaging.
Aim: This thesis investigated the feasibility of an AI-assisted POCUS workflow for breast lesion detection using prospectively acquired blind-sweep ultrasound data. The objectives were as follows: to evaluate whether standardized breast POCUS blind sweeps could be acquired by a non-expert operator within an existing clinical diagnostic workflow, and to assess the performance of an AI model for breast lesion detection in the resulting image sequences.
Method: Prospective POCUS data were acquired from 86 participants referred for breast diagnostic assessment. Using a POCUS probe, a non-expert operator performed standardized partly overlapping blind sweeps of the selected breast. In total, 833 sweeps comprising more than 185,000 frames were acquired. The data were systematically stored, processed, and retrospectively annotated using the clinical imaging findings as reference. Public breast ultrasound datasets and POCUS breast phantom acquisitions were additionally used during model development. A You Only Look Once (YOLO)-based object detection model was trained to localize breast lesions, and performance was evaluated on independently held-out POCUS study, public, and phantom test data. Temporal post-processing was applied to retain detections supported by spatially corresponding predictions in nearby frames within the same sweep.
Results: Standardized breast POCUS blind-sweep acquisition was successfully integrated into the clinical workflow, with whole-breast acquisition according to protocol completed in 85 of 86 participants. On the prospective POCUS test set comprising seven participants, the final model achieved a framelevel sensitivity of 79.0% and specificity of 98.7% after temporal post-processing. At sweep level, all 16 lesion-positive sweeps contained at least one correct detection, while 27 of 28 lesion observations were detected in at least one frame. Performance was higher on public and phantom ultrasound data than on prospectively acquired POCUS data, highlighting the greater variability and complexity of non-targeted blind-sweep imaging. Nevertheless, the results demonstrate that automated lesion localization within prospectively acquired blind sweeps is feasible.
Conclusion: This thesis provides proof of concept for combining standardized non-expert breast POCUS blind-sweep acquisition with automated lesion detection in a clinical setting. The findings support further development of AI-assisted POCUS blind sweeps as a potential breast assessment or triage approach, while emphasizing that acquisition quality is an essential prerequisite for successful downstream analysis. Future work should focus on a larger prospective dataset, validation across multiple operators and centers, automated acquisition quality control, improved use of temporal sweep information, and extension of the pipeline towards lesion segmentation and benign–malignant classification. ...
Aim: This thesis investigated the feasibility of an AI-assisted POCUS workflow for breast lesion detection using prospectively acquired blind-sweep ultrasound data. The objectives were as follows: to evaluate whether standardized breast POCUS blind sweeps could be acquired by a non-expert operator within an existing clinical diagnostic workflow, and to assess the performance of an AI model for breast lesion detection in the resulting image sequences.
Method: Prospective POCUS data were acquired from 86 participants referred for breast diagnostic assessment. Using a POCUS probe, a non-expert operator performed standardized partly overlapping blind sweeps of the selected breast. In total, 833 sweeps comprising more than 185,000 frames were acquired. The data were systematically stored, processed, and retrospectively annotated using the clinical imaging findings as reference. Public breast ultrasound datasets and POCUS breast phantom acquisitions were additionally used during model development. A You Only Look Once (YOLO)-based object detection model was trained to localize breast lesions, and performance was evaluated on independently held-out POCUS study, public, and phantom test data. Temporal post-processing was applied to retain detections supported by spatially corresponding predictions in nearby frames within the same sweep.
Results: Standardized breast POCUS blind-sweep acquisition was successfully integrated into the clinical workflow, with whole-breast acquisition according to protocol completed in 85 of 86 participants. On the prospective POCUS test set comprising seven participants, the final model achieved a framelevel sensitivity of 79.0% and specificity of 98.7% after temporal post-processing. At sweep level, all 16 lesion-positive sweeps contained at least one correct detection, while 27 of 28 lesion observations were detected in at least one frame. Performance was higher on public and phantom ultrasound data than on prospectively acquired POCUS data, highlighting the greater variability and complexity of non-targeted blind-sweep imaging. Nevertheless, the results demonstrate that automated lesion localization within prospectively acquired blind sweeps is feasible.
Conclusion: This thesis provides proof of concept for combining standardized non-expert breast POCUS blind-sweep acquisition with automated lesion detection in a clinical setting. The findings support further development of AI-assisted POCUS blind sweeps as a potential breast assessment or triage approach, while emphasizing that acquisition quality is an essential prerequisite for successful downstream analysis. Future work should focus on a larger prospective dataset, validation across multiple operators and centers, automated acquisition quality control, improved use of temporal sweep information, and extension of the pipeline towards lesion segmentation and benign–malignant classification. ...
Introduction: Breast cancer remains one of the most common cancers among women worldwide, and early detection is essential for improving clinical outcomes. Patients presenting with breast complaints in primary care are often referred to hospital-based breast imaging because clinical assessment alone cannot reliably exclude malignancy. However, many referred patients are ultimately found to have benign findings or no focal abnormality. Point-of-care ultrasound (POCUS), particularly when combined with artificial intelligence (AI), may provide an accessible first-line assessment tool to support referral decisions before specialist breast imaging.
Aim: This thesis investigated the feasibility of an AI-assisted POCUS workflow for breast lesion detection using prospectively acquired blind-sweep ultrasound data. The objectives were as follows: to evaluate whether standardized breast POCUS blind sweeps could be acquired by a non-expert operator within an existing clinical diagnostic workflow, and to assess the performance of an AI model for breast lesion detection in the resulting image sequences.
Method: Prospective POCUS data were acquired from 86 participants referred for breast diagnostic assessment. Using a POCUS probe, a non-expert operator performed standardized partly overlapping blind sweeps of the selected breast. In total, 833 sweeps comprising more than 185,000 frames were acquired. The data were systematically stored, processed, and retrospectively annotated using the clinical imaging findings as reference. Public breast ultrasound datasets and POCUS breast phantom acquisitions were additionally used during model development. A You Only Look Once (YOLO)-based object detection model was trained to localize breast lesions, and performance was evaluated on independently held-out POCUS study, public, and phantom test data. Temporal post-processing was applied to retain detections supported by spatially corresponding predictions in nearby frames within the same sweep.
Results: Standardized breast POCUS blind-sweep acquisition was successfully integrated into the clinical workflow, with whole-breast acquisition according to protocol completed in 85 of 86 participants. On the prospective POCUS test set comprising seven participants, the final model achieved a framelevel sensitivity of 79.0% and specificity of 98.7% after temporal post-processing. At sweep level, all 16 lesion-positive sweeps contained at least one correct detection, while 27 of 28 lesion observations were detected in at least one frame. Performance was higher on public and phantom ultrasound data than on prospectively acquired POCUS data, highlighting the greater variability and complexity of non-targeted blind-sweep imaging. Nevertheless, the results demonstrate that automated lesion localization within prospectively acquired blind sweeps is feasible.
Conclusion: This thesis provides proof of concept for combining standardized non-expert breast POCUS blind-sweep acquisition with automated lesion detection in a clinical setting. The findings support further development of AI-assisted POCUS blind sweeps as a potential breast assessment or triage approach, while emphasizing that acquisition quality is an essential prerequisite for successful downstream analysis. Future work should focus on a larger prospective dataset, validation across multiple operators and centers, automated acquisition quality control, improved use of temporal sweep information, and extension of the pipeline towards lesion segmentation and benign–malignant classification.
Aim: This thesis investigated the feasibility of an AI-assisted POCUS workflow for breast lesion detection using prospectively acquired blind-sweep ultrasound data. The objectives were as follows: to evaluate whether standardized breast POCUS blind sweeps could be acquired by a non-expert operator within an existing clinical diagnostic workflow, and to assess the performance of an AI model for breast lesion detection in the resulting image sequences.
Method: Prospective POCUS data were acquired from 86 participants referred for breast diagnostic assessment. Using a POCUS probe, a non-expert operator performed standardized partly overlapping blind sweeps of the selected breast. In total, 833 sweeps comprising more than 185,000 frames were acquired. The data were systematically stored, processed, and retrospectively annotated using the clinical imaging findings as reference. Public breast ultrasound datasets and POCUS breast phantom acquisitions were additionally used during model development. A You Only Look Once (YOLO)-based object detection model was trained to localize breast lesions, and performance was evaluated on independently held-out POCUS study, public, and phantom test data. Temporal post-processing was applied to retain detections supported by spatially corresponding predictions in nearby frames within the same sweep.
Results: Standardized breast POCUS blind-sweep acquisition was successfully integrated into the clinical workflow, with whole-breast acquisition according to protocol completed in 85 of 86 participants. On the prospective POCUS test set comprising seven participants, the final model achieved a framelevel sensitivity of 79.0% and specificity of 98.7% after temporal post-processing. At sweep level, all 16 lesion-positive sweeps contained at least one correct detection, while 27 of 28 lesion observations were detected in at least one frame. Performance was higher on public and phantom ultrasound data than on prospectively acquired POCUS data, highlighting the greater variability and complexity of non-targeted blind-sweep imaging. Nevertheless, the results demonstrate that automated lesion localization within prospectively acquired blind sweeps is feasible.
Conclusion: This thesis provides proof of concept for combining standardized non-expert breast POCUS blind-sweep acquisition with automated lesion detection in a clinical setting. The findings support further development of AI-assisted POCUS blind sweeps as a potential breast assessment or triage approach, while emphasizing that acquisition quality is an essential prerequisite for successful downstream analysis. Future work should focus on a larger prospective dataset, validation across multiple operators and centers, automated acquisition quality control, improved use of temporal sweep information, and extension of the pipeline towards lesion segmentation and benign–malignant classification.
Unlike modality or structure-specific segmentation models such as U-Net, foundation models like Medical Segment Anything Model (MedSAM) segment across imaging modalities and anatomical structures without retraining, at a fixed cost. MedSAM pairs a Vision Transformer Base Model (ViT-B) image encoder with a prompt encoder and a mask decoder. The encoder dominates, accounting for 95.66% of the parameters and, at 966 giga Floating Point Operations (FLOPs), for almost all of the computation of a single forward pass, incurred for every image regardless of the modality segmented. This work characterises that cost analytically, deriving a block’s FLOPs from its retained attention heads and Multi-Layer Perceptron (MLP) neurons rather than by profiling, and then asks how much of the encoder is redundant for medical image segmentation. Redundancy is studied at two levels. At the level of the individual weight, redundant weights are zeroed while the weight-matrix shapes are left unchanged, so no computation is saved. At the level of the whole computational structure, removing an attention head or an MLP neuron physically reduces the weight count. It is the only level at which inference is accelerated. Sparsity is induced during training by a penalty on the loss: the Lasso ℓ1 penalty drives individual weights towards zero, paired with unstructured pruning, whilst Sparse Group Lasso (SGL) groups the weights of each head and neuron and drives whole groups towards zero, paired with structured pruning. As both penalties are non-smooth at zero, they are imposed via Proximal Gradient Descent (PGD), in which the gradient step optimises the task loss and a closed-form shrinkage sets small weights to exactly zero. Both pruned regimes are compared against a dense task-loss-fine-tuned reference and a variant with the encoder removed entirely, the latter serving as a floor
that tests whether the encoder is needed at all for a given target. Evaluated across seven imaging modalities and eleven evaluation subsets, the encoder is found to hold substantial but unequal redundancy. At the individual-weight level, 72.49% of the prunable encoder weights can be zeroed while maintaining accuracy comparable to the baseline, but the irregular sparsity yields no computational savings. As whole structures, 91.7% of the prunable encoder weight volume can be removed, reducing the model from 93.74 to 15.82 million parameters and the encoder cost from 966.32 to 81.08 giga FLOPs, which does come at a cost of the Dice Similarity Coeffi-
cient (DSC). This cost is not uniform across structures. Some targets retain usable accuracy with the encoder fully removed, such as the ultrasound foetal head and the fundus optic disc, both above 92% DSC, whilst others depend heavily on it, most notably the Chest X-ray
(CXR) lung fields, whose median DSC falls to 54.84%. The per-structure cost ranges from
the optic cup, which improves by 3.08 points under compression, to GlaS (testB), which
falls by 8.60. The MedSAM encoder therefore carries redundancy that can be removed for
real savings in size and speed. However, the extent of removal depends on the structure
and modality being segmented. ...
that tests whether the encoder is needed at all for a given target. Evaluated across seven imaging modalities and eleven evaluation subsets, the encoder is found to hold substantial but unequal redundancy. At the individual-weight level, 72.49% of the prunable encoder weights can be zeroed while maintaining accuracy comparable to the baseline, but the irregular sparsity yields no computational savings. As whole structures, 91.7% of the prunable encoder weight volume can be removed, reducing the model from 93.74 to 15.82 million parameters and the encoder cost from 966.32 to 81.08 giga FLOPs, which does come at a cost of the Dice Similarity Coeffi-
cient (DSC). This cost is not uniform across structures. Some targets retain usable accuracy with the encoder fully removed, such as the ultrasound foetal head and the fundus optic disc, both above 92% DSC, whilst others depend heavily on it, most notably the Chest X-ray
(CXR) lung fields, whose median DSC falls to 54.84%. The per-structure cost ranges from
the optic cup, which improves by 3.08 points under compression, to GlaS (testB), which
falls by 8.60. The MedSAM encoder therefore carries redundancy that can be removed for
real savings in size and speed. However, the extent of removal depends on the structure
and modality being segmented. ...
Unlike modality or structure-specific segmentation models such as U-Net, foundation models like Medical Segment Anything Model (MedSAM) segment across imaging modalities and anatomical structures without retraining, at a fixed cost. MedSAM pairs a Vision Transformer Base Model (ViT-B) image encoder with a prompt encoder and a mask decoder. The encoder dominates, accounting for 95.66% of the parameters and, at 966 giga Floating Point Operations (FLOPs), for almost all of the computation of a single forward pass, incurred for every image regardless of the modality segmented. This work characterises that cost analytically, deriving a block’s FLOPs from its retained attention heads and Multi-Layer Perceptron (MLP) neurons rather than by profiling, and then asks how much of the encoder is redundant for medical image segmentation. Redundancy is studied at two levels. At the level of the individual weight, redundant weights are zeroed while the weight-matrix shapes are left unchanged, so no computation is saved. At the level of the whole computational structure, removing an attention head or an MLP neuron physically reduces the weight count. It is the only level at which inference is accelerated. Sparsity is induced during training by a penalty on the loss: the Lasso ℓ1 penalty drives individual weights towards zero, paired with unstructured pruning, whilst Sparse Group Lasso (SGL) groups the weights of each head and neuron and drives whole groups towards zero, paired with structured pruning. As both penalties are non-smooth at zero, they are imposed via Proximal Gradient Descent (PGD), in which the gradient step optimises the task loss and a closed-form shrinkage sets small weights to exactly zero. Both pruned regimes are compared against a dense task-loss-fine-tuned reference and a variant with the encoder removed entirely, the latter serving as a floor
that tests whether the encoder is needed at all for a given target. Evaluated across seven imaging modalities and eleven evaluation subsets, the encoder is found to hold substantial but unequal redundancy. At the individual-weight level, 72.49% of the prunable encoder weights can be zeroed while maintaining accuracy comparable to the baseline, but the irregular sparsity yields no computational savings. As whole structures, 91.7% of the prunable encoder weight volume can be removed, reducing the model from 93.74 to 15.82 million parameters and the encoder cost from 966.32 to 81.08 giga FLOPs, which does come at a cost of the Dice Similarity Coeffi-
cient (DSC). This cost is not uniform across structures. Some targets retain usable accuracy with the encoder fully removed, such as the ultrasound foetal head and the fundus optic disc, both above 92% DSC, whilst others depend heavily on it, most notably the Chest X-ray
(CXR) lung fields, whose median DSC falls to 54.84%. The per-structure cost ranges from
the optic cup, which improves by 3.08 points under compression, to GlaS (testB), which
falls by 8.60. The MedSAM encoder therefore carries redundancy that can be removed for
real savings in size and speed. However, the extent of removal depends on the structure
and modality being segmented.
that tests whether the encoder is needed at all for a given target. Evaluated across seven imaging modalities and eleven evaluation subsets, the encoder is found to hold substantial but unequal redundancy. At the individual-weight level, 72.49% of the prunable encoder weights can be zeroed while maintaining accuracy comparable to the baseline, but the irregular sparsity yields no computational savings. As whole structures, 91.7% of the prunable encoder weight volume can be removed, reducing the model from 93.74 to 15.82 million parameters and the encoder cost from 966.32 to 81.08 giga FLOPs, which does come at a cost of the Dice Similarity Coeffi-
cient (DSC). This cost is not uniform across structures. Some targets retain usable accuracy with the encoder fully removed, such as the ultrasound foetal head and the fundus optic disc, both above 92% DSC, whilst others depend heavily on it, most notably the Chest X-ray
(CXR) lung fields, whose median DSC falls to 54.84%. The per-structure cost ranges from
the optic cup, which improves by 3.08 points under compression, to GlaS (testB), which
falls by 8.60. The MedSAM encoder therefore carries redundancy that can be removed for
real savings in size and speed. However, the extent of removal depends on the structure
and modality being segmented.
Quantitative cardiac MRI is an increasingly important diagnostic tool for cardiovascular diseases. Yet, it is essential to have correct image registration for good accuracy and precision of quantitative mapping. Registering all baseline images from a quantitative cardiac MRI sequence, however, is nontrivial because the patient is moving, leading to simultaneous changes in motion, intensity, and contrast. The changes in image contrast, in particular, make it challenging to design a reliable registration metric for optimization.
In this paper, we propose a novel approach based on robust principle component analysis (rPCA) that decomposes quantitative cardiac MRI into low-rank and sparse components, in combination with a groupwise CNN-based registration backbone. The proposed framework aims for fast, robust motion correction for contrast-agnostic sequences, which benefits registration. We evaluated our proposed method on cardiac T1 mapping sequences, both pre-contrast and post-contrast. Additionally, we synthesize the numerical phantoms with gold standard to test the performance. Our experiments showed that our method effectively improved registration performance over baseline methods without rPCA, and reduced quantitative mapping error in both in-domain and out-of-domain MRI sequences. The proposed rPCA framework is generic and can be easily incorporated into existing registration methods and other clinical applications. ...
In this paper, we propose a novel approach based on robust principle component analysis (rPCA) that decomposes quantitative cardiac MRI into low-rank and sparse components, in combination with a groupwise CNN-based registration backbone. The proposed framework aims for fast, robust motion correction for contrast-agnostic sequences, which benefits registration. We evaluated our proposed method on cardiac T1 mapping sequences, both pre-contrast and post-contrast. Additionally, we synthesize the numerical phantoms with gold standard to test the performance. Our experiments showed that our method effectively improved registration performance over baseline methods without rPCA, and reduced quantitative mapping error in both in-domain and out-of-domain MRI sequences. The proposed rPCA framework is generic and can be easily incorporated into existing registration methods and other clinical applications. ...
Quantitative cardiac MRI is an increasingly important diagnostic tool for cardiovascular diseases. Yet, it is essential to have correct image registration for good accuracy and precision of quantitative mapping. Registering all baseline images from a quantitative cardiac MRI sequence, however, is nontrivial because the patient is moving, leading to simultaneous changes in motion, intensity, and contrast. The changes in image contrast, in particular, make it challenging to design a reliable registration metric for optimization.
In this paper, we propose a novel approach based on robust principle component analysis (rPCA) that decomposes quantitative cardiac MRI into low-rank and sparse components, in combination with a groupwise CNN-based registration backbone. The proposed framework aims for fast, robust motion correction for contrast-agnostic sequences, which benefits registration. We evaluated our proposed method on cardiac T1 mapping sequences, both pre-contrast and post-contrast. Additionally, we synthesize the numerical phantoms with gold standard to test the performance. Our experiments showed that our method effectively improved registration performance over baseline methods without rPCA, and reduced quantitative mapping error in both in-domain and out-of-domain MRI sequences. The proposed rPCA framework is generic and can be easily incorporated into existing registration methods and other clinical applications.
In this paper, we propose a novel approach based on robust principle component analysis (rPCA) that decomposes quantitative cardiac MRI into low-rank and sparse components, in combination with a groupwise CNN-based registration backbone. The proposed framework aims for fast, robust motion correction for contrast-agnostic sequences, which benefits registration. We evaluated our proposed method on cardiac T1 mapping sequences, both pre-contrast and post-contrast. Additionally, we synthesize the numerical phantoms with gold standard to test the performance. Our experiments showed that our method effectively improved registration performance over baseline methods without rPCA, and reduced quantitative mapping error in both in-domain and out-of-domain MRI sequences. The proposed rPCA framework is generic and can be easily incorporated into existing registration methods and other clinical applications.