X. Zhang
Please Note
18 records found
1
In acute ischemic stroke, large vessel occlusions of the anterior circulation are increasingly treated with endovascular therapy (EVT). The efficacy of this therapy depends on adequate treatment selection. Treatment decisions can be based on predictions of functional outcome. Most existing studies predict functional outcomes using clinical parameters. We set out to study functional outcome prediction performance by integrating imaging in a multimodal setting. Using a multi-center dataset containing 2927 patients, we compare the functional outcome prediction performances of clinical baseline models, including the clinically validated MR PREDICTS decision tool, image-based models with deep learning networks, and a multimodal approach combining clinical and imaging information. The predicted outcome measure is dichotomized modified Rankin Scale score 90 days after EVT. We perform sanity checks, hyperparameter optimization, and comparisons of effectiveness of using CTA, NCCT, or both images as input. Our experiments show that information extracted from CTA or NCCT images does not significantly improve the performance, as quantified using AUC, of functional outcome prediction methods compared to a baseline model. The multimodal approach may replace radiologically derived biomarkers, as its performance is non-inferior.
FocusViT
Dynamic patch focus for transformer-based gaze estimation
Anticipating daily human actions
Comparing pipelines for long-term skeleton-based prediction in real-world scenarios
Application of Augmented Reality in Robot-Assisted Mitral Valve Repair Surgery
A Feasibility Study
In mitral valve surgery, it is important to be aware of adjacent intraoperatively invisible anatomy, to avoid complications and enhance safety. In this feasibility study, we aimed to develop semi-automated intraoperative 3-dimensional (3D) augmented reality (3D-AR) overlays for robotic mitral valve repair.
Methods:
In 5 patients undergoing robot-assisted mitral valve repair, a 3D point cloud was generated, using intraoperatively recorded images from both eyes of the stereoscopic da Vinci camera (Intuitive Surgical, Sunnyvale, CA, USA). An intraoperative 3D-AR overlay was created using a scale-adaptive iterative closest point algorithm and landmarks placed on the mitral valve annulus. Finally, important anatomical structures such as the circumflex artery, Koch’s triangle, and aortic valve leaflets could be visualized as a 3D-AR overlay on top of the surgical vision. To evaluate the accuracy, these 3D point clouds were validated by calculating the 3D point cloud accuracy and landmark registration error (LRE).
Results:
The 3D point clouds and 3D-AR overlays were successfully created for all 5 patients. The 3D point clouds were accurate, with a median error of −0.92 mm, and the LRE was 5.12 mm. The time for creating the 3D-AR overlay was approximately 5 min. Besides creating the 3D-AR overlays, we could visualize the models directly within the robotic console during the surgical procedure.
Conclusions:
We present an algorithm for generating accurate semiautomatic 3D-AR overlays, visualizing essential anatomical structures during robot-assisted mitral valve repair. This may lead to automated intraoperative 3D-AR vision during robotic cardiac surgery, with the potential of increasing safety, accuracy, and efficiency. ...
In mitral valve surgery, it is important to be aware of adjacent intraoperatively invisible anatomy, to avoid complications and enhance safety. In this feasibility study, we aimed to develop semi-automated intraoperative 3-dimensional (3D) augmented reality (3D-AR) overlays for robotic mitral valve repair.
Methods:
In 5 patients undergoing robot-assisted mitral valve repair, a 3D point cloud was generated, using intraoperatively recorded images from both eyes of the stereoscopic da Vinci camera (Intuitive Surgical, Sunnyvale, CA, USA). An intraoperative 3D-AR overlay was created using a scale-adaptive iterative closest point algorithm and landmarks placed on the mitral valve annulus. Finally, important anatomical structures such as the circumflex artery, Koch’s triangle, and aortic valve leaflets could be visualized as a 3D-AR overlay on top of the surgical vision. To evaluate the accuracy, these 3D point clouds were validated by calculating the 3D point cloud accuracy and landmark registration error (LRE).
Results:
The 3D point clouds and 3D-AR overlays were successfully created for all 5 patients. The 3D point clouds were accurate, with a median error of −0.92 mm, and the LRE was 5.12 mm. The time for creating the 3D-AR overlay was approximately 5 min. Besides creating the 3D-AR overlays, we could visualize the models directly within the robotic console during the surgical procedure.
Conclusions:
We present an algorithm for generating accurate semiautomatic 3D-AR overlays, visualizing essential anatomical structures during robot-assisted mitral valve repair. This may lead to automated intraoperative 3D-AR vision during robotic cardiac surgery, with the potential of increasing safety, accuracy, and efficiency.
Along with the recent development of deep neural networks, appearance-based gaze estimation has succeeded considerably when training and testing within the same domain. Compared to the within-domain task, the variance of different domains makes the cross-domain performance drop severely, preventing gaze estimation deployment in real-world applications. Among all the factors, the ranges of head pose and gaze are believed to play significant roles in the final performance of gaze estimation, while collecting large ranges of data is expensive. This work proposes an effective model training pipeline consisting of a training data synthesis and a gaze estimation model for unsupervised domain adaptation. The proposed data synthesis leverages the single-image 3D reconstruction to expand the range of the head poses from the source domain without requiring a 3D facial shape dataset. To bridge the inevitable gap between synthetic and real images, we further propose an unsupervised domain adaptation method suitable for synthetic full-face data. We propose a disentangling autoencoder network to separate gaze-related features and introduce background augmentation consistency loss to utilize the characteristics of the synthetic source domain. Through comprehensive experiments, it shows that the model using only our synthetic training data can perform comparably to real data extended with a large label range. Our proposed domain adaptation approach further improves the performance on multiple target domains.
Human intention detection with hand motion prediction is critical to drive the upper-extremity assistive robots in neurorehabilitation applications. However, the traditional methods relying on physiological signal measurement are restrictive and often lack environmental context. We propose a novel approach that predicts future sequences of both hand poses and joint positions. This method integrates gaze information, historical hand motion sequences, and environmental object data, adapting dynamically to the assistive needs of the patient without prior knowledge of the intended object for grasping. Specifically, we use a vector-quantized variational autoencoder for robust hand pose encoding with an autoregressive generative transformer for effective hand motion sequence prediction. We demonstrate the usability of these novel techniques in a pilot study with healthy subjects. To train and evaluate the proposed method, we collect a dataset consisting of various types of grasp actions on different objects from multiple subjects. Through extensive experiments, we demonstrate that the proposed method can successfully predict sequential hand movement. Especially, the gaze information shows significant enhancements in prediction capabilities, particularly with fewer input frames, highlighting the potential of the proposed method for real-world applications.
SpeechCAT
Cross-Attentive Transformer for Audio to Motion Generation
Purpose : Stroke remains a leading cause of morbidity and mortality worldwide, despite advances in treatment modalities. Endovascular thrombectomy (EVT), a revolutionary intervention for ischemic stroke, is limited by its reliance on 2D fluoroscopic imaging, which lacks depth and comprehensive vascular detail. We propose a novel AI-driven pipeline for 3D CTA to 2D DSA cross-modality registration, termed DeepIterReg. Methods : The proposed pipeline integrates neural network-based initialization with iterative optimization to align pre-intervention and peri-intervention data. Our approach addresses the challenges of cross-modality alignment, particularly in scenarios involving limited shared vascular structures, by leveraging synthetic data, vein-centric anchoring, and differentiable rendering techniques. Results : We assess the efficacy of DeepIterReg through quantitative analysis of capture ranges and registration accuracy. Results show that our method can accurately register 70% of a test set of 20 patients and can improve capture ranges when performing an initial pose estimation using a convolutional neural network. Conclusions : DeepIterReg demonstrates promising performance for 3D-to-2D stroke intervention image registration, potentially aiding clinicians by improving spatial understanding during EVT and reducing dependence on manual adjustments.
SBM
Social Behavior Model for Human-Like Action Generation
Through the Eyes of Emotion
A Multi-faceted Eye Tracking Dataset for Emotion Recognition in Virtual Reality
PrivateGaze
Preserving User Privacy in Black-box Mobile Gaze Tracking Services
Eye gaze contains rich information about human attention and cognitive processes. This capability makes the underlying technology, known as gaze tracking, a critical enabler for many ubiquitous applications and has triggered the development of easy-to-use gaze estimation services. Indeed, by utilizing the ubiquitous cameras on tablets and smartphones, users can readily access many gaze estimation services. In using these services, users must provide their full-face images to the gaze estimator, which is often a black box. This poses significant privacy threats to the users, especially when a malicious service provider gathers a large collection of face images to classify sensitive user attributes. In this work, we present PrivateGaze, the first approach that can effectively preserve users’ privacy in black-box gaze tracking services without compromising gaze estimation performance. Specifically, we proposed a novel framework to train a privacy preserver that converts full-face images into obfuscated counterparts, which are effective for gaze estimation while containing no privacy information. Evaluation on four datasets shows that the obfuscated image can protect users’ private information, such as identity and gender, against unauthorized attribute classification. Meanwhile, when used directly by the black-box gaze estimator as inputs, the obfuscated images lead to comparable tracking performance to the conventional, unprotected full-face images.
genome-wide association study totaling 111,326 clinically diagnosed/‘proxy’ AD cases and 677,663 controls. We found 75 risk loci, of which 42 were new at the time of analysis. Pathway enrichment analyses confirmed the involvement of amyloid/tau pathways and highlighted microglia implication. Gene prioritization in the new loci identified 31 genes that were suggestive of new genetically associated processes, including the tumor necrosis factor alpha pathway through the linear ubiquitin chain assembly complex. We also built a new genetic risk score associated with the risk of future AD/dementia or progression from mild cognitive impairment to AD/dementia. The improvement in prediction led to a 1.6- to 1.9-fold increase in AD risk from the lowest to the highest decile, in addition to effects of age and the APOE ε4 allele. ...
genome-wide association study totaling 111,326 clinically diagnosed/‘proxy’ AD cases and 677,663 controls. We found 75 risk loci, of which 42 were new at the time of analysis. Pathway enrichment analyses confirmed the involvement of amyloid/tau pathways and highlighted microglia implication. Gene prioritization in the new loci identified 31 genes that were suggestive of new genetically associated processes, including the tumor necrosis factor alpha pathway through the linear ubiquitin chain assembly complex. We also built a new genetic risk score associated with the risk of future AD/dementia or progression from mild cognitive impairment to AD/dementia. The improvement in prediction led to a 1.6- to 1.9-fold increase in AD risk from the lowest to the highest decile, in addition to effects of age and the APOE ε4 allele.