Research and Publications

Powering discovery across genomics, imaging, and AI.

A single-cell, long-read, isoform-resolved case-control study of FTD reveals cell-type-specific and broad splicing dysregulation in human brain

Progranulin-deficient frontotemporal dementia (GRN-FTD) is a major cause of familial FTD with TAR DNA-binding protein 43 (TDP-43) pathology, which is linked to exon dysregulation. However, little is known about this dysregulation in glial and neuronal cells. Here, using splice-junction-covering enrichment probes, we introduce single-nuclei long-read RNA sequencing 2 (SnISOr-Seq2), targeting 3,630 high-interest genes without loss of precision, and complete the first single-cell, long-read-resolved case-control study for neurodegeneration. Exons affected by FTD-associated skipping are shorter than those whose inclusion is increased. Up to 30% of cell-(sub)type-specific splicing dysregulation is masked by other cell types or cortical layers. Surprisingly, strong splicing dysregulation events can occur in select but not all cell types. In some cases, a cell type switches in FTD to the splicing pattern of a different cell type. In addition, in separate GRN-FTD samples, the more FTD-prone frontal cortex exhibits more FTD-associated splicing patterns than the occipital cortex. Our methodologies are widely applicable to brain and other diseases.

RNA Isoform Computational Resources
View details

Generation of Synthetic Laryngoscopy Images Using StyleGAN3

This study presents the first reported generation of high-fidelity, age-stratified synthetic videoendoscopic images of the larynx using StyleGAN3. A total of 43 healthy female laryngeal videostroboscopy exams were divided into Dataset A (≥65 years, n = 33, 676 frames) and Dataset B (<65 years, n = 10, 4572 frames) based on differences in aging physiologic characteristics. Images were extracted using a deep-learning-based informative-frame classifier, manually quality-verified, and preprocessed by converting JPEGs to PNGs and resizing to 512×512 pixels. StyleGAN3 models were trained for each cohort on 6-8 GPUs for 5-9 days (target 25 000 kimg, ≈ thousand-image units per model). Each model generated 100 synthetic images evaluated using Fréchet Inception Distance (FID) and Multi-Scale Structural Similarity Index Measure (MS-SSIM). Dataset B achieved lower FID (13.22) than Dataset A (25.26), demonstrating improved fidelity with larger training sets. Dataset A revealed a strong inverse fidelity-diversity relationship (Pearson r = −0.78) and evidence of mode collapse, whereas Dataset B demonstrated stable generalization (partial r = 0.47). These results demonstrate the feasibility of age-stratified synthetic laryngeal image generation and underscore the necessity of larger datasets for maintaining both fidelity and diversity. Synthetic laryngoscopy images may enhance training, benchmarking, and dataset expansion in medical imaging research.

Computational Modeling Training
View details

Collection of biospecimens from the inspiration4 mission establishes the standards for the space omics and medical atlas (SOMA)

The SpaceX Inspiration4 mission provided a unique opportunity to study the impact of spaceflight on the human body. Biospecimen samples were collected from four crew members longitudinally before (Launch: L-92, L-44, L-3 days), during (Flight Day: FD1, FD2, FD3), and after (Return: R + 1, R + 45, R + 82, R + 194 days) spaceflight, spanning a total of 289 days across 2021-2022. The collection process included venous whole blood, capillary dried blood spot cards, saliva, urine, stool, body swabs, capsule swabs, SpaceX Dragon capsule HEPA filter, and skin biopsies. Venous whole blood was further processed to obtain aliquots of serum, plasma, extracellular vesicles and particles, and peripheral blood mononuclear cells. In total, 2,911 sample aliquots were shipped to our central lab at Weill Cornell Medicine for downstream assays and biobanking. This paper provides an overview of the extensive biospecimen collection and highlights their processing procedures and long-term biobanking techniques, facilitating future molecular tests and evaluations.As such, this study details a robust framework for obtaining and preserving high-quality human, microbial, and environmental samples for aerospace medicine in the Space Omics and Medical Atlas (SOMA) initiative, which can aid future human spaceflight and space biology experiments.

View details

Discovery of therapeutic targets in cancer using chromatin accessibility and transcriptomic data

Most cancer types lack targeted therapeutic options, and when first-line targeted therapies are available, treatment resistance is a huge challenge. Recent technological advances enable the use of assay for transposase-accessible chromatin with sequencing (ATAC-seq) and RNA sequencing (RNA-seq) on patient tissue in a high-throughput manner. Here, we present a computational approach that leverages these datasets to identify drug targets based on tumor lineage. We constructed gene regulatory networks for 371 patients of 22 cancer types using machine learning approaches trained with three-dimensional genomic data for enhancer-to-promoter contacts. Next, we identified the key transcription factors (TFs) in these networks, which are used to find therapeutic vulnerabilities, by direct targeting of either TFs or the proteins that they interact with. We validated four candidates identified for neuroendocrine, liver, and renal cancers, which have a dismal prognosis with current therapeutic options.

Machine learning RNA-seq data
View details

Rethinking clinical trials for medical AI with dynamic deployments of adaptive systems

There is a growing recognition of the need for clinical trials to safely and effectively deploy artificial intelligence (AI) in clinical settings. We introduce dynamic deployment as a framework for AI clinical trials tailored for the dynamic nature of large language models, making possible complex medical AI systems which continuously learn and adapt in situ from new data and interactions with users while enabling continuous real-time monitoring and clinical validation.

AI
View details

Statistical variability in comparing accuracy of neuroimaging based classification models via cross validation

Machine learning (ML) has significantly transformed biomedical research, leading to a growing interest in model development to advance classification accuracy in various clinical applications. However, this progress raises essential questions regarding how to rigorously compare the accuracy of different ML models. In this study, we highlight the practical challenges in quantifying the statistical significance of accuracy differences between two neuroimaging-based classification models when cross-validation (CV) is performed. Specifically, we propose an unbiased framework to assess the impact of CV setups (e.g., the number of folds) on the statistical significance. We apply this framework to three publicly available neuroimaging datasets to re-emphasize known flaws in current computation of p-values for comparing model accuracies. We further demonstrate that the likelihood of detecting significant differences among models varies substantially with the intrinsic properties of the data, testing procedures, and CV configurations of choice. Given that many of the above factors do not typically fall into the evaluation criteria of ML-based biomedical studies, we argue that such variability can potentially lead to p-hacking and inconsistent conclusions on model improvement. The obtained results from this study underscore that more rigorous practices in model comparison are urgently needed in order to mitigate the reproducibility crisis in biomedical ML research.

Machine learning
View details