Data used in "Multi-omics integration and batch correction using a modality-agnostic deep learning framework"
Jose Ignacio Alvira Larizgoitia, Gabriele Partel, Jelle Jacobs, Alejandro Sifrim
These are multimodal dataset objects and trained model parameters used in the study. The files are organized in pairs, where each multimodal dataset (.h5mu file) corresponds to a trained model parameter file (.pt) generated using the MIMA (Multimodal Integration with Modality-agnostic Autoencoders) framework. Each .h5mu file stores multiple modalities along with their corresponding metadata and annotations, serving as input to the MIMA model. The .pt files contain the trained PyTorch model parameters obtained from training MIMA on the respective dataset. Contents VISIUM_MSI_Morphology_prostate_cancer.h5muMultimodal dataset combining spatial transcriptomics (Visium), mass spectrometry imaging (MSI), and histological morphology data from prostate cancer samples. Includes global and modality-specific metadata. mima_params_prostate_cancer.ptTrained MIMA model parameters corresponding to the prostate cancer dataset above. RNA_ATAC_ATAC5kbmerge.h5muMultimodal single-cell dataset integrating RNA expression and chromatin accessibility (ATAC-seq) data. There is a second ATAC modality representing the 5 kb merged peak matrix. Includes global and modality-specific metadata. mima_params_multiome5kbmerge.ptTrained MIMA model parameters corresponding to the RNA–ATAC5kbmerge multimodal dataset. scCITEseq.h5muSingle-cell CITE-seq dataset integrating RNA expression and surface protein measurements, with relevant cell- and modality-level metadata. mima_params_cite.ptTrained MIMA model parameters corresponding to the scCITE-seq dataset. File formats .h5mu: Multimodal data container (MuData format). .pt: PyTorch serialized model parameters.
Also available via: European Organization for Nuclear Research