CORTEXA
← Browse
arxivcs.LG2026-07-09

MatBind: A Shared Embedding Space for Multimodal Materials Characterization

Le Yang, Anoop K. Chandran, Jona Östreicher, Evgenii Sovetkin, Adrian Mirza, Sebastien Bompas, Bashir Kazimi, Pascal Friederich, Stefan Kesselheim, Kevin Maik Jablonka, Stefan Sandfeld

Fully characterizing a crystalline material requires integrating heterogeneous data sources -- atomic structures, diffraction patterns, electronic density of states, and natural language -- each of which captures a different facet of the same physical object. In practice, however, these modalities are stored and analyzed in isolation, making it difficult to relate or query materials across representational boundaries. We present MatBind, a contrastive learning framework that aligns four materials modalities -- crystal structure, powder X-ray diffraction (pXRD) simulated from structures, density of states (DOS), and text -- into a unified embedding space using crystal structure as the central physical anchor. The framework induces alignment between modalities never explicitly paired during training, enabling emergent zero-shot cross-modal retrieval as a direct consequence of the shared representation. The learned embedding space organizes materials according to physically meaningful properties without explicit supervision, and retrieval performance improves systematically when modalities are combined at query time. These results demonstrate that treating heterogeneous materials data as complementary projections of a single physical reality, rather than as isolated data sources, is not a practical choice but is consistent with the underlying physics.

View free PDFSource page

Related papers

arxivcs.LGcs.AI2026-07-03

Cross-Subject Semantic Decoding with Shared-Space Alignment for Generalized Neural Representation Learning

Ji-Hoon Heo, Aleksandra Joanna Wisniewska, Seo-Hyun Lee, Seong-Whan Lee

Generalizing across subjects remains challenging in invasive neural recordings because electrode configurations, anatomical structures, and neural signal patterns vary substantially across individuals. To investigate such inter-subject variability, we propose a cross-subject sema…

View free PDFSource page
arxivcs.LG2026-07-24

LunarFM: A Shared Multimodal Representation of the Moon's Surface

Marc Girona-Mata, Jakob Gawlikowski, Sumit Goski, Gautier Bardi de Fourtou, Valentin T. Bickel, Ben Moseley, et al.

The renewed global focus on lunar exploration, driven by the prospect of in-situ resource utilization and a sustained human presence on the Moon, has created growing demand for accurate, large-scale characterization of the lunar surface. Although vast quantities of orbital remote…

View free PDFSource page
arxivcond-mat.mtrl-scics.AIcs.LG2026-07-22

Generative and multimodal AI for materials prediction and design: Progress, challenges, and perspectives

Xianyuan Liu, Charles Anjah, Benjamin E. Jolly, Jonathon F. S. Markanday, Joshua Berry, Haolin Wang, et al.

Artificial intelligence (AI) is accelerating materials prediction and design by enabling efficient exploration of chemical and structural spaces, with particular promise for novel materials discovery. However, novelty in materials discovery encompasses chemical plausibility, stru…

View free PDFSource page
arxivcs.CRcs.AIcs.LG2026-06-25

On the Inseparability of Instructions and Data in Shared-Embedding Sequence Models

Dewank Pant, Shruti Lohani, Avijit Kumar

Prompt injection is the top security risk for LLM-integrated applications, yet every defense proposed so far has been broken. We prove this is not a coincidence: in shared-embedding architectures that lack enforced control-data separation, perfect prompt-injection prevention is m…

View free PDFSource page
arxivcs.CLcs.LG2026-07-11

One mechanism for many mental spaces: a shared router over a value slot in language models

Oliver Steele, Jiangtao Wen, Yuxing Han

Language builds discourse contexts other than the actual: a painting, a belief, a memory, a hypothetical. Each is a mental space in which the same entity can take a different value, as when a flower is red in reality but purple in a portrait. Formal semantics keeps these contexts…

View free PDFSource page
arxivcs.CVcs.LG2026-07-16

AlphaWiSE: Adaptive Weight Interpolation for Continual Multimodal Representation Learning

Sarthak Jain, Qiran Hu, Zhen Zhu, Yaoyao Liu

Multimodal models such as CLIP learn a shared embedding space for cross-modal retrieval, but continual adaptation to sequentially arriving data can disrupt the cross-modal alignment acquired from earlier phases. Conventional continual-learning methods return a single checkpoint,…

View free PDFSource page