TL;DR: This work proposes a method called Local Augment (LA), designed to improve the performance of trained outlier detectors at the prediction stage without altering the trained models or accessing the training data.
In this work, the concept of test-time learning is presented, wherein Machine-Learning (ML) models are constructed by involving unlabeled test samples. Based on this concept, we propose a method called Local Augment (LA) designed to improve the performance of trained outlier dete…
Current identity customized video generation methodologies are predominantly limited to single-identity scenarios, as the lack of explicit identity separation mechanisms often leads to identity confusion in multi-identity settings. Existing multi-identity approaches, which direct…
Composition is a high-level visual intent that governs where subjects are placed and how a scene is organized, yet current unified multimodal models remain unreliable at fine-grained composition recognition and struggle to turn such intent into controllable generation. We present…