arxivcs.CVcs.AI2026-07-08
Tree-of-Thoughts Reasoning for Text-to-Image In-Context Learning
Stepanida Alekseeva, Jenifer Kalafatovich, Seong-Whan Lee
In text-to-image in-context learning (T2I-ICL), a model has to infer a latent compositional pattern from fewshot demonstrations for generating a query image. Recent studies show that state-of-the-art multimodal large language models struggle with this setting, particularly due to…