arxivcs.CV2026-07-03
Learning to Generate Multiple Objects from Dense and Occluded Layouts
Bach-Hoang Ngo, Si-Tri Ngo, Hieu Le, Trung-Nghia Le
Text-to-image diffusion models fail to generate correct object counts in dense scenes, where overlapping instances collapse into indistinguishable structures despite appearing visually plausible. We identify this as instance ownership collapse: tokens from overlapping objects int…