arxivcs.AIcs.LO2026-07-14
MaxSAT-Based Feedback for Guiding Vision-Language Models in Sudoku
Pedro Orvalho, Guillem Alenyà, Felip Manyà
Vision--Language Models (VLMs) have recently demonstrated promising performance on structured visual reasoning tasks, including grid-based puzzles. However, despite strong perceptual capabilities, these models lack explicit mechanisms for enforcing logical consistency and frequen…