MaxSAT-Based Feedback for Guiding Vision-Language Models in Sudoku @ EPIA 2026

Sep 2, 2026·
Pedro Orvalho
Pedro Orvalho
· 0 min read
Image credit: EPIA
Abstract
Vision-Language Models (VLMs) have demonstrated impressive capabilities in interpreting visual information and tackling complex reasoning tasks. However, their predictions are not guaranteed to satisfy the logical constraints of structured problems, which can lead to inconsistent or incorrect solutions. In this talk, we will present a neuro-symbolic approach that combines the perceptual and generative capabilities of VLMs with Maximum Satisfiability (MaxSAT) reasoning to improve the reliability of their outputs. Using Sudoku as a controlled visual reasoning benchmark, we introduce a MaxSAT oracle that verifies VLM-generated assignments against a formal encoding of the puzzle constraints. Rather than simply solving the problem for the model, the oracle identifies inconsistent predictions and provides targeted feedback that the VLM can use to revise its solution. We investigate different ways of integrating this feedback into the reasoning process, including step-by-step solving and full-board refinement. We evaluate the approach across multiple VLMs and Sudoku difficulty levels, showing that MaxSAT-based feedback can improve logical consistency, solution completeness, and overall solving performance. The improvements are particularly notable when refining candidate solutions that are already close to being correct. More broadly, our results illustrate the benefits of combining neural models with symbolic reasoning, highlighting how formal methods can provide structured, correctness-aware feedback to help build more reliable visual reasoning systems.
Date
Sep 2, 2026 11:00 AM — 11:25 AM
Event
Location

EPIA 2026

Funchal, Madeira, Portugal.