<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Code Understanding | Pedro Orvalho</title><link>https://pmorvalho.github.io/tags/code-understanding/</link><atom:link href="https://pmorvalho.github.io/tags/code-understanding/index.xml" rel="self" type="application/rss+xml"/><description>Code Understanding</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Sat, 18 Jul 2026 00:00:00 +0000</lastBuildDate><image><url>https://pmorvalho.github.io/media/icon_hu_449091aa0565028d.png</url><title>Code Understanding</title><link>https://pmorvalho.github.io/tags/code-understanding/</link></image><item><title>📄🤖 2 Papers accepted @ EPIA 2026!! 🎉🎉</title><link>https://pmorvalho.github.io/blog/2026-07-18-epia/</link><pubDate>Sat, 18 Jul 2026 00:00:00 +0000</pubDate><guid>https://pmorvalho.github.io/blog/2026-07-18-epia/</guid><description>&lt;p&gt;I&amp;rsquo;m very excited to share that &lt;strong&gt;two of our papers have been accepted at the 25th EPIA Conference on Artificial Intelligence (
)&lt;/strong&gt;! 🎉&lt;/p&gt;
&lt;p&gt;The accepted papers explore two complementary aspects of trustworthy AI: &lt;strong&gt;understanding the reasoning capabilities of Large Language Models for code&lt;/strong&gt; and &lt;strong&gt;improving Vision-Language Models through symbolic optimisation with MaxSAT&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id="-large-language-models-are-not-yet-robust-in-understanding-code-against-semantics-preserving-mutations"&gt;💻 Large Language Models Are Not (Yet) Robust in Understanding Code Against Semantics-Preserving Mutations&lt;/h2&gt;
&lt;p&gt;As LLMs become increasingly popular for programming assistance, an important question remains: &lt;strong&gt;do they truly understand code, or are they often relying on superficial patterns?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;In this paper, we investigate the robustness of state-of-the-art LLMs by applying a collection of &lt;strong&gt;semantics-preserving Python code transformations&lt;/strong&gt;, including variable renaming, mirrored comparisons, branch swapping, loop transformations, and loop unrolling. Although these mutations leave program behaviour unchanged, they frequently cause LLMs to alter their predictions.&lt;/p&gt;
&lt;p&gt;Our study, combining benchmark evaluation with expert analysis, shows that even the strongest proprietary models remain surprisingly fragile. We find that code-specialised LLMs can produce correct answers based on flawed reasoning in &lt;strong&gt;up to 45% of cases&lt;/strong&gt;, while performance under semantics-preserving mutations can decrease by &lt;strong&gt;as much as 70%&lt;/strong&gt;. These results highlight that current LLMs have not yet achieved stable, semantically grounded code understanding.&lt;/p&gt;
&lt;div class="text-center"&gt;
&lt;a
id="button-e5692fbea4520de871edcb0a8ea66f19"
href="https://pmorvalho.github.io/publications/epia2026-1"
target="_blank"
rel="noopener"
class="inline-flex items-center gap-2 font-medium no-underline transition-all duration-300 ease-out transform-gpu focus:outline-none focus:ring-4 focus:ring-offset-2 focus:ring-offset-white dark:focus:ring-offset-zinc-900 disabled:opacity-50 disabled:cursor-not-allowed disabled:pointer-events-none bg-gradient-to-br from-secondary-500 to-secondary-600 hover:from-secondary-600 hover:to-secondary-700 active:from-secondary-700 active:to-secondary-800 text-white shadow-lg shadow-secondary-500/25 hover:shadow-xl hover:shadow-secondary-500/30 hover:scale-105 active:scale-95 focus:ring-secondary-500/50 px-4 py-2 text-base rounded-full"
role="button"
aria-label="Read the paper"
&gt;
&lt;span&gt;Read the paper&lt;/span&gt;
&lt;/a&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;div class="pub-list-item view-citation" style="margin-bottom: 1rem"&gt;
&lt;i class="far fa-file-alt pub-icon" aria-hidden="true"&gt;&lt;/i&gt;
&lt;span class="article-metadata li-cite-author"&gt;
&lt;span class="font-bold"&gt;
Pedro Orvalho&lt;/span&gt;, &lt;span &gt;
Marta Kwiatkowska&lt;/span&gt;
&lt;/span&gt;
(2026).
&lt;a href="https://pmorvalho.github.io/publications/epia2026-1/" class="underline"&gt;Large Language Models Are Not (Yet) Robust in Understanding Code Against Semantics-Preserving Mutations&lt;/a&gt;.
In &lt;strong&gt;EPIA 2026&lt;/strong&gt;.
&lt;div class="flex flex-wrap space-x-3"&gt;
&lt;a class="hb-attachment-link hb-attachment-link-small" href="https://pmorvalho.github.io/uploads/papers/epia2026-LLMCs-Semantic-Robustness.pdf" &gt;
&lt;svg style="height: 1em" class='inline-block' xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="M19.5 14.25v-2.625a3.375 3.375 0 0 0-3.375-3.375h-1.5A1.125 1.125 0 0 1 13.5 7.125v-1.5a3.375 3.375 0 0 0-3.375-3.375H8.25m0 12.75h7.5m-7.5 3H12M10.5 2.25H5.625c-.621 0-1.125.504-1.125 1.125v17.25c0 .621.504 1.125 1.125 1.125h12.75c.621 0 1.125-.504 1.125-1.125V11.25a9 9 0 0 0-9-9"/&gt;&lt;/svg&gt;
PDF
&lt;/a&gt;
&lt;a class="hb-attachment-link hb-attachment-link-small" href="https://pmorvalho.github.io/projects/ai4code" &gt;
&lt;svg style="height: 1em" class='inline-block' xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="M13.19 8.688a4.5 4.5 0 0 1 1.242 7.244l-4.5 4.5a4.5 4.5 0 0 1-6.364-6.364l1.757-1.757m13.35-.622l1.757-1.757a4.5 4.5 0 0 0-6.364-6.364l-4.5 4.5a4.5 4.5 0 0 0 1.242 7.244"/&gt;&lt;/svg&gt;
project
&lt;/a&gt;
&lt;a class="hb-attachment-link hb-attachment-link-small" href="https://github.com/pmorvalho/EPIA26-LLMCs-Semantic-Robustness" target="_blank" rel="noopener"&gt;
&lt;svg style="height: 1em" class='inline-block' xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="M13.19 8.688a4.5 4.5 0 0 1 1.242 7.244l-4.5 4.5a4.5 4.5 0 0 1-6.364-6.364l1.757-1.757m13.35-.622l1.757-1.757a4.5 4.5 0 0 0-6.364-6.364l-4.5 4.5a4.5 4.5 0 0 0 1.242 7.244"/&gt;&lt;/svg&gt;
GitHub
&lt;/a&gt;
&lt;button class="hb-attachment-link hb-attachment-link-small js-cite-clipboard cursor-pointer" type="button" data-filename="/publications/epia2026-1/cite.bib"&gt;
&lt;svg style="height: 1em" class='inline-block' xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="M15.75 17.25v3.375c0 .621-.504 1.125-1.125 1.125h-9.75a1.125 1.125 0 0 1-1.125-1.125V7.875c0-.621.504-1.125 1.125-1.125H6.75a9.06 9.06 0 0 1 1.5.124m7.5 10.376h3.375c.621 0 1.125-.504 1.125-1.125V11.25c0-4.46-3.243-8.161-7.5-8.876a9.06 9.06 0 0 0-1.5-.124H9.375c-.621 0-1.125.504-1.125 1.125v3.5m7.5 10.375H9.375a1.125 1.125 0 0 1-1.125-1.125v-9.25m12 6.625v-1.875a3.375 3.375 0 0 0-3.375-3.375h-1.5a1.125 1.125 0 0 1-1.125-1.125v-1.5a3.375 3.375 0 0 0-3.375-3.375H9.75"/&gt;&lt;/svg&gt;
&lt;span&gt;Cite&lt;/span&gt;
&lt;/button&gt;
&lt;a class="hb-attachment-link hb-attachment-link-small" href="https://arxiv.org/abs/2505.10443" target="_blank" rel="noopener"&gt;
&lt;svg style="height: 1em" class='inline-block' xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="M19.5 14.25v-2.625a3.375 3.375 0 0 0-3.375-3.375h-1.5A1.125 1.125 0 0 1 13.5 7.125v-1.5a3.375 3.375 0 0 0-3.375-3.375H8.25m0 12.75h7.5m-7.5 3H12M10.5 2.25H5.625c-.621 0-1.125.504-1.125 1.125v17.25c0 .621.504 1.125 1.125 1.125h12.75c.621 0 1.125-.504 1.125-1.125V11.25a9 9 0 0 0-9-9"/&gt;&lt;/svg&gt;
Preprint
&lt;/a&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="-maxsat-based-feedback-for-guiding-vision-language-models-in-sudoku"&gt;🧩 MaxSAT-Based Feedback for Guiding Vision-Language Models in Sudoku&lt;/h2&gt;
&lt;p&gt;Can symbolic reasoning help Vision-Language Models solve structured reasoning problems more reliably?&lt;/p&gt;
&lt;p&gt;In this paper, we present a &lt;strong&gt;neuro-symbolic framework&lt;/strong&gt; that combines Vision-Language Models with a &lt;strong&gt;Maximum Satisfiability (MaxSAT)&lt;/strong&gt; solver for Sudoku. Rather than replacing the VLM, the symbolic component acts as a logical consistency checker, identifying conflicting assignments and generating structured textual and visual feedback that guides the model towards better solutions.&lt;/p&gt;
&lt;p&gt;By integrating formal optimisation into the reasoning loop, our approach significantly improves logical consistency and increases the number of correctly solved Sudoku instances across multiple open-source and proprietary VLMs. The results demonstrate the potential of symbolic optimisation as an effective mechanism for enhancing the reliability of multimodal AI systems.&lt;/p&gt;
&lt;div class="text-center"&gt;
&lt;a
id="button-08f46ef5b834729f5bb3c89233d3a366"
href="https://pmorvalho.github.io/publications/epia2026-2"
target="_blank"
rel="noopener"
class="inline-flex items-center gap-2 font-medium no-underline transition-all duration-300 ease-out transform-gpu focus:outline-none focus:ring-4 focus:ring-offset-2 focus:ring-offset-white dark:focus:ring-offset-zinc-900 disabled:opacity-50 disabled:cursor-not-allowed disabled:pointer-events-none bg-gradient-to-br from-secondary-500 to-secondary-600 hover:from-secondary-600 hover:to-secondary-700 active:from-secondary-700 active:to-secondary-800 text-white shadow-lg shadow-secondary-500/25 hover:shadow-xl hover:shadow-secondary-500/30 hover:scale-105 active:scale-95 focus:ring-secondary-500/50 px-4 py-2 text-base rounded-full"
role="button"
aria-label="Read the paper"
&gt;
&lt;span&gt;Read the paper&lt;/span&gt;
&lt;/a&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;div class="pub-list-item view-citation" style="margin-bottom: 1rem"&gt;
&lt;i class="far fa-file-alt pub-icon" aria-hidden="true"&gt;&lt;/i&gt;
&lt;span class="article-metadata li-cite-author"&gt;
&lt;span class="font-bold"&gt;
Pedro Orvalho&lt;/span&gt;, &lt;span &gt;
Guillem Alenyà&lt;/span&gt;, &lt;span &gt;
Felip Manyà&lt;/span&gt;
&lt;/span&gt;
(2026).
&lt;a href="https://pmorvalho.github.io/publications/epia2026-2/" class="underline"&gt;MaxSAT-Based Feedback for Guiding Vision-Language Models in Sudoku&lt;/a&gt;.
In &lt;strong&gt;EPIA 2026&lt;/strong&gt;.
&lt;div class="flex flex-wrap space-x-3"&gt;
&lt;a class="hb-attachment-link hb-attachment-link-small" href="https://pmorvalho.github.io/uploads/papers/epia2026-MaxSAT-VLMs-Sudoku.pdf" &gt;
&lt;svg style="height: 1em" class='inline-block' xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="M19.5 14.25v-2.625a3.375 3.375 0 0 0-3.375-3.375h-1.5A1.125 1.125 0 0 1 13.5 7.125v-1.5a3.375 3.375 0 0 0-3.375-3.375H8.25m0 12.75h7.5m-7.5 3H12M10.5 2.25H5.625c-.621 0-1.125.504-1.125 1.125v17.25c0 .621.504 1.125 1.125 1.125h12.75c.621 0 1.125-.504 1.125-1.125V11.25a9 9 0 0 0-9-9"/&gt;&lt;/svg&gt;
PDF
&lt;/a&gt;
&lt;button class="hb-attachment-link hb-attachment-link-small js-cite-clipboard cursor-pointer" type="button" data-filename="/publications/epia2026-2/cite.bib"&gt;
&lt;svg style="height: 1em" class='inline-block' xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="M15.75 17.25v3.375c0 .621-.504 1.125-1.125 1.125h-9.75a1.125 1.125 0 0 1-1.125-1.125V7.875c0-.621.504-1.125 1.125-1.125H6.75a9.06 9.06 0 0 1 1.5.124m7.5 10.376h3.375c.621 0 1.125-.504 1.125-1.125V11.25c0-4.46-3.243-8.161-7.5-8.876a9.06 9.06 0 0 0-1.5-.124H9.375c-.621 0-1.125.504-1.125 1.125v3.5m7.5 10.375H9.375a1.125 1.125 0 0 1-1.125-1.125v-9.25m12 6.625v-1.875a3.375 3.375 0 0 0-3.375-3.375h-1.5a1.125 1.125 0 0 1-1.125-1.125v-1.5a3.375 3.375 0 0 0-3.375-3.375H9.75"/&gt;&lt;/svg&gt;
&lt;span&gt;Cite&lt;/span&gt;
&lt;/button&gt;
&lt;a class="hb-attachment-link hb-attachment-link-small" href="https://arxiv.org/abs/2607.12711" target="_blank" rel="noopener"&gt;
&lt;svg style="height: 1em" class='inline-block' xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="M19.5 14.25v-2.625a3.375 3.375 0 0 0-3.375-3.375h-1.5A1.125 1.125 0 0 1 13.5 7.125v-1.5a3.375 3.375 0 0 0-3.375-3.375H8.25m0 12.75h7.5m-7.5 3H12M10.5 2.25H5.625c-.621 0-1.125.504-1.125 1.125v17.25c0 .621.504 1.125 1.125 1.125h12.75c.621 0 1.125-.504 1.125-1.125V11.25a9 9 0 0 0-9-9"/&gt;&lt;/svg&gt;
Preprint
&lt;/a&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;I&amp;rsquo;m very grateful to my collaborators for making these works possible!&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;See you in Madeira at EPIA 2026!&lt;/strong&gt; 🇵🇹🎉&lt;/p&gt;</description></item></channel></rss>