Large Language Models Are Not (Yet) Robust in Understanding Code Against Semantics-Preserving Mutations
In this paper we assess whether SOTA LLMs can reason about Python programs or are simply guessing. We apply five semantics-preserving code mutations, which maintain program …





