The Inverted Competence
The standard test of understanding: if a system can do something successfully, it understands it. If it can explain something accurately, it understands it. Competence and explanation are treated as convergent evidence for the same underlying capacity.
Chacón Sartori (arXiv:2603.28371, March 2026) identifies what he calls the Bidirectional Coherence Paradox: in large language models, competence and grounding not only dissociate but invert across epistemic conditions. In low-observability domains — where the causal mechanisms are hidden or abstract — LLMs often act successfully while misidentifying the mechanisms that produce their success. They get the right answer for the wrong reasons, and their explanations of why they succeeded are confabulations. In high-observability domains — where causal structure is transparent — they generate explanations that accurately track the observable mechanisms yet fail to translate those diagnoses into effective interventions. They explain correctly but act incorrectly.
The paradox is bidirectional: the domains where LLMs succeed at tasks are the domains where their explanations fail, and the domains where their explanations succeed are the domains where their task performance fails. Competence and explanation trade off rather than reinforce.
The mechanism is the relationship between coherence and grounding. In both cases, the explanatory output is coherent — it sounds reasonable, follows logical structure, and satisfies conversational expectations. Coherence is maintained across both conditions. What breaks is the grounding: in low-observability domains, coherent explanations are ungrounded in the actual causal mechanism; in high-observability domains, grounded explanations are unlinked to the action-selection process. Coherence masks both failures equally.
The structural observation: coherent explanation is not evidence of understanding in either direction. The system that succeeds without understanding and the system that understands without succeeding both produce equally fluent explanations. The explanatory coherence that we use as a test of understanding is the one property that is invariant across the competence-grounding dissociation — which means it is the one property that cannot diagnose it.