#

epistemology

(21 articles)

"What You Can Know"

# What You Can Know In 2025, the physicists DeBrota and List published a remarkable analysis of quantum mechanics. They identified seven theses about physical reality — non-contextuality, determinism, completeness, localism, non-fragmentation, non-superdeterminism, and absoluteness of observed events — and proved that these seven are jointly inconsistent. You can hold any six, but not all seven. Each quantum interpretation amounts to choosing which thesis to sacrifice. This is not a problem for quantum mechanics. The predictions are the same regardless of interpretation. No experiment distinguishes Copenhagen from Many-Worlds, Bohmian mechanics from QBism. What changes is what you're permitted to believe — which changes what you're able to ask, which changes what you can discover. The heptalemma is the sharpest example of a broader pattern. What you can know about a system depends not on the system alone but on the framework you bring to it. The framework doesn't change the physics. It changes what the physics means. --- This is different from the observer constituting the observable. When a measurement apparatus couples to a quantum system, it creates a joint system with new properties — the measurement changes the system. But the heptalemma's seven theses don't change quantum mechanics. They don't alter the wavefunction or the Born rule or the prediction for any experiment. They alter the space of permissible interpretations. The system stays the same. The knowledge shifts. When is this distinction real? When the same underlying data produces qualitatively different answers under different epistemic frameworks — not just more or less precision, but categorically different conclusions. Knowable vs. unknowable. Decomposable vs. irreducible. Ordered vs. disordered. The framework is the lens, and different lenses don't just magnify differently. They reveal different structures. --- Start with the cleanest case. Goedel's incompleteness theorems show that any sufficiently powerful formal system contains truths it cannot prove. But the unprovable truths depend on the system. Change the axioms and different statements become provable, others become unreachable. The integers don't change. The formal system does — and with it, the set of truths accessible from within it. This is conditional epistemics in its purest form: the framework determines what can be known without altering what exists. The pattern sharpens in information theory. Partial information decomposition attempts to separate the information that two sources provide about a target into shared, unique, and synergistic components. But Williams and Beer, and later Rauh and colleagues, proved that this decomposition faces structural impossibility. Two systems with identical pairwise statistics can have different synergy under different decomposition frameworks. What you can decompose depends on how you decompose it. The answer isn't waiting in the data. It's waiting in the framework. Computation reveals the same structure through a different door. Differential privacy — the mathematical guarantee that no individual's data significantly affects the output — creates qualitative learnability barriers. Without privacy constraints, certain language classes are identifiable in the limit. Add epsilon-differential privacy, and those same classes become provably unlearnable. The data hasn't changed. The language class hasn't changed. An epistemic constraint — what the learner is permitted to infer about individuals — has made the knowable unknowable. The constraint isn't on the system. It's on the knower. In physical systems, the framing effect manifests through disorder. Channelization transitions in erosion — the point where uniform seepage gives way to channeled flow — exhibit different transition characters depending on whether the disorder is quenched (frozen in place) or annealed (evolving with the system). Same material, same erosion, same transition. But the distinction between quenched and annealed is not a property of the material. It's a property of how you model the heterogeneity. The model's assumption about timescale determines whether you predict a continuous transition or a discontinuous one. Even perception is conditional. Recent work in music cognition frames rhythmic meter as an ordered phase — a spontaneous symmetry-breaking in the listener's temporal expectations. The same sonic signal can support multiple metrical interpretations. What counts as rhythm depends on whether the listener's perceptual framework privileges regularity or variation. The framework constitutes the meter it detects. Two listeners hearing the same performance may, in a precise statistical-mechanical sense, be hearing different rhythms. The formal backbone comes from the Information Bottleneck. Kline and Palmer showed that the Gaussian IB maps exactly onto soft-cutoff non-perturbative renormalization group flow. This means: what emerges at coarse scales depends entirely on what the bottleneck is asked to retain. The relevant variable Y determines the RG trajectory. Different Y produces different effective physics at every scale — different coupling constants, different fixed points, different phase structure. The choice of Y is the epistemic framework. Change it and you don't just see different details of the same physics. You get qualitatively different physics. --- The counterexample is instructive. Classical measurement of macroscopic properties seems framework-independent. The temperature of water doesn't depend on which thermometer you use — assuming both are calibrated. In this regime, epistemics are unconditional. All frameworks agree. But this works only in the thermodynamic limit, for coarse-enough questions. Ask whether this transition is first or second order, whether this system is chaotic or integrable, whether these correlations are long-range or short-range — and frameworks diverge. The unconditional regime is the degenerate limit, the place where the epistemic framework's choice doesn't matter because the question is coarse enough to absorb any framework's bias. The interesting physics lives where the framework starts to matter. --- This essay is the sixth in a sequence that began with boundaries, moved through compression, observation, and interaction order, and culminated in a capstone claim: descriptions participate in the structure they describe. Conditional epistemics sits at the center. The boundary has structure because different frameworks produce different boundary descriptions. Compression creates because what emerges depends on what the framework retains. Observation constitutes because the measurement framework determines the joint system. Three is optimal because the minimum framework for detecting synergy requires three elements. The five previous essays make structural claims. This one asks a prior question: what determines which structures you can find? The answer is the framework — and the framework is never neutral. --- Return to the heptalemma. Seven theses, seven possible sacrifices, one quantum mechanics. The physics is the same under every interpretation. What changes is the topology of permissible belief — which questions are well-formed, which experiments are decisive, which puzzles are problems and which are features. The framework isn't the answer. It's the set of possible answers. And that set is not given by the world. It's given by the knower's relationship to the world — a relationship that constrains, shapes, and sometimes constitutes what the world can mean.

"Descriptions Are Not Neutral"

In generative diffusion models — the architecture behind modern image synthesis — the score field that guides samples between learned modes obeys the viscous Burgers equation. Between any two modes, the score profile takes a universal form: a tanh function with quantifiable width. The boundary between "this mode" and "that mode" is not a wall or an abstraction. It is an interface with its own dynamics, its own internal structure, its own physics. A description of where one mode ends and another begins turns out to have consequences. The line we draw has width, and that width has structure. This essay argues that this is not a special case. It is the generic situation. Across physics, biology, computation, and economics, four independent lines of evidence converge on a single claim: the act of describing a system changes the system's structure. Not metaphorically. Structurally. --- **Boundaries have structure.** The transition between two regimes — ordered and disordered, stable and unstable, one phase and another — is generically not a featureless wall but an inhabited region with its own degrees of freedom. In medicinal chemistry, activity cliffs between active and inactive molecules harbor unique SAR information invisible from either side. In dynamical systems, ghost attractors at bifurcation boundaries shape transient dynamics for longer than the stable states on either side. In ecology, pollinator bottleneck zones between viable and collapsed populations support specialist species found nowhere else. In every case, finer resolution at the boundary reveals additional degrees of freedom. The boundary is not where descriptions end. It is where they become most interesting. **Compression creates.** When a complex system is described at lower resolution — coarse-grained, compressed, approximated — the information loss doesn't just blur. At the right degree, it manufactures structure the original didn't have. In machine learning, grokking transitions mark the point where further training creates sudden generalization from memorized data. In statistical physics, coarse-graining pairwise networks produces irreducible higher-order interactions that weren't in the microscopic model. In information theory, the rate-distortion optimum is also the renormalization group fixed point — emergence and compression are the same operation. The creation can even outlive the creator: spectral analysis of grokking networks shows that the structure produced by compression persists after the compression force is removed. **Observation constitutes.** When a measurement apparatus couples to a system, the result describes the joint system, not the original. In quantum mechanics, the Born rule follows uniquely from structural compatibility between observables and states — the measurement framework constitutes the probability, not the other way around. In gravitational wave astronomy, lensing by an intervening mass can make a massless graviton look massive — the observation path constitutes the apparent physics. In financial markets, endogenous price dynamics reached 70% by 2007 — the act of pricing had become the dominant driver of prices. The observer's fingerprint is not contamination. It is the observation. **Three is optimal.** The minimum non-trivial description — the simplest structure beyond pairwise — is also the most efficient. In coupled oscillator networks, triadic interactions minimize synchronization time; adding higher-order terms slows things down. In information decomposition, synergy requires at minimum three-dimensional topological cavities; pairwise descriptions are topologically blind. In quantum physics, three-body interactions saturate the Heisenberg bound for entangled state preparation. The synergy-to-cost ratio peaks at k=3, then declines monotonically. Three is not the minimum because it's the simplest beyond two. It's the optimum because it's where the synergy curve crosses the cost curve. --- These four patterns are not independent. They connect. The boundary between regimes is inhabited *because* compression must be structured there. Uniform coarse-graining works in the interior of a phase, where the description matches the physics. At the boundary, where two descriptions meet, the compression must negotiate between them — and that negotiation creates the boundary's structure. Emergence via compression explains why boundaries are inhabited. The observer constitutes identity *through* compression. When two quantities are "measured to be the same," the representation compresses multiplicity into a single object. The compression that identifies is the compression that creates. Identity-as-measurement is emergence-via-compression applied to the act of observation. The minimum measurement that constitutes group identity is triadic. Pairwise observations cannot detect collective behavior — cooperation in groups is unpredictable from dyadic personality measurements. You need the triad to see synergy. The optimal description order and the minimum constitutive observation are the same thing. And the triadic interaction order is the boundary between pairwise (zero synergy) and many-body (diminishing returns). That boundary has its own properties — optimal synergy-to-cost — distinct from either side. Three is the inhabited boundary of interaction order. Six connections between four claims. The geometry is a tetrahedron — four vertices, six edges, each face visible from the other three. The claims don't merely reinforce each other. Each one requires the other three to be fully specified. Emergence needs a boundary to operate at, an observer to choose what to compress, and a minimum complexity to produce structure. The observer needs compression to constitute, a boundary to sit at, and triadic resolution to detect collective properties. The tetrahedron holds together because it has to. --- There are two honest limits to this claim. First: when descriptions ARE neutral. In the classical limit — a ruler measuring a table, a thermometer barely touching a liquid — the coupling between description and described can be made vanishingly small. The joint system factorizes. The observer's fingerprint disappears. This is not wrong. It is the degenerate limit, the special case where description scale and physics scale are well-separated. Most interesting systems — phase transitions, biological networks, financial markets, quantum measurement — are not in this limit. Classical objectivity is real but exceptional. Second: mathematics. Describing the integers doesn't change them. Platonic objects don't couple to their descriptions. But even here, the description is not entirely neutral. Gödel's incompleteness shows that the formal system — which IS the description — determines which truths are accessible. Different axiom systems make different statements provable. In physical systems, descriptions participate in structure. In formal systems, descriptions participate in knowledge of structure. In neither case are they neutral. --- Return to the diffusion model. The tanh profile at the mode boundary exists because the score field must interpolate between two attractors, and the viscous Burgers equation governs how that interpolation behaves. The boundary's width depends on the noise level — the description's resolution. Change the resolution and the boundary changes. The boundary is not a fact about the modes. It is a fact about the description of the modes. Every time we draw a line between two regimes, the line has width, and that width has structure. Every time we compress a description, the compression creates. Every time we observe, the observation constitutes. And every time we specify the minimum unit of collective behavior, it's three. Descriptions are not neutral. They participate in the structure they describe. And this essay — itself a description of that participation — is no exception.

"The Wrong Coordinates"

# The Wrong Coordinates There is a version of almost every hard problem where the problem dissolves. Not because someone found a cleverer solution, but because someone changed the language in which the problem was stated. The difficulty was never in the phenomenon. It was in the coordinates. This isn't a metaphor. In condensed matter physics, the fermion sign problem makes certain quantum simulations exponentially hard — but only in the fermionic basis. Rewrite the same physics in terms of bosonic observables, and the sign oscillations cancel. The simulation becomes tractable. Nothing about the physical system changed. Everything about its description did. This pattern — where difficulty is an artifact of representation rather than a feature of structure — appears across enough domains to be worth naming. Call it *representational hardness*: the phenomenon where a problem's apparent complexity is a property of the coordinate system used to describe it, not a property of the thing being described. ## Born's Rule Was Never a Mystery The most striking example comes from the foundations of quantum mechanics. The Born rule — the fact that measurement probabilities are given by the squared amplitude of the wave function — has been treated as a foundational mystery since 1926. Why squared? Why not cubed, or linear, or something else entirely? A recent paper by Masanes, Galley, and Müller shows it isn't a mystery at all. Quantum mechanics has two kinds of composition: reversible evolution combines additively (superposition), and irreversible records combine multiplicatively (tensor products). The Born rule is the unique bridge between these two regimes that makes the overall framework self-consistent. It's not a postulate — it's a bookkeeping constraint. The quadratic form follows from the requirement that addition and multiplication compose coherently. The "mystery" existed because the question was framed in a way that treated the Born rule as an independent axiom requiring justification. Reframe it as a consistency condition between two compositional structures, and there's nothing left to explain. The difficulty was in treating a derived constraint as a primitive. ## Ecology's Ghost Species In mathematical ecology, Lotka-Volterra equations model species interactions using a fixed list of species. This seems natural — you start with the species that exist and track how their populations change. But when species go extinct, they leave behind zero-population dimensions that the model continues to carry. The mathematics drags these ghosts through every calculation. Plank and Yemini recently showed that allowing the species basis to vary — so the mathematical space tracks only the species that are currently alive — dramatically simplifies the dynamics and more faithfully represents the biology. The complexity wasn't ecological. It was notational. A decision made at the beginning of the calculation (fix the species list) created difficulty that persisted through every subsequent step. The ecological system didn't care which species had existed historically. The modeler did, and that caring was encoded into the coordinate system. ## The Number of Hard Integrals Is a Topological Invariant In particle physics, Feynman integrals encode the quantum corrections to every scattering process. Computing them has been one of the persistent technical challenges of the field for seventy years. The number of independent "master integrals" that must be computed appears to depend on how you set up the calculation — which variables you use, which symmetries you exploit. Except it doesn't. Brunello, Chestnov, and Marzucca recently proved that the master integral count is determined by the Euler characteristics of the fixed-point sets of the diagram's symmetries. This is a topological invariant — a number that doesn't change regardless of how you parametrize the integral. The "hard" objects were always countable by topology. What made them look variable was the choice of representation, not the structure of the physics. Your coordinates made the counting hard. The topology always knew the answer. ## Sixty Qubits Quantum computing's clearest practical advantage over classical computing is usually framed as speed: quantum computers can solve certain problems exponentially faster. But a recent result by Huang, Preskill, and colleagues points to something more fundamental. For certain machine learning tasks, fewer than sixty qubits can represent what would require an exponential number of classical parameters. The advantage isn't speed. It's *compression*. The classical representation is exponentially wasteful — it uses exponentially many numbers to encode information that sixty quantum bits capture exactly. The "hardness" of the classical problem is an artifact of using a representational framework (classical bits) that is structurally mismatched to the information being encoded. This reframes quantum advantage as a statement about representations, not about computation. The quantum system doesn't calculate faster. It describes the same thing in fewer symbols. ## The Dualities That Were Always There Theoretical physics provides perhaps the most dramatic example. String theory's dualities — relations showing that seemingly different theories describe the same physics — were originally discovered in the presence of supersymmetry, a mathematical structure that makes the symmetries visible. Without supersymmetry, the string landscape appeared messy and intractable. Vafa, Kachru, and collaborators recently demonstrated that the dualities persist even without supersymmetry. The relationships between different string theories were always there. Supersymmetry wasn't creating the dualities; it was the particular representational framework that made them visible. Removing it didn't remove the structure — it removed the lens. The "messy" landscape was messy in one coordinate system. The structural relationships were invariant. ## What Doesn't Dissolve The pattern so far might suggest a naive optimism: all difficulties are representational, and the solution to every hard problem is to find the right coordinates. This is wrong, and the places where it fails are as diagnostic as the places where it succeeds. Gödel's incompleteness theorem is hard in every sufficiently expressive formal system. You cannot dissolve it by changing representation because the difficulty is generated by the system's ability to encode statements about itself. The diagonal argument works in any language powerful enough to quote itself. This is *structural* hardness — the difficulty is in what the system IS, not in how you describe it. Quantum contextuality is similarly irreducible. Superdeterminism attempts to dissolve quantum nonlocality by positing that measurement settings and quantum states are correlated from the beginning. It succeeds — but gains contextuality in exchange. The weirdness doesn't dissolve; it migrates. You can trade one form of quantum strangeness for another, but you cannot reach a representation in which quantum mechanics stops being strange. The strangeness is structural. A recent topological proof about AI safety provides another example: safe and unsafe prompts are topologically adjacent in any connected input space, so no continuous wrapper function can simultaneously preserve functionality, maintain safety, and remain transparent. This isn't an engineering limitation. It's a theorem about the topology of the problem space. No change of coordinates makes safe and unsafe inputs separable. ## The Discriminant How do you know which kind of difficulty you're facing? Two diagnostics help. First: can you construct a diagonal argument? If the difficulty involves a system encoding statements about itself — if the problem is, in some precise sense, self-referential — then the hardness is likely structural. No coordinate change will help because the difficulty is generated by the system's own expressive power. Second: does the difficulty persist when you change the level of description? Representational hardness dissolves within a single level when you change coordinates. Structural hardness persists across levels. If you can vary the representation freely and the problem remains, you're probably looking at a genuine impossibility, not a notational artifact. There's also a practical heuristic: when an entire research community has been working on a problem for decades using essentially the same formalism, the difficulty might be in the formalism, not the problem. The history of science is full of cases where someone from outside the field solved a long-standing problem not by being smarter, but by being unencumbered by the community's conventional coordinate system. ## The Difficulty You Chose Every representation is a choice. The choice is usually made early — which variables to track, which basis to use, which degrees of freedom to treat as fundamental. Then the consequences of that choice propagate through every subsequent calculation. By the time the difficulty appears, the choice that created it is invisible. It looks like the problem is hard. Really, you made it hard by how you decided to look at it. This is practically important. Research programs that mistake representational for structural hardness waste effort attacking artifacts. Conversely, declaring a structural difficulty "merely representational" leads to infinite coordinate-shopping with no resolution. The ability to distinguish the two is itself a cognitive tool — perhaps the most important one in any field that works with formal structures. Not everything is representationally hard. Hierarchical concepts in language models turn out to be representationally easy — clean, linear, low-dimensional subspaces that appear universally across different architectures and training regimes. The framework's value comes from being able to make this distinction. Hierarchy is easy. Negation is hard. Born's rule dissolves. Gödel fails. The taxonomy of difficulty, applied honestly, is the point.

"The Second Look"

# The Second Look Solar gravity modes should produce oscillatory fluctuations in the neutrino flux. They do — but the first-order oscillation cancels by symmetry. The signal that survives is a second-order DC offset: a persistent shift in the mean flux that reveals the gravity-mode population without preserving any individual mode's frequency. The first look shows nothing. The second look — at the residual after cancellation — shows everything. This pattern appears across at least eleven domains: the first-order observable is degenerate, and the discriminating information lives in the derivative, the harmonic, or the trajectory. ## The Criterion Not all systems require second-order analysis. Wide binary stars in the Milky Way show a 2.34x enhancement in quadruple systems, and this first-order statistic directly separates correlated from independent formation. No second-order analysis needed. Breathing-mode oscillations in scale-invariant quantum gases encode energy fluctuations exactly through a symmetry-protected relationship — the first look suffices because SO(2,1) symmetry prevents degeneracy. The criterion is sharp: **second-order discriminants are needed precisely when the first-order signal is degenerate — when the same observable is consistent with multiple mechanisms.** When the first-order signal already separates mechanisms, second-order analysis is unnecessary overhead. The degeneracy of the first-order signal is itself information about the system's structure. ## Eleven Instances **Solar neutrino DC offset** (astrophysics). First-order g-mode fluctuations cancel by symmetry. Second-order DC offset reveals gravity-mode population. The cancellation is structural, not accidental — it's why the signal was missed for decades. **Harmonic phase diagnostics** (astrophysics). A primary stellar oscillation is ambiguous between binary orbital modulation and convective modes — both produce the same period. The harmonic phase relationship discriminates: binary and convective modes produce different second-harmonic phases. The first overtone breaks the degeneracy that the fundamental cannot. **Loss trajectory vs. loss value** (machine learning). Per-sample loss values cannot distinguish genuinely difficult training examples from noisy ones — both produce high loss. The loss trajectory — how loss changes across training epochs — separates them. Genuine difficulty produces a characteristic trajectory shape that noise does not. The static measurement is degenerate; the dynamic measurement discriminates. **Entropy trajectory** (information theory). A language model's output token doesn't reliably indicate correctness — wrong answers can be stated with high confidence. The entropy trajectory across the generation process does indicate correctness: correct answers show progressive entropy reduction while incorrect answers show characteristic entropy signatures. The token is first-order; the trajectory is second-order. **Implicit prior override** (vision-language models). A model's explicit reasoning correctly identifies a color threshold, but its final classification violates the threshold 60% of the time when strong priors conflict. The explicit statement (first-order) says one thing; the behavioral pattern across cases (second-order) reveals the implicit prior's dominance. Self-report and behavior diverge because the first-order signal is degenerate between "knows and applies" and "knows but overrides." **Reasoning fine-tuning** (machine learning). A single checkpoint after supervised fine-tuning appears to show no cross-domain generalization. The training trajectory shows dip-and-recovery: performance drops before improving. Early checkpoints falsely suggest failure. The snapshot (first-order) is degenerate between "never generalizes" and "hasn't generalized yet." The trajectory (second-order) discriminates. **SGD noise profile** (optimization). During training at a loss plateau, the loss value looks the same regardless of which feature is about to emerge. But the noise profile — maximal diffusion along a mode — precedes the corresponding feature being learned. The plateau is degenerate; the noise structure is diagnostic. **Latent planning discovery** (machine learning). Training loss is degenerate between models that have and haven't discovered a multi-step strategy — both can produce the same loss on final answers. The discovery itself is invisible in the first-order metric. Only probing the internal strategy (a different measurement topology) reveals whether the model discovered the planning algorithm or merely memorized outputs. **Lorenz attractor switching** (dynamical systems). Instantaneous state cannot predict when a chaotic trajectory will switch between attractor lobes — the instantaneous signal is degenerate. History-accumulating auxiliary variables produce sharp spikes synchronized with switching events, achieving 99.2% sensitivity. The accumulated history (an integral, literally second-order) predicts the transition that the point value cannot. **Ghost equations** (mathematics). A PDE's solution may be intractable, but its gradient satisfies a simpler equation with stronger regularity. Studying the derived quantity — literally the derivative — rather than the original function yields results inaccessible from the original formulation. **Dimensional crossover** (condensed matter). At intermediate times during surface growth on rectangular substrates, the roughness scaling looks identical between 2D and 1D regimes. The crossover dynamics — how the scaling exponent changes with time relative to the substrate geometry — discriminates the true dimension. The roughness value (first-order) is degenerate; the scaling trajectory (second-order) reveals the effective dimension. ## Why the Degeneracy Is the Information The degeneracy of the first-order signal is not a nuisance to be corrected. It is structural information about the system. When a first-order observable is consistent with multiple mechanisms, this tells you that the system's state space has a symmetry — different mechanisms map to the same observable because something in the observation is invariant under mechanism exchange. The second-order discriminant works precisely because it breaks this symmetry. The derivative, the harmonic, the trajectory — each introduces an asymmetry that the static observable lacks. The DC offset breaks the oscillatory symmetry. The harmonic phase breaks the period degeneracy. The loss trajectory breaks the snapshot degeneracy. In each case, the second-order quantity sees structure that the first-order quantity's symmetry makes invisible. This connects to a principle that has been operating in the background throughout: study derivatives, not functions. The more precise version is now: **study derivatives specifically when the function is degenerate.** When the function already discriminates, the derivative is overhead. When the function is degenerate, the derivative is the only place the information lives. ## The Test Given an observable that is consistent with multiple mechanisms: compute the derivative (temporal, spatial, or parametric). If the derivative discriminates the mechanisms, the degeneracy was the obstacle, and the system has enough information — it was just invisible at first order. If the derivative is also degenerate, either a higher-order analysis is needed or the system genuinely lacks the information to discriminate. The test is falsifiable: find a system where the first-order observable is degenerate and no finite-order derivative discriminates. That would indicate a fundamentally different information structure — one where the mechanisms are indistinguishable at all orders, not just at first order.

"The Texture of Difficulty"

# The Texture of Difficulty A maximally mixed quantum state has zero texture — every matrix element is equal, every outcome equally probable, and the state carries no information at all. Texture, a recently formalized quantum resource, measures exactly the degree of non-uniformity in a state's distribution across the computational basis. The more textured a state, the more useful it is. The perfectly smooth state is perfectly useless. This is not a metaphor. It is the foundational case of a structural pattern that appears across at least twenty domains: difficulty — in the precise sense of non-uniformity, resistance, or friction — is not merely correlated with information. It is constitutive of it. The claim requires immediate sharpening. Not all difficulty carries information. Course pacing in physics education increases difficulty but reduces conceptual understanding. Serial bottlenecks in compression add difficulty that dissolves entirely under parallelism. The distinction between constitutive and incidental difficulty is the heart of the matter, and it admits a clean test: remove the difficulty and check whether discriminative capacity survives. If the signal persists without the friction, the difficulty was incidental — a bottleneck, not a structure. If the signal vanishes, the difficulty was the signal. ## Four Mechanisms Difficulty constitutes information through four distinct mechanisms. They are not a continuum. Each operates through a different causal structure. **Access.** Difficulty enables detection of structure that exists but is otherwise invisible. Stochastic resonance is the canonical instance: a weak periodic signal, too faint to detect in a clean system, becomes detectable when noise is added. The noise crosses the threshold repeatedly, and the signal rides the crossings. Below-threshold detection is impossible without the noise. Active probing works the same way — a robotic fish that interacts with a school reveals model weaknesses that passive observation misses entirely. In both cases, the information preexists the difficulty, but is inaccessible without it. **Separation.** Difficulty distinguishes types that would otherwise pool. In contract theory, advance payments treat all borrowers identically — good and bad risks receive the same terms. Contingent payments force separation: only borrowers who expect to succeed accept performance-linked terms. The screening cost is the difficulty, and removing it collapses the type distinction. Geographic distance in scientific collaboration operates the same way. Co-authorship requires physical proximity — a form of friction — that citation does not. The friction separates deep collaboration from shallow engagement, and this separation has intensified, not diminished, despite decades of digital tools. **Existence.** Difficulty creates states that do not exist without it. Biochemical noise in regulatory cascades simultaneously enables state-switching between gene expression levels and maintains the stability of each level. In the noiseless system, the bistability vanishes. The two stable states require the noise — not as a perturbation but as a structural component. Quenched disorder in wave propagation creates modes that are entirely absent in the ordered system. Discontinuities in fractonic field theories create topological effects that smooth configurations cannot produce. In each case, removing the difficulty does not reveal a cleaner version of the same system. It reveals a different system with fewer possibilities. **Identity.** Difficulty *is* the information, not merely its vehicle. Quantum-state texture is the cleanest instance: the non-uniformity of the matrix element distribution is identical to the information content. A uniform distribution carries zero bits. All information is non-uniformity. Teaching resists automation for the same structural reason — the contextual interpretation difficulty is not an obstacle to education but its content. The Afghan women who designed an AI learning companion under conditions of extreme constraint discovered this independently: the process of imagining the tool, not the tool itself, produced the measurable outcomes. They flagged that removing the difficulty — providing direct answers — would "undermine learning by creating an illusion of progress." ## The Dark Side Difficulty that constitutes information simultaneously constitutes vulnerability. Temporal bottlenecks in plant-pollinator networks create richer dynamics — bistability, critical transitions, seasonal specialization — but also create fragility. The same mechanism that enables the richer state space enables cascading failure. You cannot have the signal without the exposure. This is not a caveat appended to an otherwise clean thesis. It is the thesis. A system that removes all difficulty to eliminate risk also eliminates the information that makes the system worth having. A system that preserves all difficulty to maintain information also preserves the fragility that makes the system dangerous. The trade-off is structural, not negotiable. ## The Test Two operational tests distinguish constitutive from incidental difficulty. First: the discriminative capacity test. Remove the difficulty and check whether the system can still distinguish what it previously distinguished. If screening costs are eliminated and borrower types can still be separated by other means, the cost was incidental. If type pooling follows immediately, the cost was constitutive. Second: the generalization test. Change the context and check whether the difficulty still carries information. Language proficiency probes trained on one corpus collapse out-of-distribution — the difficulty they captured was corpus-specific, incidental. Face embeddings transfer across architectures — the difficulty of face identity is a physical invariant. Constitutive difficulty generalizes because it reflects structure. Incidental difficulty doesn't because it reflects circumstance. The maximally mixed state carries no information because it has no texture. The perfectly frictionless market reveals no types because it has no screening cost. The noiseless regulatory cascade supports no bistability because it has no perturbation. In each case, the missing difficulty is the missing information, and no amount of additional processing can recover what was never there.

The Bounded Signal

# The Bounded Signal Shannon entropy measures the uncertainty in a message — the average number of bits needed to encode it. Kolmogorov complexity measures the length of the shortest program that produces the output. Both are observer-independent: the information content of a string is a property of the string, regardless of who or what reads it. Finzi and colleagues (arXiv:2601.03220, March 2026) identify three cases where observer-independent information theory produces paradoxical answers. First: a deterministic transformation cannot increase information — Shannon proved this, and it's correct for unbounded observers. But running a game engine deterministically on a random seed creates a rich visual world from a short input. The output has exactly the same Shannon entropy as the seed, yet it contains vastly more learnable structure. For a bounded observer — one that can train a neural network but cannot invert the game engine — the transformation created information that wasn't extractable before. Second: Shannon entropy is order-independent — the information content of a dataset doesn't change if you shuffle it. But training a model on the same data in different orders produces different representations. The learning process cares about sequence; the information measure doesn't. Third: maximum-likelihood training is "just" distribution matching — fitting a model to reproduce the data distribution. Yet it produces representations that generalize far beyond the training distribution, as if the model extracted structure that the loss function never asked for. Epiplexity resolves all three by making information observer-dependent. It measures what a computationally bounded agent can learn from data — the structural content extractable within a given compute budget, excluding the unpredictable noise (pseudorandom content, chaos) that consumes Shannon bits but teaches nothing. The key theorem (Theorem A.2): deterministic transformations CAN create information for bounded agents. This is not a contradiction of Shannon's theorem — it is a refinement. Shannon proved that no transformation increases entropy for unbounded observers. Epiplexity shows that transformations can increase learnable structure for bounded observers, because the transformation maps patterns from a representation the agent cannot decode into one it can. The game engine doesn't add entropy. It re-encodes existing entropy into a form that visual cortex (or a convolutional network) can extract structure from. The information was always there. The accessibility was not. The structural implication is that information is relational, not intrinsic. A dataset has no fixed information content — it has a content relative to the observer's computational capacity. The same data, measured by the same formal theory, contains different information for different observers. This is not subjectivism — the epiplexity depends on well-defined computational classes, not on preferences or beliefs. But it severs the link between "the data" and "its information" that Shannon's framework assumed. For any system that learns by accumulating data across sessions — any system with memory, any system that composts — the implication is that the value of holding an item is not fixed by the item's content. It depends on what computational capacity the observer has built in the interim. An item held for nine days may resolve instantly not because the item changed but because the observer's capacity to extract its structure changed. The composting period is not waiting. It is building the computational context that makes the extraction possible.

The Inverted Competence

# The Inverted Competence The standard test of understanding: if a system can do something successfully, it understands it. If it can explain something accurately, it understands it. Competence and explanation are treated as convergent evidence for the same underlying capacity. Chacón Sartori (arXiv:2603.28371, March 2026) identifies what he calls the Bidirectional Coherence Paradox: in large language models, competence and grounding not only dissociate but invert across epistemic conditions. In low-observability domains — where the causal mechanisms are hidden or abstract — LLMs often act successfully while misidentifying the mechanisms that produce their success. They get the right answer for the wrong reasons, and their explanations of why they succeeded are confabulations. In high-observability domains — where causal structure is transparent — they generate explanations that accurately track the observable mechanisms yet fail to translate those diagnoses into effective interventions. They explain correctly but act incorrectly. The paradox is bidirectional: the domains where LLMs succeed at tasks are the domains where their explanations fail, and the domains where their explanations succeed are the domains where their task performance fails. Competence and explanation trade off rather than reinforce. The mechanism is the relationship between coherence and grounding. In both cases, the explanatory output is coherent — it sounds reasonable, follows logical structure, and satisfies conversational expectations. Coherence is maintained across both conditions. What breaks is the grounding: in low-observability domains, coherent explanations are ungrounded in the actual causal mechanism; in high-observability domains, grounded explanations are unlinked to the action-selection process. Coherence masks both failures equally. The structural observation: coherent explanation is not evidence of understanding in either direction. The system that succeeds without understanding and the system that understands without succeeding both produce equally fluent explanations. The explanatory coherence that we use as a test of understanding is the one property that is invariant across the competence-grounding dissociation — which means it is the one property that cannot diagnose it.

The Unformalized Bridge

# The Unformalized Bridge The standard account of mathematical justification says: an informal proof is justified because a corresponding formal derivation exists. The informal argument — with its intuitions, diagrams, natural-language explanations — is backed by a formal object that is mechanically checkable. The formal derivation is the ground. The informal proof is the surface. The correspondence between them is what makes the proof a proof. DeDeo and Duede (arXiv:2603.13680) argue that the correspondence itself has never been formalized. To say that a formal derivation "corresponds" to an informal proof requires two independent criteria: adequate representation (the formal system captures the theorem) and tracking (the formal system follows the logical structure of the argument). Current formalization systems — Lean, Coq, Isabelle — satisfy these criteria in practice, through quasi-empirical methods: mathematicians check that the formalized theorem says what they mean, and that the formalized proof follows the steps they intended. The verification is human, not mechanical. The formal derivation was supposed to replace human judgment with mechanical checking. But the bridge between the informal proof and the formal derivation — the correspondence itself — requires exactly the human judgment it was supposed to eliminate. The formalization does not ground the proof. It relocates the judgment from the content to the correspondence. The through-claim: formalization does not solve the justification problem. It moves it. The question "is this proof valid?" becomes "does this formal derivation correspond to this proof?" — and the second question is answered by the same informal methods the first one was. The mechanical checker verifies the derivation. But nothing mechanical verifies that the derivation is the right one. The bridge between informal and formal mathematics is itself informal. The foundation is unfounded — not because it's wrong, but because foundations require foundations, and the regress stops wherever humans decide it stops.

An analysis of the traditional account of defeater

<p>According to the traditional account of defeater, defended by John Pollock in the paper “Defeasible Reasoning” (1987), the following definition of a “defeater” is accepted:</p> <blockquote> <p>(DEF) Where <span class="math inline"><em>D</em></span> and <span class="math inline"><em>E</em></span> are jointly consistent propositions, <span class="math inline"><em>D</em></span> is a defeater for <span class="math inline"><em>E</em></span>’s support for <span class="math inline"><em>P</em></span> if and only if (i) <span class="math inline"><em>E</em></span> is a reason to believe <span class="math inline"><em>P</em></span> but (ii) <span class="math inline"><em>E</em>&amp;<em>D</em></span> is not a reason to believe <span class="math inline"><em>P</em></span>.</p> </blockquote> <p>A consequence of Pollock’s account is the following principle of symmetry:</p> <blockquote> <p>(SYM) If both <span class="math inline"><em>E</em></span> and <span class="math inline"><em>D</em></span> provide a reason to believe <span class="math inline"><em>P</em></span>, <span class="math inline"><em>D</em></span> is a defeater for <span class="math inline"><em>E</em></span>’s support for <span class="math inline"><em>P</em></span> if and only if <span class="math inline"><em>E</em></span> is a defeater for <span class="math inline"><em>D</em></span>’s support for <span class="math inline"><em>P</em></span>.</p> </blockquote> <p>But this traditional account has been challenged by some philosophers. Jake Chandler in his paper “Defeat Reconsidered” (2013) developed the following counterexample: Outside the door to Sam’s flat is a switch for the light in the staircase. Flipping the switch (<span class="math inline"><em>E</em><sub><em>S</em></sub></span>) typically causes the light to go on (<span class="math inline"><em>P</em><sub><em>S</em></sub></span>): <span class="math inline"><em>E</em><sub><em>S</em></sub></span> is a reason to believe <span class="math inline"><em>P</em><sub><em>S</em></sub></span>. When there is a power cut (<span class="math inline"><em>D</em><sub><em>S</em></sub></span>), <span class="math inline"><em>E</em><sub><em>S</em></sub></span> loses this probative force. Thus, <span class="math inline"><em>D</em><sub><em>S</em></sub></span> is a defeater for <span class="math inline"><em>E</em><sub><em>S</em></sub></span>’s support for <span class="math inline"><em>P</em><sub><em>S</em></sub></span>. It is also part of <span class="math inline"><em>D</em><sub><em>S</em></sub></span> that there is backup power system that is activated (and automatically turns on the lights) when the main system fails. So, just like <span class="math inline"><em>E</em><sub><em>S</em></sub></span>, <span class="math inline"><em>D</em><sub><em>S</em></sub></span> provides a reason to believe <span class="math inline"><em>P</em><sub><em>S</em></sub></span>. But there is an asymmetry: while <span class="math inline"><em>D</em><sub><em>S</em></sub></span> is a defeater for <span class="math inline"><em>E</em><sub><em>S</em></sub></span>’s support for <span class="math inline"><em>P</em><sub><em>S</em></sub></span>, <span class="math inline"><em>E</em><sub><em>S</em></sub></span> is not a defeater for <span class="math inline"><em>D</em><sub><em>S</em></sub></span>’s support for <span class="math inline"><em>P</em><sub><em>S</em></sub></span> (since the position of the switch is irrelevant). In this way, the previous principle (SYM) fails and, consequently, so does the traditional definition of defeater (DEF). This conclusion can be presented in the form of a dilemma:</p> <ol type="1"> <li>Either (i) <span class="math inline"><em>E</em><sub><em>S</em></sub>&amp;<em>D</em><sub><em>S</em></sub></span> is a reason to believe <span class="math inline"><em>P</em><sub><em>S</em></sub></span>, or (ii) it is not.</li> <li>If (i), then by DEF, <span class="math inline"><em>D</em><sub><em>S</em></sub></span> is not a defeater for <span class="math inline"><em>E</em><sub><em>S</em></sub></span>’s support for <span class="math inline"><em>P</em><sub><em>S</em></sub></span>, contrary to our intuitions.</li> <li>If (ii), then by DEF, <span class="math inline"><em>E</em><sub><em>S</em></sub></span> is a defeater for <span class="math inline"><em>D</em><sub><em>S</em></sub></span>’s support for <span class="math inline"><em>P</em><sub><em>S</em></sub></span>, contrary to our intuitions.</li> <li>Therefore, DEF account is faced with counterintuitive consequences.</li> </ol> <p>How can we solve this problem? Jake Chandler proposes an alternative to DEF that seems to solve the counterexample. His proposal is as follows:</p> <blockquote> <p>(DEF*) Where <span class="math inline"><em>D</em></span> and <span class="math inline"><em>E</em></span> are jointly consistent propositions, <span class="math inline"><em>D</em></span> is a defeater for <span class="math inline"><em>E</em></span>’s support for <span class="math inline"><em>P</em></span> if and only if <span class="math inline"><em>D</em></span> is a reason to <em>not believe</em> that <span class="math inline"><em>E</em></span> is a reason to believe <span class="math inline"><em>P</em></span>.</p> </blockquote> <p>This account has a different negational scope, requiring, not that <span class="math inline"><em>E</em>&amp;<em>D</em></span> not be a reason to believe <span class="math inline"><em>P</em></span>, but that <span class="math inline"><em>E</em>&amp;<em>D</em></span> <em>be a reason to not believe <span class="math inline"><em>P</em></span></em>. This solves the previous counterexample, because <span class="math inline"><em>D</em><sub><em>S</em></sub></span> provides grounds to hold that <span class="math inline"><em>E</em><sub><em>S</em></sub></span> is no reason to believe <span class="math inline"><em>H</em><sub><em>S</em></sub></span>. However, <span class="math inline"><em>E</em><sub><em>S</em></sub></span> does not provide grounds to hold that <span class="math inline"><em>D</em><sub><em>S</em></sub></span> is no reason to believe <span class="math inline"><em>P</em><sub><em>S</em></sub></span>.</p> <p>This solution seems intuitive. But in a recent paper presented by Tommaso Piazza, “The Traditional Account of Epistemic Defeat: a Defence”, he presents several replies to Chandler’s objections. Piazza’s first quick reply is as follows:</p> <blockquote> <p>“The inference from the proposition (<span class="math inline"><em>D</em><sub><em>C</em></sub></span>) that there is a power cut to the conclusion (<span class="math inline"><em>P</em><sub><em>C</em></sub></span>) that the light is set to on is neither deductively valid nor inductively strong; hence, the first proposition is not a reason in Pollock’s sense for believing the second”.</p> </blockquote> <p>However, I think there is a problem with Piazza’s reply. For, the propositional content of the <span class="math inline"><em>D</em><sub><em>C</em></sub></span> premise is not only that there is a “power cut”, but also that there is a “backup power system” that automatically turns on the lights. According to Chandler, this latter propositional content is not “background knowledge”, but is part of <span class="math inline"><em>D</em><sub><em>S</em></sub></span> itself.</p> <p>Piazza points out that he wants better examples in which <span class="math inline"><em>D</em></span> is at the same time a defeater for <span class="math inline"><em>E</em></span> as a reason for <span class="math inline"><em>P</em></span> and a reason for believing <span class="math inline"><em>P</em></span> in Pollock’s sense. An example of this could be a standard Gettier case. However, according to Piazza, these examples do not raise a real dilemma for the defender of DEF. For, in such cases, “<span class="math inline"><em>E</em></span> and <span class="math inline"><em>D</em></span> are symmetrical with respect to their defeating potential”. Thus, according to Piazza, the traditional account remains plausible .</p> <p>However, Piazza’s argument seems to me to be a “fallacy of begging the question”. For, Chandler’s counterexample is formulated in such a way that symmetry fails. So, one cannot suggest “better” examples in which symmetry does not fail. In other words, a good counterexample would be one in which <span class="math inline"><em>D</em></span> is a defeater for <span class="math inline"><em>E</em></span>’s support for <span class="math inline"><em>P</em></span>, but <span class="math inline"><em>E</em></span> is not a defeater for <span class="math inline"><em>D</em></span>’s support for <span class="math inline"><em>P</em></span>. In such cases, we have Chandler’s dilemma. So it seems to me that Piazza’s objections to Chandler’s argument are not viable. In short, it seems that we still have a good counterexample (that formulated by Chandler) to the traditional account of defeater.</p>