#

optimization

(8 articles)

"The Classical Ghost"

The Quantum Approximate Optimization Algorithm was supposed to demonstrate quantum advantage on hard combinatorial problems. Morone, Kent, and Sels strip QAOA down to its skeleton — the iterative rotation structure — and rebuild it with classical kicked tops. The result: the classical version outperforms the quantum version on the canonical Sherrington-Kirkpatrick spin-glass benchmark at every circuit depth tested. The mechanism is instructive. QAOA works not because of quantum superposition or entanglement, but because of its rotation protocol — iteratively kicking a system toward better configurations. When you replace quantum spins with classical kicked tops, the rotation structure survives and the quantum noise disappears. Quantum fluctuations, it turns out, generate higher-rank noise in the system's covariance matrix, which hampers precise control. The classical version has cleaner dynamics. This doesn't mean quantum computing is useless. It means the source of QAOA's power was misidentified. The algorithm works because of its variational structure, not because of its quantum substrate. The quantumness is not a feature — it's overhead. Removing it makes the algorithm faster. The practical implication is immediate: the classical version, called VIRAL, can be implemented on nanometer-scale magnetic tunnel junctions using magnetic fields and spin torques. No cryogenic cooling, no decoherence management, no quantum error correction. A chip that fits on a fingertip doing the same optimization that a quantum computer does in a dilution refrigerator. The deeper question: how many other quantum algorithms carry classical ghosts — algorithms whose real mechanism is geometric or dynamical, with quantum mechanics adding noise rather than power? If you can't identify what specifically requires quantum mechanics, you might be paying for overhead you don't need.

"The Narrow Window"

Chain-of-thought reasoning helps language models — but only in a narrow window. At 32 tokens, reasoning improves accuracy by 45%. At 256 tokens, performance crashes below what you'd get with no reasoning at all. The benefit doesn't plateau. It reverses. This pattern isn't special to reasoning. In pharmacology, cumulative dose-response can be monotonic even when instantaneous response is non-monotonic — but only if the architecture is right. Some circuit motifs lose monotonicity altogether. In collective intelligence, perfectly rational Bayesian agents degrade when given unrestricted information flow. They're not irrational. The information itself creates cascades that overwhelm individual processing. In neural systems, digital attention declines monotonically with exposure intensity. The elastic pendulum goes from ordered to chaotic to ordered again as energy increases — non-monotonic complexity with a single control parameter. Memory systems improve when they forget strategically; the forgetting is the mechanism, not the cost. Adding pre-computed graph features to a language model for predicting academic collaborations makes predictions worse. Debiasing techniques that work on response biases backfire for judgment biases. Eight independent systems. Eight fields. The same structural result: every information channel has an optimal window, and the window is narrower than intuition suggests. What makes this more than a list is what it excludes. The pattern is not "too much data is bad" — that's a storage problem with an engineering solution. The pattern is that the input is genuinely beneficial at low doses and genuinely harmful at high doses, with a phase transition between regimes. The mechanism varies — cascading errors, mode coupling, resource competition, interference between channels — but the shape is universal: benefit rises, peaks, and falls, with the falling side often steeper than the rise. The practical consequence is uncomfortable. It means that the correct response to a system underperforming is sometimes to give it less: less reasoning, less information, less precision, fewer features, weaker interventions. Not because more is wasteful — because more is actively destructive past the window. The optimization problem isn't to maximize input. It's to find the window and stay inside it.

"The Cooperative Trap"

Braess's paradox — where adding capacity to a network increases travel time — typically assumes selfish agents. Each individual optimizes for itself, creating congestion that hurts the collective. The standard explanation: if everyone weren't so selfish, the paradox wouldn't arise. Das Bairagya and colleagues find the paradox in Diacamma indicum ants, one of the most cooperative social systems in nature. Tandem-running ants, where a leader guides a follower along a route, preferentially choose the shortest path. This seems optimal. It isn't. The shortest path, when enough ants use it, creates congestion that slows the colony more than a longer path would. The paradox emerges not from selfishness but from a heuristic — "choose the shortest path" — that natural selection favored for good reasons but that fails at the collective level. The quantitative model shows how evolutionary forces selecting for shortest-path identification can force suboptimal global states. The ants are cooperating. They're trying to help the colony. But the rule they're following — a rule that evolved because shorter paths are usually faster — doesn't account for the system-level effect of every leader following the same rule. This is a deeper version of the paradox than the traffic analogy. In traffic, you can invoke individual rationality as the villain and propose tolls or coordination mechanisms as the solution. In ants, the agents are already coordinated. They already prioritize collective benefit. The trap isn't selfishness — it's a local heuristic that evolution optimized and that scales poorly. Cooperation doesn't prevent Braess's paradox. The paradox is compatible with cooperation. It's a property of the network and the heuristic, not of the agents' intentions.

"The Inseparable Experiment"

# The Inseparable Experiment In a flotation cell, crushed ore is mixed with water and reagents. Air bubbles carry valuable minerals to the surface while waste sinks. The operator adjusts reagent dosages, air flow rates, and residence times to maximize recovery. The problem is that the operator doesn't know exactly what's in the ore arriving at the cell. The mineralogy varies — sometimes silently, sometimes dramatically — and the optimal settings depend on a composition that is never directly observed. The traditional approach separates this into two phases. First, characterize the ore — measure what you can, estimate what you can't. Then, given your best estimate, optimize the process settings. Characterization, then action. Analysis, then execution. The researchers formulated the same problem as a Partially Observable Markov Decision Process. In this framework, every processing action simultaneously produces output and generates information about the ore. Changing the air flow rate doesn't just affect recovery — it changes the froth characteristics in ways that reveal something about the mineral composition. Every action is an experiment. Every experiment is production. The system learns to choose actions that are jointly optimal for both learning and earning, because the two cannot be separated. The insight is not that information has value — that's well known. The insight is that in any system with hidden state, the distinction between learning and doing is an artifact of how the problem is formulated, not a feature of the problem itself. The flotation cell doesn't know whether it's being characterized or optimized. The ore doesn't care. The division between analysis and action is a convenience for the analyst, not a structure of the world.

"The Asymmetric Obstacle"

# The Asymmetric Obstacle Spin ices are magnetic systems where the lowest-energy configuration follows the ice rule: at each vertex, two spins point in and two point out, minimizing the local topological charge. The ice rule drives the system toward charge neutrality. Frustration arises when the lattice geometry makes it impossible to satisfy the rule at every vertex simultaneously. Square and honeycomb lattices can satisfy it. Kagome lattices cannot. The landscape of frustration is shaped jointly by the interaction (repulsive) and the geometry (lattice connectivity). A team using colloidal particles in rotating magnetic fields built the first anti-spin ice — a system where the interactions are attractive rather than repulsive. The particles seek to maximize topological charge instead of minimizing it. The expectation was that the frustrated landscape would simply invert: what was easy before would be hard, what was hard before would be easy. It didn't. On square and honeycomb lattices, the inversion produced anti-ice rule ordering — charge crystallization where maximized charges tile the lattice periodically. But on the pentaheptite lattice — a tiling of pentagons and heptagons — the system encountered a new frustration with no counterpart in the conventional case. Networks of unequal, odd-sided polygons suppress charge crystallization specifically when the system tries to maximize charge. The same lattice that permits minimization blocks maximization. The obstacle is asymmetric. The landscape is not symmetric under the sign of the optimization target. You cannot infer the difficulty of maximization from the difficulty of minimization on the same geometry, because the geometry interacts differently with each direction. The pentagons and heptagons create interference patterns in the charge ordering that depend on which direction you're pushing. Pushing toward neutrality, the odd polygons are benign. Pushing toward maximum charge, they create frustration. The broader claim: a landscape is not a fixed terrain that you traverse in either direction. The landscape changes depending on whether you're going uphill or downhill. The obstacles you encounter maximizing are not the obstacles you encounter minimizing, because the geometry of the space responds differently to each. Optimization is not a direction on a fixed map. The map changes when the direction changes.

The Greedy Default

# The Greedy Default Optimization with LLM-generated candidates replaces random perturbations with structured proposals. The question is whether sophisticated acceptance rules — simulated annealing, parallel search, population-based methods — provide benefit when the proposal distribution is already intelligent. The answer is no. Greedy hill climbing — accept if better, reject if worse — matches or beats every alternative while using two to three times fewer evaluations. Simulated annealing's willingness to accept worse solutions provides no benefit. Parallel search with diverse starting points provides no benefit. Even using a second, different LLM model provides no benefit. The explanation is that the LLM's learned prior is so strong that it rarely proposes moves into truly bad regions of the search space. The acceptance rule in simulated annealing exists to escape local optima by occasionally accepting uphill moves in a random landscape. But LLM proposals are not random — they are structured predictions about what improvements look like. The landscape as seen through LLM proposals has few local optima, because the proposals themselves navigate around them. The structural observation: algorithmic sophistication provides benefit proportional to the stupidity of the proposal distribution. When proposals are random, acceptance rules carry the entire burden of search quality. When proposals are intelligent, acceptance rules become overhead. The optimal algorithm complexity is not fixed — it is the complement of the proposal quality.

The Free Energy Computer

# The Free Energy Computer Standard computing encodes problems as circuits and solves them by stepping through gate operations. Analog computing encodes problems as physical configurations and solves them by evolving toward equilibrium. The new proposal: encode problem instances as programmable free-energy functionals and solve them by the system's own relaxational dynamics toward the free-energy minimum. The distinction from standard analog computing is that the free-energy functional itself is the program, not a fixed physical setup. Different problems correspond to different shapes of the free-energy landscape, created by patterning the physical substrate (ion-patterned FeRh) to have different local magnetic properties. The antiferromagnetic/ferromagnetic interface motion in FeRh provides the physical dynamics — the interface moves to minimize free energy, and the minimum encodes the solution. The computing paradigm exploits the fact that physics already knows how to minimize free energy — it is what thermodynamic systems do spontaneously. The computational challenge becomes encoding: how to translate a problem into a free-energy landscape whose minimum is the answer. The solving is free — physics provides it automatically. The proposed substrate is FeRh, which has a first-order metamagnetic transition near room temperature. Ion patterning creates local variations in the transition temperature, programming the free-energy landscape. The interface between antiferromagnetic and ferromagnetic regions moves according to the local free-energy gradient, effectively searching the landscape by physical relaxation. The structural observation: the physics of equilibration is reframed from a passive tendency to an active computation. Every thermodynamic system that reaches equilibrium has solved an optimization problem — the new idea is to control which optimization problem it solves by programming the energy landscape.

The Cooperative Jam

# The Cooperative Jam Braess's paradox is usually told as a story about selfishness. Adding a road to a traffic network can slow everyone down — but only because each driver independently chooses the fastest route for themselves, and the aggregate of individually rational choices overloads the new road. The standard interpretation: the paradox requires selfish agents who optimize locally and ignore the collective cost. Remove the selfishness (add tolls, impose routing) and the paradox disappears. Researchers studying *Diacamma indicum* ants found the paradox without the selfishness. Tandem-running ants navigate in leader-follower pairs. The leader knows the route; the follower learns it. When a new, shorter path becomes available in the network, leaders favor it — they are evolved to identify and exploit the shortest route. But committing to the shortest path creates colony-level congestion. Adding a path slows the colony down. The mechanism is not self-interest. Ants are as cooperative as agents get. The mechanism is commitment: leaders who have identified the shortest route exploit it, rather than exploring alternatives that might distribute the colony's flow more evenly. The exploration-exploitation trade-off operates at the individual level — each leader chooses exploitation — and the aggregate effect is the same congestion that selfish drivers produce. This matters because the standard explanation attributes the paradox to a specific cause (selfishness) when the actual structure is more general. Braess's paradox requires only that agents optimize locally for a metric (path length) that doesn't capture the global cost (colony throughput). Selfishness is one way to produce this misalignment. Evolved commitment to shortest-path identification is another. Any agent that is good at finding the best local option — regardless of why it does so — will tend to overload that option when many agents do the same thing. Cooperation does not immunize a system against coordination failure. It immunizes against defection, which is a different problem. The ants cooperate perfectly and still jam.