0. The claim
Three fields that rarely cite each other describe their central mysteries in the same grammar. Neuroscientists say a memory is an attractor. Economists say a crash is a phase transition. Hardware engineers say an Ising machine "finds" a solution by relaxing into a minimum. In each case the interesting behavior is not written down anywhere in the components. It appears when many coupled elements settle, oscillate, or teeter together.
The shared grammar is dynamical systems theory: state spaces, flows, fixed points, limit cycles, bifurcations, slow and fast variables, noise. This essay takes each domain seriously on its own terms, then asks which parts of the shared vocabulary are real transfers and which are metaphors wearing equations.
1. Brains: computation as flow on a manifold
1.1 Attractors as memory
John Hopfield's 1982 paper made the analogy exact. A network of binary units with symmetric weights has an energy function that never increases under asynchronous updates. Stored patterns are minima. Recall is descent: start near a pattern, fall into it. Memory is not an address, it is a basin. Amit, Gutfreund, and Sompolinsky (1985) used spin-glass methods to compute the capacity, roughly 0.14 patterns per neuron before spurious minima overwhelm retrieval. Modern dense associative memories (Krotov and Hopfield, 2016) and the "Hopfield networks is all you need" line (Ramsauer et al., 2020) raise capacity to exponential in dimension and show that transformer attention is a single step of such a retrieval dynamics. That connection matters for Section 5.
Point attractors handle discrete memories. Continuous attractors handle analog quantities. The head-direction system in rodents behaves like a ring attractor: a bump of activity on a ring of neurons, whose position encodes heading, moved by velocity inputs. Zhang (1996) gave the theory; Seelig and Jayaraman (2015) found a literal ring of activity in the fly ellipsoid body; Kim et al. (2017) showed it behaves as the theory predicts. Grid cells are the two-dimensional analog: Burak and Fiete (2009) built a continuous attractor on a torus, and Gardner et al. (2022) recorded thousands of grid cells and found population activity lying on a toroidal manifold, one of the cleaner confirmations of a dynamical hypothesis in systems neuroscience.
Working memory was long modeled as persistent activity held by recurrent excitation, following Amit and Brunel (1997) and Compte, Wang, et al. (2000). The alternative view, that information can be held in "activity-silent" synaptic traces such as short-term facilitation (Mongillo, Barak, Tsodyks, 2008; Stokes, 2015; Wolff et al., 2017), is a genuine controversy. Both are attractor stories, but they differ in which variables are slow. In the silent view, the relevant state lives in synaptic efficacies, and firing is a readout perturbation. This is a nice example of how "where is the memory" is a question about which dimensions of the state space carry the slow variable.
1.2 Metastability and itinerancy
Real cortex does not sit in a fixed point. Rabinovich, Huerta, Varona, and Afraimovich (2008, "Transient cognitive dynamics, metastability, and decision making") proposed stable heteroclinic channels: sequences of saddle points, each briefly visited, connected by their unstable manifolds. Cognition is itinerancy along the channel, not convergence. Tsuda's "chaotic itinerancy" is an older cousin. Tognoli and Kelso (2014, "The metastable brain") frame the same idea via coordination dynamics: in the extended HKB model, when coupling weakens, the phase-locked fixed points disappear through a saddle-node bifurcation, leaving "ghosts" where the system lingers before escaping. Metastability here has a precise meaning: dwelling near remnants of attractors without settling.
1.3 Criticality
Beggs and Plenz (2003) recorded cortical slices and cultures and found "neuronal avalanches" with power-law size distributions and exponent near minus 1.5, matching a critical branching process. The hypothesis that the brain tunes itself to a critical point is attractive because critical systems maximize dynamic range (Kinouchi and Copelli, 2006), information transmission, and susceptibility.
The honest status: contested. Power laws alone are weak evidence; Touboul and Destexhe (2010, 2017) showed non-critical models produce them. Better tests use exponent relations and shape collapse (Friedman et al., 2012; Fontenele et al., 2019). Priesemann and colleagues (2014, and Wilting and Priesemann, 2018) argue for a "slightly subcritical" reverberating regime that keeps a safety margin from runaway activity. Moretti and Muñoz (2013) proposed that heterogeneity in connectivity smears the transition into a Griffiths phase, an extended region with critical-like properties and no fine tuning. My reading: the brain is near, not at, a transition, and "critical-like" behavior may be generic for heterogeneous networks rather than a tuned achievement. That is still dynamically meaningful, but it is a weaker claim than the popular version.
1.4 Synchronization and coherence
Kuramoto (1975) gave the minimal model: N phase oscillators with all-to-all sinusoidal coupling. Below a critical coupling they drift; above it, an order parameter (the mean phase vector) grows continuously from zero. This is a second-order phase transition in a deterministic system, and it is the template for every "sync" story since. In cortex, Fries's communication-through-coherence hypothesis (2005, revised 2015) proposes that effective connectivity between regions is gated by phase alignment of gamma rhythms, so that inputs arrive during the receiver's excitable phase. Evidence is supportive but not decisive; Ray and Maunsell have argued gamma is too variable to serve as a reliable clock. The dynamical content is clear regardless: phase relations are a control variable that can be changed faster than synaptic weights.
1.5 Low-dimensional population dynamics
The most consequential shift in the last fifteen years is treating a population's activity as a trajectory on a low-dimensional manifold. Churchland, Cunningham, Kaufman, Shenoy, and colleagues (2012) applied jPCA to motor cortex during reaching and found rotational dynamics, as if the population were a damped oscillator whose initial condition (set during preparation) determines the reach. Sussillo and Barak (2013) showed how to find fixed points and slow points of trained recurrent networks and use them to explain computation. Mante, Sussillo, Shenoy, and Newsome (2013) showed a trained RNN reproduced context-dependent selection in prefrontal cortex via a line attractor plus a selection vector. LFADS (Pandarinath et al., 2018) fit a latent dynamical system to spike trains directly. Gallego, Perich, Miller, and Solla (2017) argued that "neural manifolds" are the right unit of analysis, and Gallego et al. (2020) showed manifolds are stable across months even as individual neurons drift.
Caveats: rotational structure can arise from trivial sources (Lebedev et al. raised this; the reply is that shape and timing constraints matter). And a manifold found by PCA is a statistical object; whether it is a dynamical invariant is a separate claim.
1.6 Edge of chaos and reservoirs
Jaeger (2001) and Maass, Natschläger, and Markram (2002) independently proposed reservoir computing: a fixed random recurrent network with a trained linear readout. It works because a high-dimensional nonlinear system with fading memory separates input histories. Bertschinger and Natschläger (2004) showed computational performance peaks near the order-chaos transition, where the network has long memory without losing sensitivity to input; Legenstein and Maass (2007) refined this. Langton's older "edge of chaos" for cellular automata is the ancestor. This is one of the few places where "criticality helps computation" has a clean quantitative demonstration, and it transfers directly to hardware in Section 3.
1.7 Free energy as gradient flow
Friston's free energy principle (2010 review; many papers since) says biological systems minimize variational free energy, a bound on surprise, and that perception and action are gradient descent on it. As a dynamical statement, it says neural states flow down a Lyapunov function defined by a generative model. Critics (Colombo and Wright, 2018; Biehl, Pollock, and Kanai, 2021, on the Markov blanket derivations; Bruineberg et al., 2022) argue that in its general form the principle is close to unfalsifiable, and that the substantive content lives in specific process theories like predictive coding. I think that is right: the useful part is the local claim that cortical circuits implement approximate Bayesian inference by relaxation, which is testable, not the universal claim.
2. Systems of people: order parameters without a controller
2.1 Synergetics and the slaving principle
Hermann Haken, working from laser physics in the 1970s, formalized what later fields rediscovered. Near an instability, a few slow modes (order parameters) become unstable while many fast modes remain stable. Adiabatic elimination lets you express the fast modes as functions of the slow ones: the fast are "slaved." The result is a closed, low-dimensional equation for the order parameters. This is the mathematical core of emergence in the weak sense: the macroscopic description is not added on, it is derived by separation of time scales. Haken applied it to pattern formation, then to brain (with Kelso, the HKB model of 1985), then to sociology. It is the most transferable single idea in this essay.
2.2 Schelling and agent-based models
Schelling (1971) showed that agents with mild in-group preference produce sharp segregation on a grid. The macroscopic pattern is not in any agent's preferences. Later work (Vinković and Kirman, 2006) mapped Schelling dynamics onto phase separation in physics; Stauffer and Solomon and others linked it to Ising-type models. Epstein and Axtell's Sugarscape (1996) generalized the method. Agent-based modeling is essentially numerical dynamical systems where the state space is too high-dimensional and heterogeneous for closed-form analysis.
2.3 Complexity economics
Brian Arthur (1989, "Competing Technologies, Increasing Returns, and Lock-In") used stochastic path-dependence to show that with increasing returns, markets can lock into inferior technologies. Formally, this is a nonlinear Pólya urn with multiple stable fixed points; early noise picks the basin. The Santa Fe Institute program (Arthur, Durlauf, Lane, 1997) framed the economy as an evolving, out-of-equilibrium system. The Santa Fe artificial stock market (Arthur, Holland, LeBaron, Palmer, Tayler, 1997) showed that heterogeneous adaptive agents generate fat tails and volatility clustering absent from rational-expectations equilibria. Farmer and Foley's 2009 Nature comment argued ABMs should inform policy; adoption has been slow.
2.4 Instability and crashes
Minsky's financial instability hypothesis (1970s; "Stabilizing an Unstable Economy," 1986) is a verbal dynamical model: stability breeds leverage, leverage breeds fragility. Sornette and collaborators (Johansen, Ledoit, Sornette, 2000; Sornette's 2003 book "Why Stock Markets Crash") formalized bubbles as regimes of faster-than-exponential growth with log-periodic oscillations, decorated by a discrete scale invariance, ending at a finite-time singularity. The model has had some out-of-sample successes and notable failures; the community is split, and I would treat the log-periodic signature as suggestive rather than established. What is uncontroversial is that herding and positive feedback produce super-exponential growth that must end.
2.5 Reflexivity
Soros's reflexivity (1987, "The Alchemy of Finance") says participants' beliefs affect fundamentals, which affect beliefs. In dynamical terms, the observation function feeds back into the state equation, so there is no fixed point defined independently of expectations. Equilibrium economics assumes this loop is absent or contracting. Reflexivity says it can be expanding. Beinhocker and others have argued this is what makes economies non-ergodic. This is a real structural difference from physics: the components model the system.
2.6 Opinion dynamics
Deffuant, Neau, Amblard, and Weisbuch (2000) and Hegselmann and Krause (2002) introduced bounded-confidence models: agents average opinions only with others within a threshold. Above a critical threshold, consensus; below, fragmentation into clusters. This is a bifurcation controlled by a single parameter and produces polarization without any polarizing agent. Kuramoto-style models of opinion synchronization exist (Pluchino, Latora, Rapisarda, 2005) but the phase-oscillator picture is a looser fit for opinions than for neurons. Castellano, Fortunato, and Loreto's 2009 review is the standard survey.
2.7 Early warning signals
Scheffer et al. (2009, Nature, "Early-warning signals for critical transitions") observed that near a fold bifurcation, the dominant eigenvalue approaches zero, so recovery from perturbations slows. Critical slowing down shows up as rising autocorrelation and variance. Dakos et al. and Lenton et al. applied this to lakes, climate, and ecosystems. Applications to finance (Diks, Hommes, Wang, 2019) and social systems are more tentative. The problem is that many social transitions are not fold bifurcations, and the signals also arise from increased noise or changing measurement. Boettiger and Hastings (2012) documented false positives. Actionability is discussed in Section 5.
2.8 Mean fields and weak emergence
When agents are many and interactions are dilute, one can pass to a mean-field limit and obtain differential equations for population fractions. This is why epidemiological SIR models and Kuramoto both work. The condition is that correlations between individuals are negligible, which is exactly what fails in bubbles and mobs.
Bedau (1997, 2002) distinguished weak emergence, where macro behavior is derivable only by simulation but is fully determined by the micro rules, from strong emergence, where macro behavior has irreducible causal powers. Social emergence is weak emergence. This has a sharp implication for prediction: weakly emergent systems can be computationally irreducible (Wolfram's term), so the fact that we understand the rules does not mean we can forecast. Prediction is limited by sensitivity and by reflexivity, not by ignorance of mechanism.
3. In situ memory: when the physics does the computing
3.1 The von Neumann bottleneck
Backus's 1978 Turing lecture named it: separating memory from processor means most energy and time go to moving data. The brain does not do this. Synapses are simultaneously the memory and the multiply. Whether this is a principled advantage or an accident of biology is worth asking. The claim holds for energy: shuttling a word across a chip costs orders of magnitude more than a floating-point operation, and DRAM access dwarfs both. It fails as a universal principle in one sense: separation buys programmability and error correction, which is why digital won. In-memory computing is an attempt to get co-location without giving up too much of either.
3.2 Memristive crossbars
Chua (1971) predicted the memristor; Strukov, Snider, Stewart, and Williams (2008) reported one in titanium dioxide. A crossbar of memristive elements performs a matrix-vector product in one step by Ohm's and Kirchhoff's laws: apply voltages to rows, read currents on columns, conductances are the weights. Prezioso et al. (2015) trained a small perceptron on such an array. Ambrogio et al. (2018, IBM) demonstrated mixed-precision training with phase-change memory at software-equivalent accuracy. The physics is doing multiply-accumulate; the weights sit where the computation happens. Limitations are real: device variability, drift, limited write endurance, and the analog-to-digital conversion at the periphery often dominates energy.
3.3 Neuromorphic hardware
Carver Mead's analog VLSI program (Mead, 1989, "Analog VLSI and Neural Systems"; Mahowald and Mead's silicon retina) used transistors in subthreshold, where currents are exponential in voltage, to emulate ion-channel dynamics directly. The computation is the device physics. Digital neuromorphic systems went a different way: SpiNNaker (Furber et al., Manchester, 2014) uses many small ARM cores passing spikes as packets; Intel's Loihi (Davies et al., 2018) and Loihi 2 (2021) implement spiking neurons and on-chip plasticity in digital logic, gaining event-driven sparsity while keeping determinism. BrainScaleS (Heidelberg) is mixed-signal and runs faster than biological time. IBM's TrueNorth (2014) and NorthPole (2023) push memory-adjacent digital compute. The lesson: co-location is a spectrum, and the gain comes from locality and sparsity more than from analog per se.
3.4 Physical reservoir computing
If a reservoir just needs to be a high-dimensional nonlinear system with fading memory, almost anything qualifies. Fernando and Sojakka (2003) used a bucket of water. Appeltant et al. (2011) used a single delayed-feedback optoelectronic node. Torrejon et al. (2017, Nature) used a spintronic nano-oscillator for spoken digit recognition. Photonic reservoirs (Vandoorne et al., 2014; Brunner, Fischer, and colleagues) exploit passive silicon photonics. Nakajima et al. (2015) used a soft silicone octopus arm. Tanaka et al. (2019) reviewed the field. The important dynamical point: these systems work best when tuned near their own edge of chaos, exactly as Bertschinger and Natschläger found in software. The substrate's transient dynamics is the computation, and the readout is the only learned component.
3.5 Energy-based computation by relaxation
Hopfield's insight runs backward too: if a physical system has an energy function, letting it relax solves the minimization. Ising machines encode a combinatorial problem in spin couplings and let the hardware find low-energy states. D-Wave's quantum annealers are the best-known; whether they provide quantum advantage remains disputed. Coherent Ising machines (Inagaki et al., 2016; McMahon et al., 2016, Science) use networks of optical parametric oscillators whose phases settle to a spin configuration. Oscillator-based Ising machines (Wang and Roychowdhury, 2019) use subharmonic injection locking so that coupled electronic oscillators' phases binarize and minimize an Ising Hamiltonian. Memristive Hopfield networks (Cai et al., 2020, Nature Electronics) do the same with crossbars plus noise for annealing. In every case, memory (the coupling matrix) and computation (the relaxation) are the same physical object.
3.6 Learning in the physics
Backpropagation needs a separate computational graph. Can the physics learn as well as compute? Scellier and Bengio (2017) proposed equilibrium propagation: let an energy-based network relax to a free equilibrium, then nudge the outputs toward the target and relax again; the difference in local quantities between the two equilibria approximates the gradient. Stern, Hexner, Rocks, and Liu (2021) and Stern and Murugan (2023 review, "Learning without neurons in physical systems") built "coupled learning" into physical networks, including electronic and mechanical ones, where local rules on edges update conductances or stiffnesses using only locally measurable quantities from free and clamped states. Dillavou et al. (2022) built a self-learning resistor network. Wright et al. (2022, Nature, "Deep physical neural networks") trained arbitrary physical systems via physics-aware training. Lopez-Pastor and Marquardt (2023) proposed Hamiltonian echo backpropagation for time-reversible physical systems. This is still small-scale, but conceptually it closes the loop: the substrate's own dynamics does both inference and gradient estimation.
3.7 Thermodynamic limits
Landauer (1961) showed erasing one bit dissipates at least kT ln 2. Bennett (1982) showed reversible computation avoids it in principle. Modern digital logic sits roughly three to four orders of magnitude above the limit per operation; the brain, at about 20 W for perhaps 10^15 synaptic events per second, is closer but still far off. Wolpert and others have extended stochastic thermodynamics to compute the minimal cost of specific computations. The relevance here: relaxation-based computing pays its thermodynamic cost through dissipation during descent, and there is an unavoidable tradeoff between speed, accuracy, and heat. Memory as an attractor costs energy to maintain against noise and to overwrite.
3.8 Memory as attractor, not address
Pull these together and a definition falls out. In a store, a memory is a location whose contents are arbitrary with respect to the dynamics. In a Hopfield network, a memristive crossbar, or a cortex, a memory is a stable set of the substrate's own flow. Retrieval is convergence. Generalization is basin shape. Forgetting is loss of stability or basin capture by another attractor. Interference is spurious minima. This is not just terminology: it predicts that content-addressability comes for free, that capacity scales with structure not with bits, and that memory and inference are inseparable because both are the same relaxation.
4. The shared vocabulary, and where it breaks
4.1 What transfers cleanly
State space and flows. In all three domains, the useful move is to stop asking "what does each component do" and ask "what is the geometry of the trajectory."
Attractors and basins. Hopfield memories, market equilibria, Ising ground states. The formal object is identical. What differs is whether the system actually reaches it (brains and markets mostly do not sit still).
Order parameters and time-scale separation. Haken's slaving principle is why a population of neurons can be described by a few latent variables (Section 1.5), why a market can be described by prices rather than by every trader, and why a reservoir's readout only needs a few modes. The existence of an effective low-dimensional description is not an assumption; it is a theorem given a spectral gap.
Bifurcations. Kuramoto's onset of sync, the bounded-confidence fragmentation threshold, the HKB loss of phase locking, the fold that early-warning signals detect. Bifurcation theory is the classification of ways qualitative behavior can change, and it is substrate-independent.
Symmetry breaking. A ring attractor's bump breaks rotational symmetry; Schelling segregation breaks spatial symmetry; an Ising machine's ground state breaks spin-flip symmetry. In each case, noise or initial condition selects among degenerate solutions, and history matters afterward.
Noise-induced transitions. Kramers escape from a basin is the same calculation whether the well is a cortical attractor, a market regime, or a memristor state. Stochastic resonance and annealing exploit it.
Coarse-graining. Renormalization asks which microscopic details survive to the macroscopic description. The answer, relevant operators, is why so few parameters matter near a transition, and why very different systems share critical exponents.
4.2 Where it is analogy
Criticality is the clearest case of unequal evidence. Neuronal avalanches are measured with millisecond resolution over thousands of events with controlled exponents and shape-collapse tests; the debate is about interpretation, not data. "Market criticality" rests on fat tails and volatility clustering, which many non-critical mechanisms produce, and on log-periodic fits that have not survived rigorous out-of-sample testing. Brain criticality is a contested empirical hypothesis. Market criticality is closer to a heuristic.
Order parameters in social systems are often chosen rather than derived. Haken derived his from the spectrum of the linearized dynamics. "Polarization" or "sentiment" are picked because they are measurable, with no guarantee of closure.
Kuramoto for opinions is a metaphor. Neurons have phases; opinions do not, and the coupling function is not sinusoidal or symmetric. The bounded-confidence models are more honest because they were built from social assumptions, not borrowed.
Reflexivity has no clean physical analog. Spins do not model the magnet. This is the deepest asymmetry between social systems and the other two, and it undercuts any transfer of predictability claims.
Hardware relaxation is the domain where the vocabulary is literally true rather than a model. An Ising machine's energy function is not an analogy for its dynamics; it is its dynamics. The cost is that we get exactly the emergence we engineer and no more.
4.3 Emergence as a fact about descriptions
Dynamical systems theory makes emergence non-mysterious by making it about closure. A coarse-grained description is "emergent" when the coarse variables evolve according to a rule that is (approximately) closed: you can predict their future from their present without going back to the micro level. Slaving is one route to closure. Renormalization fixed points are another.
Erik Hoel, Albantakis, and Tononi (2013) made this quantitative with causal emergence: compute effective information at each level of coarse-graining; the macro level can have higher effective information than the micro when the micro is noisy or degenerate. Hoel (2017) tied this to error-correcting codes. Rosas, Mediano, Jensen, Seth, Barrett, Carhart-Harris, and Bor (2020) used partial information decomposition to define "causal decoupling," where a macro variable predicts its own future better than any micro variable does, and applied it to Conway's Life and to neural data. Both frameworks have critics: Dewhurst and Eva and others argue the effective-information gain depends on the choice of intervention distribution, and PID has non-unique definitions. But the framing is right. Emergence is a claim that some level has its own effective dynamics. It is testable, it comes in degrees, and it needs no new physics. Bedau's weak emergence, Haken's slaving, and Hoel's causal emergence are three views of one idea: macro variables that close over themselves.
5. Open questions at the intersections
Can hardware exploit metastability the way cortex seems to?
Every engineered relaxation system wants a fixed point. Cortex appears to want the opposite: saddles, ghosts, heteroclinic sequences that produce sequences and decisions. Rabinovich's channels have been simulated but not built. An oscillator-based or memristive system designed to follow a controlled itinerary between quasi-stable states, rather than to converge, would be a new computing primitive. The design question is how to make saddle dwell times programmable and robust to device noise.
Can social early-warning signals be made actionable?
Critical slowing down requires knowing the bifurcation type and having a long stationary baseline; social systems provide neither. Two directions look promising. First, use information-theoretic emergence measures (Rosas-style) to detect when a macro variable is becoming causally decoupled, which would indicate a regime is forming rather than simply that variance rose. Second, exploit reflexivity deliberately: an early warning that is published changes the system. Whether that stabilizes or destabilizes is an open control-theory question with real policy stakes.
What does "no separation of memory and compute" imply for language models and agents?
A transformer's attention is one step of Hopfield retrieval over the context, but its parameters are a frozen store and its context is a separate address space with explicit retrieval. Agents bolt on vector databases, which is a von Neumann architecture with extra steps. The dynamical view suggests memory should be a stable set of the network's own recurrent dynamics, updated by local rules during inference. Test-time training, fast weights (Schmidhuber's 1992 idea, revived by Ba et al. 2016 and in recent linear-attention work), and state-space models are partial moves. The open question is whether an agent whose memory is an attractor landscape can be made auditable, or whether address-based memory is the price of oversight.
Is near-criticality generic or tuned, and does it matter for design?
If Griffiths-phase behavior arises from heterogeneity alone, then reservoirs and neuromorphic chips could get critical-like benefits by design of connectivity statistics, without homeostatic tuning. If the brain actively regulates its distance to criticality, hardware would need an equivalent of synaptic scaling. Experiments that manipulate heterogeneity while holding mean coupling fixed, in both cultures and physical reservoirs, would discriminate the two.
Can physical learning rules discover the slow variables?
Equilibrium propagation and coupled learning adjust couplings given targets. Haken's slaving says the useful description is in the slow modes. A physical system that reshapes its own spectrum to expose a few slow, task-relevant modes would be doing dimensionality reduction in its dynamics, which is what motor cortex appears to do. Whether local rules can produce this without a global objective is unresolved, and it is the same question as whether societies can institutionalize their own order parameters.
Coda
Emergence is not a thing that happens. It is what a description looks like when time scales separate and a few variables close over themselves. Neurons, traders, and memristors are different in almost every particular, and identical in this one respect: the interesting variables are not the ones you started with. The discipline the shared language imposes is to say, in each case, which variables, which time scales, which bifurcation, and what evidence. Where that can be done, the transfer is real. Where it cannot, we are telling a story, and should say so.