Is Natural Selection a Product of Physics?
Life as entropy maximizing has the answer backwards.
For years I have been circling one question: is natural selection a phenomenon that stems from biology, or from physics that biology happens to exhibit? I wrote an essay on it in 2008, another in 2023, and last year a computational paper that was, in hindsight, a mess but made me rethink the premise I had been working on from the start. I had an assumption about entropy wrong, in the same way many people who love this topic have it wrong. What resulted is much simpler and has more implications.
The max entropy intuition, and why it is wrong
The assumption was that life is order i.e. low entropy and that this low entropy must be compensated for by higher entropy output (we shit more than we eat) and so life is perhaps entropy maximizing. Living things are just the universe’s most efficient way of turning sunlight into waste heat via complex, low entropy mechanisms.
Max entropy is elegant. It is also wrong, and the way it is wrong leads somewhere more interesting.
Each square metre of earth absorbs 240 watts of sunlight (averaged over day and night). This energy arrives as low entropy photons from a 5800 degree sun and is reradiated into space as many more, higher entropy, lower frequency infrared photons at about minus 18 degrees C. Call that entropy production per second sigma.
Sigma depends on two things: how much energy is input and the output temperature. The output temperature is fixed by the energy and the fact that energy in has to equal energy out. A body radiating 240 watts per square metre is at minus 18. This is the physics of glowing.
If there is life, such as a plant on the square metre, its leaves catch photons that would have hit dirt. The photon’s energy goes into sugar, then into soil, then into heat. Every step is a partial degradation. But at the end of the chain the same 240 watts leave as the same minus 18 degree infrared. Entropy in: unchanged. Entropy out: unchanged. Sigma: unchanged.
What life did was change where inside the square metre the degradation happens, and over what time. The rock did it all at the surface in an instant. The forest does some in leaves, some in creatures, some in fungus, over years. Same total.
There is no entropy maxing because life may be low entropy relative to dirt but it is much higher entropy than sunlight. Sunlight is absurdly low entropy; even very ordered things like diamonds, plants, animals and buildings have many times more entropy per joule. So when a joule of light becomes a joule of leaf, its entropy has gone up, not down. Everything we call order on Earth is more disordered, per joule, than the light that paid for it. Life is a step on the way down from sun to space, not up.
Dirt, Diamond and Vortex Earths
Imagine three planets, all absorbing the same 240 watts per square metre, all radiating at minus 18: one made of high entropy dirt, the other of low entropy diamond and the third a low entropy replenishing vortex, like a whirlpool.
The dirt planet: high entropy, nothing interesting. Entropy production rate is sigma, as sunlight warms the surface and it glows into space in the infrared.
The second is a single, massive diamond. Entropy production rate is still sigma, as sunlight warms the surface and it glows into space in the infrared. The diamond’s low entropy is a state, paid for once when it formed, and it costs nothing per second because a diamond does not decay on any timescale that matters. The diamond planet has vastly lower entropy than the dirt planet but exactly the same entropy production.
The third planet is a giant sustained whirlpool, driven by sunlight, an object made of nothing but flow, no diamond-like trapped order at all. Its low entropy is a standing pattern that decays continuously and is rebuilt continuously out of the absorbed 240 watts. Every second some of its order is lost to friction and every second the flow re-imposes it, dumping the disorder as heat. Its entropy per second: minus something, plus the same something, net zero. Its entropy production: still sigma, spread through the vortex over its turnover time rather than dumped at the surface. What is different about the whirlpool planet is not the total. It is that this arrangement has a running cost, and the running cost is paid out of the energy input, the flux. You can only have as much whirlpool as 240 watts will keep spinning.
Three planets, three wildly different entropies, one entropy production rate. But the third one is interesting. Unlike the diamond it is an object made purely of the shape of energy-driven flows rather than something sitting in them. Rather than being something that dictates that entropy output is maximised, a constant entropy output for a given energy flow dictates a cap on how much order it can have, so that the first and second laws of thermodynamics (energy conservation, entropy increase) are preserved for the overall sun, planet and space system. What makes the last case particularly interesting is that a channel in a flow turns out to be the only case where anything like selection can happen, where there are multiple channels competing for a share of energy flows.
While the diamond and the whirlpool are growing
The steady state is the easy case, but the same rules against increased entropy from localised production apply to growing systems.
While the diamond grows, slightly less than 240 watts leaves the planet; some of the sunlight is spent doing the pulling and ends up locked in bonds instead of degrading to heat, the shortfall being banked. Less energy is radiated, so LESS entropy is carried out with it. The planet’s own entropy is falling, the export is falling, and the second law is still satisfied, because the sunlight that drove the reaction was degraded on the way, from 5800 degrees to the 300 degrees of chemistry, and that degradation is produced inside, at the moment the photon’s low entropy is spent. When growth stops, the diamond is trapped order, costs nothing, and entropy production returns to sigma. Burn it and the banked joules come back as heat and the deficit is repaid. At no point does the planet export more entropy than the dirt one.
The whirlpool growing is the same in a different currency: energy goes into spinning up coherent motion instead of into heat, so production dips while it spins up. But the whirlpool is also decaying and being repaired every second even while it grows, and that repair is a slice of sigma spent inside the vortex. The diamond banks and then rests. The whirlpool banks a little and then rents.
In terms of steam engine physics, earth is a heat engine between the sun and space and life is the working fluid. A heat engine placed between a hot reservoir and a cold one, with a fixed heat flow between them, cannot increase the total entropy production. It can only match it, by wasting everything, or fall below it, by extracting work and storing or exporting order. Nothing life does at fixed inflow can push the total above what dead planets already achieve.
The 2nd law requirement for compensation for local order is real, but not extra. It’s drawn from a budget that was already being spent in full. A leaf pays for its order by degrading sunlight that would have been degraded on the dirt. It moved the spending; it did not raise it.
It becomes easier to comprehend once you consider that sunlight is lower entropy than anything it makes: a diamond, a leaf, a brain, a language: every one of them is a degraded form of the light that paid for it, a step down on the way from sun to space. “Life violates the second law” is not a paradox to be resolved; it is a category error. Life is what the second law looks like when light takes the scenic route from stars to space.
Increased entropy from increased input energy
There is one apparent way out of the fixed budget, and it matters. Structure can change how much comes in. A dark forest on pale ground absorbs light that would have been reflected, so absorbed power rises and sigma rises with it. Altering the earth’s albedo, and therefore how much sunlight is absorbed and reradiated later rather than directly reflected into space, can alter the entropy production, but it does that with more energy input. Fire spreading into fresh fuel unlocks an energy gradient that would otherwise have sat dormant.
But all these are increasing the entropy production by increasing the energy consumption on the input side. Structure acts on inputs and routes. It cannot act on outputs, because outputs are pinned by where the flow ends up, the sea, or space at minus 18. A water mill creates structured flows. It can dam a stream and deepen its channel. It cannot lower the sea.
Why this fixes what selection is made of
Natural selection is survival of the fittest in a competition for finite resources among organisms that are randomly mutated. Once you consider that entropy production is a fixed property of the energy gradient, the same for every arrangement of channels on it (where life can be seen as channels), you see that it cannot be what selection sorts by. Selection can only act on what differs between competitors, and total dissipation does not. What differs is share: how much of the flow passes through this channel rather than that one. This turns out to be enough.
Here is the mechanism, in a picture that has no biology in it.
Rain falls in a basin and many channels carry water out. Each channel has a width, so wider channels carry more water, in proportion. Each channel’s banks get knocked about at random, silting or scouring, and a channel carrying more water gets disturbed more. Now count two things.
Count channels, and over time most are narrow. A channel that silts up carries little, is rarely disturbed, and stays as it is. The population of channels drifts to where it is least disturbed. This is a result Rolf Landauer proved in 1975, sometimes called the blowtorch theorem, and it is all he claimed: where things pile up is set by the noise along the way, not by how good the destinations are.
Count water, and the answer is the opposite. Water follows width, and the few channels that happen to be wide at any moment carry most of it, precisely because they are the ones still being disturbed, still changing, still in play. A narrow channel is safe and carries nothing. A wide channel is at risk and carries the flow. So the channel census says trickles and the water census says a handful of rivers, and the gap between the two censuses widens over time, as long as the disturbance grows at least in proportion to width.
The tension between population and fitness is the shape Darwin described: many are born, a few carry the future, and the few are not a random sample. Here it appears with nothing alive, from a fixed flow, undirected noise, and one physical premise, that busier channels are disturbed more. Fitness is nothing but share of the flow.
And notice what the disturbance is. Each channel’s width is a local rule: how much of what arrives passes. Knocking the width is rewriting the rule at random. That is mutation, of a rule, at a rate set by the flow the rule carries. The channels are a population of rules under throughput-scaled mutation, which is also what a population of catalysts is: each catalyst sets how a flow runs, is not consumed by it, and is varied by the very activity it shapes.
Rain on a noisy plateau. Left and middle: when cutting rises with the water passing, a channel network carves itself and a few channels end up carrying most of the drainage. Right: when cutting depends on slope alone, the plateau erodes as a smooth sheet and nothing concentrates. Same rain, same land, one exponent.
Two more steps, still without biology. Let a channel wider than its stable size split, each branch inheriting its parent’s cross-section, or let a wide channel erode into its neighbour’s catchment and take its water. Both are copying at a rate that rises with share, and both make the basin’s channels trend wider. That is descent with modification, and it is Price’s equation with nothing alive. And if a channel’s own flow can rewrite its width, then among the channels whose flow scours them wider and the channels whose flow silts them narrower, the first kind takes over. There was no external input here: reading, in the sense of a channel consulting a state it wrote itself, is what selection does to any feedback whose sign is free.
In the simulations, give each channel a coupling that lets its own flow rewrite its state, with the sign free and starting at random, and within a few thousand steps nearly every surviving channel carries the positive sign. Nobody chose it. It is the sign whose consequences keep flowing.
I created small simulations, forty channels at a time, and each step behaved as the argument said.
Forty channels sharing a fixed flow, each knocked up or down at random. When the knocking rate rises with a channel’s share, three of the forty end up carrying half the flow and the fraction keeps climbing. When the knocking rate is the same for all, it stops at a fifth. Same kicks, same channels; the only difference is where the noise is.
Add one thing, a busy channel occasionally copying its state to a neighbour, and the whole population’s average state rises steadily instead of merely concentrating. Take away either the copying or the activity-scaled noise and it goes flat. That climb is descent with modification, with nothing alive.
To be more precise, what matters is not how often a channel is disturbed but how much: how often times how hard, squared, which I call “jitter”. Selection happens inside a window: the jitter must grow at least in proportion to the state (here, width), and not more than one power faster than the flow does. Too little state-dependent mutation from noise and nothing concentrates. Too much, and the high states are emptied so thoroughly that even the flow ends up down low. That is a two-sided phase diagram with two regimes where the theory predicts no selection at all, which is what makes it testable. In a real system you measure three things separately: how flow depends on state, how disturbance frequency depends on state, and how disturbance size does, and the theory says in advance whether selection should appear. It reproduces two known results, Zipf’s law for cities and the power laws of growing networks.
The whole result in one picture. Horizontal: how steeply a channel’s share of flow rises with its state. Vertical: how steeply its jitter rises. Yellow: a vanishing few channels carry half the flow. Dark: they don’t. The red line is where a real system sits if busier channels jitter more; it runs through the yellow wedge for every m above one. Cities, firms and growing networks, from published numbers, all land inside the wedge.Where life starts
Flows of water in channels can show almost everything in terms of selection. What they can’t do is carry state, like width, as a number. A channel’s state travels only with the channel, by splitting or conquest. It cannot be handed to a different channel, stored, varied, and handed on again. A living thing like grass does that: a tuft drops seed, and the seed carries the root plan without carrying the root. The state has become portable, and once it is portable it accumulates. Not just a few wide channels, but ones that are wider than any parent, generation after generation.
That single point is where life begins. Not at metabolism, not at reproduction, both of which a river manages in its own way, but at the point where a channel’s state can travel without the channel.
Back to the entropy discussion
I thought life was on the output side of the ledger, that it existed to maximise entropy, and that evolution’s arrow was toward more dissipation. It is on the input side. Life takes a share of a flow that was going to happen anyway, and competes for more of it. Total dissipation on a fixed gradient does not move; what moves is who is in the way of the flow, and how much of the leak they have taken.
Complexity rises not because complexity is more dissipative but because capturing share buys structure: a channel’s share of the energy flux is the budget from which it pays to maintain its own structure, and structure captures more share, until the leak is gone or the gradient fails, and then it stops, or reverses. To be precise about the budget: it is a share of the conserved energy flow. Entropy is not conserved and is not the budget; it is the bill.
Life is invisible in the planet’s entropy budget and loud in its spectrum. Invisible: a living Earth and a dead one at the same absorbed power radiate the same total infrared and produce the same entropy. Loud: a living Earth’s atmosphere holds oxygen and methane together, far out of chemical equilibrium, and its continents reflect sharply more near-infrared than visible light, the red edge of vegetation, both visible from light years away. Structure hides in the outputs and shows in the routes, which is also why searches for life elsewhere look for chemistry and colour, not waste heat.
Is this natural selection as physics? In the smallest systems that can carry it, via the simulations I have done, it is provably so. Natural selection is a theorem of non-equilibrium physics with a stated boundary, and the currency of that theorem is share of the energy flow, not the entropy produced.
The next step is to test against real world examples where you can weigh flow, disturbance rate and disturbance size against each other.
What is borrowed and what is new
Most of this is old. The occupancy law is Landauer’s. That selection tends to increase the energy flux through a system was conjectured by Alfred Lotka in 1922, and he went further than I do here, claiming the total rises, where my result is about how a fixed flux is shared. The three regimes of the window are known from the mathematics of growing networks. That where you put the mutation rate decides who wins has been known in evolutionary game theory since the 1990s. And the thermodynamics of the three Earths is textbook physics that I, and many others, had backwards.
What is new is small and specific. Multiply Landauer’s occupancy by the share law and the two censuses come apart; that gap has the form of selection; it survives only inside a finite window; and if jitter is pinned to throughput by physics rather than chosen by the modeller, the arbitrariness the game theorists found goes away and the exponent becomes something to measure. That, and reading a channel’s state as a mutable, heritable rule. It is a note with a test attached, not a revolution. It is also, I think, correct.
What is not here, and why
This essay is the linear half of the story, deliberately. The two lines in the next section need only that share rises at least in proportion to state, and nothing else about the channels. That is what makes them provable and that is why they are the part I am willing to stand on.
But almost everything interesting about real channels is nonlinear, and I have left it out because it is not yet at the same standard. In the simulations behind this piece, a flow pushed past a threshold cannot stay smooth and is forced to make a loop, and a flow with the momentum term removed never loops at any drive. A memory that is continuously rewritten only forgets more slowly; only a gated memory, one that writes big changes and ignores small ones, holds a pattern. Rain on a plateau carves channels only when cutting rises steeply enough with the water passing; a linear medium under a blinking gradient stores energy but builds no pattern at all. Four times over, structure needed a kink, and linear systems could not loop, remember, concentrate or compete. That is a set of boundaries, and it says what a substrate must have before any of this essay applies to it. It is the next piece.
Behind that sits the larger picture, which I have been circling since 2008 and which this work has made sharper without making it a theorem: that everything persistent in a flow is a pattern the flow keeps re-making; that a whirlpool and a strand of DNA are both ships of Theseus, and differ only in whether the rebuild consults the past; that a loop is the smallest thing with an inside; and that the place where a channel’s state can travel without the channel is the place where life starts.
Summary
The provable part of this essay, the linear result about shares and jitter, is two lines of mathematics.
Terms: g is a channel’s state, its width or conductance, whatever sets its share. p(g) is the fraction of channels at state g. D(g) is jitter: disturbance rate times disturbance size squared. share(g) is the fraction of the flow a channel at state g carries, rising as g to the power m. Jitter rises as g to the power q.
p(g) ∝ 1 / D(g)
Where the channels are. How often you find a channel at a given state is one over how much it jitters there. For a randomly kicked variable with no drift this is the standard stationary distribution, and it is the content of Landauer’s 1975 paper, where he stated it for a system with two stable states and a hot patch in between: occupancy follows the kinetics along the path, not the depth of the wells.
share(g) · p(g) ∝ g^(m − q)
Where the flow is. Multiply how much a channel at that state carries by how many channels are there, and you get where the flow actually is. Flow rises with state as g^m; jitter rises as g^q. When they rise together, q equals m, the product is flat: the flow is spread evenly across every state while the channels are piled at the bottom. Population and flow disagree, and the disagreement is selection. It holds when jitter grows at least in proportion to state and not more than one power faster than the flow: 1 ≤ q ≤ m + 1. Outside that window, nothing separates.
Everything else in this essay is a consequence of those two lines and one physical hypothesis: that busier channels jitter more. That is not a law; near equilibrium the fluctuation-dissipation theorem makes it plausible, far from equilibrium it is something to measure, and where it has been measured (cities, firms, networks) it holds. Landauer gave us the first line fifty years ago. The second is what you get when you remember that the channels are sharing something.
The full logic, the code, and the entropy ledger are on GitHub. Corrections welcome, especially from people who can measure jitter.




