Sources are cited in full below. The literature/ archive the citations refer to is a local folder of verified full texts, not redistributed here.

Three Schools, Three Shocks

Physics, system, and industry have each asked their own question of life. Three times a discovery let all three ask it at once. The third time stalled, and the reason is narrower than it looks.

Part I

The frame

Three questions, and the long prehistory they came from.

1Three questions, not three fields

Physics asks what life is as an object. System asks how life works as a machine. Industry asks how life could be useful as a tool.

These are not three disciplines. They are three questions, and each has its own idea of what would count as an answer.

The physics question wants an invariant. Something true of life the way thermodynamics is true of gases: a bound, a scaling, a law that does not care which organism you brought. The systems question wants a reason. Given that a cell does X, why is it built this way and not the other way, and what would go wrong if it were? The industry question wants a yield. Given that we want Y, what do we feed the cells and what do we change in them.

Because the answers look so different, the three schools can work in the same building for decades without meeting. What makes the history interesting is that three times they were forced to meet, each time by a discovery that made the same object visible to all three at once. What follows tracks those three meetings, and then argues that the third one broke down for one identifiable technical reason.

PHYSICS What is life, as an object? SYSTEM How does life work, as a machine? INDUSTRY How can life be useful, as a tool? ANTIQUITY 1900s 1970s 2000s form, flow, vital fluid the object has a nature to be named Schrödinger: an aperiodic crystal order from order; negative entropy kinetic proofreading Hopfield 1974 · Ninio 1975 equilibrium bound on fidelity is violated growth laws, proteome partition Schaechter–Maaløe–Kjeldgaard 1958 → Scott–Hwa 2010 fire, humours, the pump borrow the newest machine you have homeostasis, cybernetics Cannon, Wiener: regulation by feedback demand theory Savageau 1974, 1977 why activator here and repressor there network motifs, design principles Milo 2002 · Shen-Orr 2002 · Alon 2007 the cell as a circuit diagram beer, soy sauce, bread, the horse harness it long before explaining it industrial fermentation Pasteur, Buchner, penicillin, amino acids the cell as a chemical plant MCA 1973–74 · flux balance mid-1980s stoichiometry, yields, control coefficients genome-scale metabolic models Edwards & Palsson 2000 every annotated reaction, one matrix systems biology · synthetic biology · virtual cells
Figure 1. Three questions and their lineages. Each column is a question, not a department. Each question is answered differently in each era, and the era's answer is usually borrowed from whatever the era can build. The convergence at the bottom is the subject of Parts III and IV.
How this history is sourced

Every claim below that carries weight rests on a full text that was read, not on an abstract, a reference list, or a second-hand quotation. Papers are named in place and quoted where the wording matters, and the 107 entries in the reference list are all full texts.

Where the field's own memory of an episode is wrong about a date, a place or an attribution, the correction is made in the text at the point of use. Those corrections are usually the interesting part, so they are not tidied into footnotes. Three of them change the argument rather than just the details: §11 on where the growth laws came from, §12 on when the cooling started, and §18 on what the scaling failure actually measures.

2The questions are old, and the answers are borrowed

Each school's ancient form is not a quaint prelude. It is the same question with a worse vocabulary.

Take the systems question first, because its pattern is the clearest. "How does life work as a machine" has always been answered by reaching for the most impressive machine available. Life works like fire, an elemental transformation. Like a system of fluids, so that draining some blood might rebalance it. Like a pump, once Harvey had one. Like a clockwork automaton, once Descartes had one. Like a steam engine, once there were engines. Like a servo loop, once there were servos. Like a circuit, then a computer, then a neural network.

It is tempting to read this as a series of mistakes. It is better read as a series of instruments. Each machine metaphor imports a whole analytical apparatus along with the image, and the apparatus is what does the work. When Cannon named homeostasis and Wiener named cybernetics, the payload was not the word but the mathematics of negative feedback. The cost is that you also import the metaphor's blind spots, and those are invisible from inside. That is precisely the trap Part IV is about.

1971, both ideas in one book
Chapter titles: "self-constructing machines … self-reproducing machines … strange properties: invariance and teleonomy." And later: "the chemical invariants … DNA as the fundamental invariant … the translation of the code."
Jacques Monod, Chance and Necessity, 1971 PDF

That is worth pausing on. The two ideas this whole history turns on, that life is a machine and that life is chemistry all the way down, were canonised together, in print, by one of the architects of molecular biology, in 1970 and 1971. Monod's table of contents is the outline of the next fifty years.

The physics question has the same borrowing pattern with a different flavour. Its answers are attempts to name what kind of object life is: a form, a flow, a vital fluid, and then, once statistical mechanics existed, an ordered structure that resists the second law. Schrödinger's aperiodic crystal is the famous instance. What matters is the move, not the answer: pick a class of physical object, and see what follows from life belonging to it.

The industry question is the oldest of the three and the least troubled by any of this. Beer, wine, soy sauce, fermented fish, bread, cheese, the harnessed horse. All of it is life put to work with no theory of life at all. This is the school with the longest unbroken record of success, and it is the school least dependent on the other two. Remember that when we get to §17. When the theory-driven programme stalled, the field did not collapse. It fell back on the school that never needed the theory.

A caution about the machine metaphor

The machine framing has a serious modern critic. Nicholson (2019) gives the position a name, the machine conception of the cell, and states the target precisely: the conception "grounds the conviction that a cell's organization can be explained reductionistically, as well as the idea that its molecular pathways can be construed as deterministic circuits". His case is that single-molecule data have accumulated against it across four domains, and that what is emerging instead "emphasizes the dynamic, self-organizing nature of its constitution, the fluidity and plasticity of its components, and the stochasticity and non-linearity of its underlying processes".

His two stated reasons are worth keeping in view, because neither is about complexity. Cells "unlike machines, are self-organizing, fluid systems that maintain themselves in a steady state far from thermodynamic equilibrium", and at their scale they are "subject to very different physical conditions compared to macroscopic objects, like machines".

That critique is engaged rather than dismissed here, because the specific failure it names, stochastic and non-modular molecular behaviour construed as deterministic circuitry, is the same failure Part IV documents from inside the engineering programme. Note where the disagreement actually lies. Nicholson's target is the circuit analogy, and Part IV is 20 years of evidence that he is right about it. The reply offered here, in §21, is narrower than a defence of the metaphor: what survives is not the machine as circuit but the machine as binding and catalysis, a model class in which far-from-equilibrium operation and occupancy-dependent coupling are what the equations are made of rather than what they neglect.

Part II · Shock one

Biochemistry becomes writable

The 1970s did not discover that life is chemistry. It made the chemistry nameable, measurable, and cuttable. That is what let all three schools move at once.

3What the 1970s actually settled

The thesis was old. What was new was that you could now name the enzyme, write the reaction, and cut the DNA.

The 1970s are remembered as the decade the universality of biochemistry became a foundational consensus. The shape of that is right and the date needs care, because the intellectual content was won much earlier, in stages. Wöhler synthesised urea in 1828. Buchner fermented sugar with cell-free extract in 1897. Sumner crystallised urease in 1926 and enzymes became proteins. Watson and Crick in 1953. The genetic code closed by 1966. By any reasonable reading, "life is chemistry" was settled science before 1970.

So why does the 1970s feel like the moment? Three things converged, and none of them is the thesis itself.

  1. The last big mechanism closed. Mitchell's chemiosmotic hypothesis, proposed in 1961 and resisted for a decade, carried the field through the 1970s and took the 1978 Nobel. With oxidative phosphorylation explained, bioenergetics had no remaining place for anything but chemistry.
  2. The toolkit became universal. Restriction enzymes (1978 Nobel), recombinant DNA in 1972 and 1973, Asilomar in 1975, Sanger and Maxam–Gilbert sequencing in 1977. Chemistry stopped being only an explanation and became an intervention.
  3. The thesis got written down by its architects. Monod's Chance and Necessity (1970 in French, 1971 in English) and Jacob's La logique du vivant (1970) are the canonical statements, and they arrive precisely at the turn of the decade.
What the decade actually delivered

Read the 1970s not as the decade biology learned that life is chemistry, but as the decade that fact became operational. Before it, you could believe life was chemistry. After it, you could write the reaction down with names and numbers in it, and then go change the DNA that encoded it. Every one of the three programmes below depends on that, not on the metaphysics.

4Physics: kinetic proofreading

Take an equilibrium bound seriously, notice the measured value violates it, and conclude that something must be burning fuel.

This is the cleanest example in biology of the physics move. Here is the argument as Hopfield states it.

Discrimination between a correct substrate C and a near-identical wrong one D happens at a recognition site that makes the pathway energetically more favourable for C. In a simple scheme, the error frequency is bounded below by exp(−ΔGCD/RT), where ΔGCD is the largest free energy difference available along the path. That bound is a thermodynamic statement, and it is not negotiable.

Now put in the measured numbers. Protein synthesis inserts the wrong amino acid at most about 1 in 104 times. To buy that from a single equilibrium discrimination you need 5.5 kcal/mol, which Hopfield says is "often difficult to justify" for codon–anticodon binding or amino acid recognition. DNA replication is worse: the error rate is about 10−9.

The contradiction, stated
"It is often difficult to justify the 5.5 kcal (23 kJ) necessary to explain the known low error rates of 10−4 in protein synthesis … The situation is much worse in the case of DNA replication, where the error-rate is about 10−9."
Hopfield 1974, PNAS 71:4135, Department of Physics, Princeton, and Bell Laboratories PDF

The resolution: if the reaction is "strongly but nonspecifically driven, e.g., by phosphate hydrolysis", the same discrimination step can be applied twice, and the error becomes the square of the equilibrium floor. Two facts about this deserve emphasis. First, the mechanism is forced by the contradiction, not fitted to data. Second, Hopfield notes that the scheme he is displacing is "based on Michaelis kinetics", so the physics critique lands squarely on a Michaelis-kinetics account of the mechanism. Hold that thought until §19.

Equilibrium discrimination one recognition step, no fuel correct C bound wrong D bound ΔG error ≥ exp(−ΔG / RT) a hard floor. 10⁻⁴ costs 5.5 kcal/mol. measured: replication reaches 10⁻⁹. floor violated. Kinetic proofreading same recognition, applied twice, driven by NTP hydrolysis E·S E*·S P NTP → NDP reject 1 reject 2 error ≈ [exp(−ΔG / RT)]² two independent chances to reject the same wrong substrate the price is fuel. accuracy is bought, not given.
Figure 2. The proofreading argument. Left, the equilibrium floor. Right, the driven scheme that beats it by applying the same discrimination twice. The logic is entirely physics: an inequality is stated, a measurement violates it, and the violation forces a mechanism. Hopfield explicitly notes that the scheme being displaced was "based on Michaelis kinetics."

Ninio, and what "independently" means

Ninio's paper is real, contemporaneous, and different. It was submitted 12 December 1974, four months after Hopfield's 6 August 1974 contribution, from the Salk Institute, and published in Biochimie in 1975. Its own summary says it discusses "the relationship between our scheme for a delayed reaction and Hopfield's scheme", citing Ninio's own prior scheme in a CNRS volume then in press.

The nuance that matters: Ninio's argument is probabilistic and kinetic, framed as a "kinetic amplifier" arising from a delay in one step. He does not make the thermodynamic-contradiction argument. So "discovered independently" is fair on priority, and "the same argument" is not. Hopfield brought the physics. Ninio brought the mechanism-space analysis. It is a useful reminder that the two temperaments coexist even inside a single school.

The thread continued

The argument extends. If fidelity requires a nonequilibrium drive, the drive need not be chemical. Spatial gradients will do. That is Galstyan, Husain, Xiao, Murugan & Phillips (2020, eLife), with the time-domain companion in Xiao & Galstyan (2024, PNAS). The move is the same one Hopfield made: keep the bound, ask what currency can pay for beating it.

5System: demand theory and the inverse problem

See a design, ask what function it is optimal for, and let the answer be checkable against a census of real operons.

The systems move needs a target, and Savageau found an unusually clean one. A cell must turn a gene on when a signal is present. Two architectures do that. Activate an activator, or de-repress a repressor. Functionally they are interchangeable. So why does nature use one here and the other there?

The methodological problem is stated exactly in the 1977 paper, and it is the heart of the systems school.

Why you cannot just compare two operons
"One cannot draw conclusions about the differences between repressor and activator mechanisms by comparing directly two representative systems such as the inducible lactose (repressor-controlled) and maltose (activator-controlled) operons because there may be other (unknown) elements involved in their control, and because the systems differ in many ways that are irrelevant to the comparison … Ideally, one would like a controlled comparison in which the two systems are identical in every respect except one: the type of control mechanism utilized. Although this is difficult to obtain experimentally, it can be simulated by appropriate mathematical analysis."
Savageau 1977, PNAS 74:5647 PDF

The model is the instrument that makes the controlled comparison possible. That is the systems school's characteristic use of theory, and it is worth naming because it is not prediction. The model holds function fixed so that architecture can vary.

And the answer, in the paper's own words: regulation by a repressor is selected when there is low demand for expression in the organism's natural environment, and regulation by an activator when there is high demand. Use it or lose it.

The reasoning behind that rule is the interesting part, because three candidate objectives are eliminated before the surviving one.

F1 energy cost of making the regulator F2 individual robustness to loss-of-function mutation F3 population robustness to mutational takeover observed rule lac and mal operons ACTIVATOR signal present → make activator REPRESSOR signal absent → make repressor high demand low demand high demand low demand high costlow cost fragilerobust robustfragile ACTIVATORnot observed low costhigh cost robustfragile fragilerobust not observedREPRESSOR predicts the reverse predicts the reverse matches the target
Figure 3. Demand theory as an inverse optimality problem. Fix the primary function (signal on → gene on), then score each architecture against candidate secondary objectives. Energy cost and individual mutational robustness both predict the reverse of what lac and mal actually do. Only robustness at the population level survives: under high demand, an activator that mutates kills its own carrier, so the mutant is purged rather than accumulating. The theory's content is the elimination, not the surviving row.

Two refinements the full texts force.

The population argument is present in 1977, but qualitative. The 1977 paper argues through the fate of regulatory mutants and their "selective disadvantage". The explicit population-dynamic model, with mutation rates and growth rates for wild type and each mutant class across two environments, is the 1998 quantitative development. Both are in Part I and Part II of the 1998 pair.

The rule is contested, in a productive way. Gerland & Hwa (2009) examine how far the lethality assumption holds and find regimes where individual-level robustness dominates instead. Shinar, Dekel, Tlusty & Alon (2006) derive a related rule from error minimisation, and Sasson et al. (2012) tie mode of regulation to insulation of expression. The rule survives. Its unique explanation does not.

What the school lost by not doing experiments

Savageau's later work moved toward modelling and simulation and toward the power-law (S-system) formalism rather than toward experiments the theory implied. That is a real cost. Once molecular characterisation became routine in the 1980s, a theory that only postdicts operon census data has no remaining leverage. The place where a design rule still buys something is at the frontier of what can be built, because there the space of options is too large to search. One concrete form that could take: cycle antibiotic demand high and low, alternately introducing activator- and repressor-architecture competitor strains, and drive resistance out of a population by making each architecture accumulate the mutations its own demand regime tolerates. The rule stops being a postdiction about operons and becomes a design.

6Industry: the cell as a chemical plant

One idea, two programmes, two decades, and two opposite attitudes to rate constants.

The cell-as-chemical-plant move is usually placed in the late 1970s and early 1980s, producing flux balance analysis and metabolic control analysis together. It splits cleanly into two distinct programmes, with different mathematics and different decades.

Metabolic control analysis is 1973–74, and it is about sensitivity. Kacser & Burns and, independently, Heinrich & Rapoport asked how a change in one enzyme's activity propagates to the flux through a whole pathway. The answer is the control coefficient and the summation theorem, which say that control over flux is distributed and sums to one. This is a physiology question answered with calculus, and its most famous downstream consequence is Kacser & Burns's structural explanation of genetic dominance.

Flux balance analysis is the mid-1980s, and it is about stoichiometry. Varma & Palsson's 1994 review, which is the consolidating statement, traces the lineage to Papoutsakis 1984 on butyric-acid fermentations, Fell & Small 1986 on fat synthesis under stoichiometric constraints, and Watson 1984 and 1986. The framing is unmistakably industrial: "Flux balance methods only require information about metabolic reaction stoichiometry, metabolic requirements for growth, and the measurement of a few strain-specific parameters."

Why the split matters

MCA needs kinetics and gives you sensitivities. FBA refuses kinetics and gives you a feasible polytope of fluxes. The industrial school chose FBA, and it chose it precisely because it does not require rate laws or parameters. That choice is the single most consequential methodological decision in this history, and it looks prescient in §16 and constraining in §22. It also anticipates, by fifteen years, exactly the complaint Bailey would file in "Complex biology with no parameters" (2001).

The engineering discipline that grew from this was named by Bailey in 1991 as metabolic engineering, and codified in Stephanopoulos, Aristidou & Nielsen (1998). Bailey's own 2001 commentary is worth reading against this whole document, because it draws the distinction the rest of the argument turns on: the information needed to predict a biological system's behaviour splits into the systems-structural (which components, and which interactions among them) and the parametric (the rate constants and equilibrium constants). His bet was that structure carries much more than people assumed. Part V is an argument that he was right, and that the field bet the other way.

Part III · Shock two

The genome, and the whole-system view

In the 1970s the unit of analysis was a pathway. After the genome it was a cell. Each school reached for the whole at once, and two of them reached for it together.

7What the genome bought

Not new physics. A complete parts list, which is a different and in some ways more dangerous gift.

In the 2000s it became possible to see the whole cell, and that licensed a new level of ambition in all three schools. The mechanism is worth being precise about, because the parts list is exactly what makes the third shock's disappointment intelligible.

A sequenced, annotated genome gives you three things. An enumeration of the components. For metabolism, the stoichiometry of nearly every reaction, because gene annotation names the enzyme and the enzyme fixes the reaction. And a coordinate system in which to place measurements, which is what made the omics era possible.

What it does not give you is any rate law, any binding constant, any copy number, or any indication of which reactions matter under which condition. The genome delivers the structural half of Bailey's split and none of the parametric half. Every programme in this part is an attempt to get by on the structural half, and the differences between them are differences in how they handle that gap.

195019701980 2000201020162026 SHOCK 1 SHOCK 2 SHOCK 3 biochemistry becomes writable the genome the cooling PHYSICS SYSTEM INDUSTRY Schaechter–Maaløe 1958 proofreading 1974 chemotaxis, integral feedback 2000 growth laws 2010 demand theory 1977 motifs 2002 toggle + repressilator 2000 retroactivity, burden 2008–15 MCA 1973 flux balance mid-1980s genome-scale models 2000 six products on the market 2020 systems biology synthetic biology
Figure 4. The three shocks, and who moved when. Dots are the works discussed here, placed in their school's lane. Note the two convergences at 2000: the physics and systems lanes meeting over chemotaxis, and the systems and industry lanes meeting over engineered circuits. Note also that the industry lane's 2020 marker is the only one that is unambiguously a delivered product.

8Physics meets system: chemotaxis, end to end

One pathway followed from receptor to flagellar motor, with a system-level property proved to be architectural rather than tuned.

Chemotaxis is the culmination of the 1999–2000 moment, and the best-documented episode in the whole story. The sequence is worth laying out because it is a template.

  1. Behaviour, quantified. Berg & Brown 1972 track single cells in three dimensions. Segall, Block & Berg 1986 measure the impulse response and find its integral is zero. Adaptation is a measured fact, in 1986.
  2. The property is claimed to be robust. Barkai & Leibler 1997 argue that exact adaptation cannot come from parameter tuning, because it survives parameter variation.
  3. The claim is tested. Alon, Surette, Barkai & Leibler 1999 vary CheR expression over a 100-fold range. Adaptation time moves more than 20-fold, from 23 ± 2 minutes down to about 1 minute. Adaptation precision stays flat. One property is structural. Its neighbours are fragile.
  4. The structure is named. Yi, Huang, Simon & Doyle 2000 identify methylation as an integrator, with CheR and CheB saturated, and invoke the internal model principle: perfect adaptation if and only if integral feedback.

That is the physics school and the systems school doing one job. Step 3 is a physicist's move, distinguishing a robust observable from a tuned one. Step 4 is a control engineer's move, matching a measured invariance to a theorem about architecture. Neither is possible without the parts list, because the theorem is about the whole loop from receptor to motor.

The template, and its limits

This template is the high-water mark of the whole story, and it is worth asking why it did not generalise. It required a pathway with about a dozen named proteins, a single measurable output, a decade of quantitative behavioural data, and a system-level property sharp enough to state as a theorem. Very few biological systems come with all four. The full argument about which properties are structural and which are tuned, and what reaction orders have to do with it, is the subject of the companion essay, The Biomachine Perspective.

9System meets industry: the VLSI roadmap

The founding papers used Hill functions. The founding manifesto cited the VLSI textbook by name. Both facts matter later.

Two papers in Nature in January 2000 started the engineering programme. Elowitz & Leibler built the repressilator, three repressors in a cycle, and Gardner, Cantor & Collins built the toggle switch, two mutually repressing genes.

Both are modelled with Hill functions, and the details are worth having exactly, because they come back in §20.

The founding models, in their own parameters

Repressilator. mRNA and protein for each of three genes, with repression through a Hill term. Reported parameters: Hill coefficient n = 2, 20 proteins per transcript, protein half-life 10 min, mRNA half-life 2 min, and KM = 40 monomers per cell.

Toggle switch. Two repressors with cooperativity exponents β and γ. The paper's own bistability condition is a statement about those exponents: "at least one of the inhibitors must repress expression with cooperativity greater than one", with higher-order cooperativity enlarging the bistable region. The mechanism cited is cooperative binding of repressor multimers to multiple operator sites.

And in the textbook that taught the field, Alon's Introduction to Systems Biology, the Hill function appears about twenty times and is introduced exactly as you would expect: "The Hill function can be derived from considering the equilibrium binding of …".

So the field's dynamical vocabulary was fixed, from the first week, as Hill functions composed additively. Now the second half of the roadmap.

Endy's 2005 manifesto is the programmatic statement: standardisation, decoupling, abstraction, so that biology can be engineered the way other things are engineered. It cites Mead & Conway, Introduction to VLSI Systems, Addison-Wesley, 1980 by name, and the lesson it draws from that book is specific. Decoupling is illustrated by "very-large scale integrated (VLSI) electronics, which is an engineering technology that only became practical once rules were worked out to enable the separation of chip design from chip fabrication".

That provenance is not Endy's alone. Way, Collins, Keasling & Silver (2014) record that the framework came out of a 2003 DARPA-sponsored study which "resulted in a recommendation that the field of synthetic biology should learn from the earlier developments in the microchip industry", explicitly including MOSIS and the separation of design from manufacturing.

The roadmap was explicit, and it was self-consciously provisional

That the 2000s programme took electronics as its model, down to expecting a VLSI-style scaling curve, is not a retrospective simplification. The founding manifesto put the VLSI textbook in its bibliography and named design-fabrication decoupling as the transferable lesson. The rest of the programme follows faithfully: a parts registry, standard assembly (BioBricks), a reference unit for promoter strength (relative promoter units, Kelly et al. 2009), and component datasheets (Brophy & Voigt 2014).

One correction the full text forces, in Endy's favour. He hedges the whole triad in the paper itself: "To be clear, these ideas and my prioritization of them could be wrong; I would explicitly encourage the widespread invention and discussion of alternative ideas." The manifesto is not the naive document the field's later self-criticism sometimes implies. It also names evolution as the fourth challenge and says plainly that it "is largely unaddressed within past engineering experience".

Two details in the manifesto are worth extracting, because later sections are about their failure.

Two things Endy 2005 asked for, in his own words

The chassis, defined as independence. "One engineer might develop standard 'power supply and chassis' cells that provide known rates of nucleotides, amino acids and other resources to any engineered biological system placed within the cell, independent of the details of the system." That last clause is the assumption. §14 is six independent measurements of it failing, and §19 shows it is the same inequality as retroactivity. Endy's Figure 2 makes the strong version of the claim visually: in his abstraction hierarchy, "abstraction barriers block all exchange of information between levels".

A target of one thousand. "Implicit in this hierarchy are formidable molecular engineering challenges; for example, engineering a set of 1,000 synthetic transcription factors … each recognizing a unique cognate DNA binding site with >99% specificity." That is a numeric goal, set in 2005, for the parts layer. Twenty years later, Cello's characterised gate library is about twenty repressors and the largest circuit uses ten of them. §18 is that arithmetic.

Endy also noticed, correctly, the one curve that did behave: "Bulk DNA synthesis capacity appears to have doubled every 18 months or so for the last ten years." That trend continued for another two decades. The design layer built on top of it did not. Figure 7 is that contrast.

It is worth being fair to the analogy, and the fair point deserves more than a concession. A resistor is not a law of nature. It is a manufactured object whose linearity is the product of decades of materials work, and the same is true of a transistor with a clean transfer curve. The 2000s programme was not wrong to want components. It was wrong about what stage of the process it was at. Electronics had its abstraction because someone first built the parts that deserved it.

10All three at once: genome-scale metabolism

The parts list plus stoichiometry plus linear programming, and a real prediction accuracy to report.

Edwards & Palsson (2000, PNAS) is the moment the industrial school cashed the genome. The annotated E. coli MG1655 sequence plus biochemical information gives a genome-specific stoichiometric matrix, and that matrix plus capacity constraints defines the space of achievable flux distributions. They then knock out genes in central metabolism, in silico and in vivo, and compare.

The number is 86% of examined cases predicted qualitatively correctly for growth or no growth. That is a real, checkable, useful result, and it was obtained with no rate laws at all. The refusal of kinetics was a strength here.

It is also the ceiling. A stoichiometric model can tell you what is impossible and what the best conceivable yield is. It cannot tell you what a cell will actually do, when, or how fast, because those are the questions kinetics answers. Orth, Thiele & Palsson (2010) are clear about this in the canonical primer, and §16 returns to what the ceiling has meant in practice.

11The fourth programme: growth laws

Copenhagen, not Amsterdam. Salmonella, not E. coli. Exponential in 1958, linear in the right ratio by 2010. And it matters more to this story than it looks.

The story usually told is that a Maaløe-era group showed RNA fraction is linearly proportional to bacterial growth rate across media, that the thread went quiet, and that Hwa picked it up in the 2000s with omics tools. The shape is right. Four details are wrong, and one of them changes the argument.

Four corrections to the usual telling
  • The place. Schaechter, Maaløe & Kjeldgaard were at the State Serum Institute, Copenhagen, Denmark. Not AMOLF, which is in Amsterdam and is a different institution with a different lineage. The confusion is easy and the correction is firm: the paper's own byline says Copenhagen.
  • The organism. Salmonella typhimurium, not E. coli.
  • The functional form. The 1958 papers report cell mass, nuclei per cell, and RNA and DNA content as exponential functions of growth rate. The linear law is a later and different statement.
  • What the linear law actually is. Scott, Gunderson, Mateescu, Zhang & Hwa (2010) establish that the RNA/protein ratio is linearly correlated with growth rate, and explain why: about 85% of RNA is rRNA in ribosomes, so the ratio measures ribosome content, and the linear relation follows if ribosomes are growth-limiting and translate at a constant rate.

Now the part that changes the argument. Scott et al. 2010 is not a side thread. Read its abstract as a document about synthetic biology, because that is partly what it is.

A growth-law paper, on the subject of engineered circuits
"Elucidating these relations is important both for understanding the physiological functions of endogenous genetic circuits and for designing robust synthetic systems. … Endogenous and synthetic genetic circuits can be strongly affected by the physiological states of the organism, resulting in unpredictable outcomes. … The use of such empirical relations, analogous to phenomenological laws, may facilitate our understanding and manipulation of complex biological systems before underlying regulatory circuits are elucidated."
Scott et al. 2010, Science 330:1099 PDF

That last clause is a complete research strategy, and it is a rival to the one argued for here. Hwa's answer to "we have no foundational model class" is: do not wait for one. Find phenomenological laws that hold across conditions, and use them. It worked. The linear growth law, overflow metabolism, proteome partition, and the quantitative response to translation-inhibiting antibiotics all came out of it, and Klumpp, Zhang & Hwa (2009) turned it directly on the gene-expression side effects that were about to blindside the circuit engineers.

So the correct account of the 2000s has four programmes, not three, and the fourth had a different theory of what theory is for. §22 puts it side by side with the others.

Part IV · Shock three

The cooling

The third shock is not a discovery. It is the arrival of systematic evidence that the parts do not compose and the host is not a chassis. This part gives that evidence with its numbers, and then gives the counter-evidence too.

12Dating the turn

The plateau was measured in 2009 and named in Nature in January 2010. What 2012 to 2016 added was not the diagnosis but the numbers.

The cooling is usually dated to roughly 2013–2016. That is right about when the quantitative evidence landed, and about three years late for the diagnosis. The turn has three dates, not one, and the first is the one that matters most.

2009, the measurement. Purnick & Weiss counted regulatory regions per published circuit from 2000 to 2008, found a plateau, and proposed that the design principles were too simplistic. §18 treats this as the central piece of evidence in the whole story. Everything that follows in this part is the field working out why that plateau was there.

2010, the year it went public and the year the field split. Two things happened, both in Nature, pointing in opposite directions. In January, Kwok's news feature "Five hard truths for synthetic biology" reported the plateau as news and enumerated the reasons. In December, Elowitz & Lim published "Build life to understand it", which reframes the entire enterprise.

Kwok's five truths deserve to be listed in full, because four of them are the entire diagnosis this essay reaches, and because they were in print in January 2010.

Table 1. Kwok's five hard truths, January 2010. Her article is a news feature, so each truth is a reported consensus rather than a new result. All five survive the next decade's evidence, and that evidence is what the right-hand column and the section links point to.
Kwok 2010What the later evidence did to it
Many of the parts are undefinedQuantified in 2013 and 2016. Kosuri et al. measured 12,563 promoter-and-RBS combinations, Beal et al. ran an interlaboratory reproducibility study, and Mutalik et al. had to engineer reliability in rather than find it (§13).
The circuitry is unpredictableKwok's example is Collins spending "about three years of tweaking" to move a toggle switch into yeast, because the two promoters were unbalanced. That is composition failing at n = 2 (§13, §16).
The complexity is unwieldyKwok reports Keasling's estimate of roughly 150 person-years for the artemisinin pathway, about a dozen genes. Ten years later the largest dynamical circuit is fourteen regulators (§18).
Many parts are incompatibleThis is growth burden, already correctly identified. Kwok cites Tan, Marguet & You (2009), where a circuit slows growth, growth slows dilution, and bistability appears from nowhere (§14, §15).
Variability crashes the systemNoise, plus evolutionary loss of function. Sleight & Sauro (2013) evolved 192 populations and found circuits expressed below 10% of maximum are significantly more stable, a fitness threshold for keeping a circuit at all (§17).
The line, and who actually said it
"The field has had its hype phase. … Now it needs to deliver."
Martin Fussenegger, quoted in Kwok 2010, Nature 463:288 PDF. The full text corrects an attribution: the line is Fussenegger's, in Kwok's closing section, not Kwok's own. Kwok's own summary of the same section is blunter, that "the field has yet to deliver much of practical use". Voigt (2020) calls the piece "an infamous article in 2010" and adds: "Early research struggled to design cells and physically build DNA with pre-2010 projects often failing due to uncertainty and variability."

And the reframing, from two of the field's most prominent circuit builders, in the same year:

The pivot, dated 2010
"Synthetic biology is redefining the discipline of biology and helping people reach a deeper understanding of how life works. … Traditional biologists seek to reverse engineer natural biological systems. Synthetic biologists seem to do the opposite. … Both communities face the same daunting challenge: how to relate the architecture of a gene circuit to its behaviour in a cell or tissue."
Elowitz & Lim 2010, Nature 468:889 TXT

So the field's justification inverted, from "understand biology in order to engineer it" to "engineer it in order to understand biology". The inversion has a precise date: December 2010, in a Nature comment, argued for on the merits rather than as a retreat. Note also that Khalil & Collins published "Synthetic biology: applications come of age" in the same year, saying the opposite. 2010 is the year the field split on whether it was delivering.

There is one genuine complication, and it is worth more than the tidy version. The field's own histories do not agree on when the trouble was. Cameron, Bashor & Collins (2014) divide the field into four periods, and their label for 2004 to 2007 is "expansion and growing pains", "characterized by an expansion of the field but a lag in engineering advances". Their label for 2008 to 2013 is "increase in pace and scale". So Collins's group puts the lag before Purnick & Weiss measured it, and calls the years usually identified as the cooling an acceleration.

Three datings, one resolution

Purnick & Weiss say complexity plateaued through 2008. Cameron et al. say engineering lagged in 2004 to 2007 and then accelerated. The field's working memory says it cooled in 2013 to 2016. These are not three guesses at one date. They are measurements of different layers.

Read what Cameron et al. actually list as the acceleration: Gibson assembly, MAGE, complete Boolean gate sets, genome synthesis. That is throughput, construction and library scale, and it did accelerate, exactly as Figure 7's sequencing and synthesis curves show. Purnick & Weiss are counting regulators in one designed circuit, which is the design layer, and that is the curve that stayed flat. The 2013-to-2016 dating is a third thing again: it is when the explanations for the flat curve were quantified.

So the field's chronologies conflict only if you assume it has one scaling curve. It has at least two, and they diverge. That divergence is the subject of §18.

The honest arc

The diagnosis was not discovered in 2013. It was available, in Nature, in January 2010, complete with the growth-burden mechanism and a measured plateau behind it. Truths 2, 3 and 4 are the whole account of what went wrong.

So the honest arc is not "the field cooled when it discovered the problem". It is closer to: the field named the problem early, correctly, and in public, and then spent a decade quantifying it while continuing to build. That is a more interesting story, and it puts more weight on the question of why naming the problem did not solve it. §19 onward is an answer to that.

2012 to 2016, the quantification. This is where the usual dating is right, and the next five sections are that evidence. What arrives in these years is not a new insight. It is measurement: how far off composition actually is, how much a circuit actually costs its host, and how fast an unburdened design is actually lost to evolution.

13The parts do not compose

12,563 promoter and ribosome-binding-site combinations, measured. The conclusion the authors drew was to stop predicting.

The single most decisive paper is Kosuri, Goodman, Cambray, Mutalik, Gao, Arkin, Endy & Church (2013, PNAS). Note that author list before reading the result. Arkin, Endy and Church, three of the architects of the standardisation programme, are on the paper that measures whether standardisation works.

They synthesised 12,563 combinations of common promoters and RBSs and measured DNA, RNA and protein for the whole library at once. Against a simple multiplicative model:

Kosuri et al. 2013 · 12,563 promoter × RBS combinations fraction of constructs whose measured level fell within twofold of the model prediction RNA 80% protein 64% worst 5% off by 13× on average, not worst case "could hinder large-scale genetic engineering" Beal et al. 2016 · 88 institutions, three constitutive constructs standard deviation of the measured ratio between two promoters, across labs 1.54× two strong promoters 5.75× strongest / weakest promoter Host strain did not matter. Choice of instrument did. This is the simplest possible measurement in the field.
Figure 5. Two composability results. Top, the direct test of whether characterised parts combine predictably. Bottom, the first large interlaboratory study, which asked only whether 88 labs measuring the same three constitutive constructs would agree on the ratio between them. The answer in both cases is "roughly, most of the time, with a bad tail."

The most important sentence in Kosuri et al. is not a number. It is the conclusion:

The strategic surrender
"The ease and scale of this approach indicates that rather than relying on prediction or standardization, we can screen synthetic libraries for desired behavior."
Kosuri et al. 2013, PNAS 110:14024 PDF

Read that against Endy 2005. The 2005 programme was standardisation so that composition would be predictable. The 2013 verdict, from partly the same people, is that composition is not predictable enough, so build a library and screen. The field did not merely find composition hard. It changed strategy away from prediction.

The companion papers from the same BIOFAB collaboration, Mutalik et al. (2013a) and (2013b), are the constructive answer, and the full texts sharpen the point rather than softening it. Two numbers side by side.

Prediction against insulation, 2013

Predicting. "The best available computational tool for designing context-optimized translation control elements for use in E. coli gives an ~47% chance to design elements that express proteins to within twofold of a target expression level." At the height of the standardisation programme, the state of the art in forward design was a coin flip.

Insulating. Mutalik et al. reach ~93% within a twofold window. They get there by changing the architecture, not the model: their bicistronic design buries the ribosome-binding site inside a short leader peptide's coding sequence, so that the site cannot form secondary structure with whatever gene follows it. The interference is not predicted. It is engineered out of reach.

That pattern is worth naming, because it recurs. Faced with a composition failure, the field's successful response has repeatedly been to build a physical insulator rather than a better model: bicistronic designs here, insulating load drivers for retroactivity in §15, copy-number-invariant promoters for plasmid variation, orthogonal ribosomes for host coupling. Each works. None of them composes, in the sense that none makes the next circuit predictable.

The field also has a precise name for what is missing here, coined by four of its leaders. Way, Collins, Keasling & Silver (2014) propose distinguishing concept modularity from engineering modularity:

The distinction the field drew for itself
"An open question is whether biology is genuinely modular in an engineering sense or whether modularity is only a human construct that helps us understand biology. … For example, to understand the process of translation, it is useful to think of a ribosome binding site and a coding sequence as separate modules. However, from an engineering point of view, these elements are not necessarily distinct modules because they can interact through mRNA secondary structure, with levels of translation resulting from a nonlinear combination of the two elements."
Way, Collins, Keasling & Silver 2014, Cell 157:151 PDF

That is the sharpest available statement of "genes turned out not to be modular", and it locates the failure correctly. The gene is a real conceptual unit. It is not an interface with a defined signature. The same paper adds an observation that explains why some things are composable: the modularity of operons and pathogenicity islands "is a consequence of natural selection", because those units were selected for horizontal transfer between hosts. Where biology is modular, selection put the modularity there. It is not a generic property of genes, and it is not guaranteed to sit where a designer wants a boundary.

14The host is not a chassis

Put a circuit in a cell and the cell pushes back. Six independent measurements of the same fact, against an assumption the field wrote down in 2005.

The 2000s picture was that a wild-type cell is a chassis, and the circuit is what you bolt on. That picture inverted, and the chassis turned out to be doing the work. This is the one place where the assumption was written down explicitly at the start, which makes the test unusually clean. Endy 2005 defines the chassis as a cell providing resources "to any engineered biological system placed within the cell, independent of the details of the system". Every result below measures a dependence on the details of the system.

Here is the evidence, in increasing order of how much it hurts the modular picture.

1. Expression costs capacity, measurably. Ceroni, Algar, Stan & Ellis (2015) built an in vivo capacity monitor and used it to rank constructs by burden. Their framing: "all heterologous expression represents an unnatural load, consuming cellular resources and leading to decreased growth rates that can predispose synthetic constructs to evolutionary instability." Growth burden and mutational escape are the same problem seen at two timescales.

2. Two unconnected genes are quantitatively coupled. Gyorgy et al. (2015) put one constitutive and one inducible reporter on the same plasmid, with no regulatory path between them, and found the attainable pairs of protein concentrations constrained by a linear relation. They call it an isocost line, borrowing from microeconomics. Changing the inducible gene's RBS strength rotates the line. Changing plasmid copy number shifts it.

3. The coupling can invert a circuit's function. Qian, Huang, Jiménez & Del Vecchio (2017) built a library of genetic activation cascades and tuned each gene's resource demand explicitly. Resource competition creates "non-regulatory interactions among genes", and in an activation cascade these produce responses that are "biphasic or monotonically decreasing". An activation cascade that goes down when you induce it is not a mis-tuned module. It is a module whose sign is not its own.

4. Growth itself is a feedback loop through the circuit. Tan, Marguet & You (2009) found bistability emerging from a circuit that has no bistable motif, purely because expression slows growth and slower growth dilutes less. Klumpp, Zhang & Hwa (2009) had already worked out the general growth-rate dependence of gene expression from the physiology side.

5. Past a threshold, both the circuit and the cell collapse together. Nikolados, Weiße, Ceroni & Oyarzún (2019) coupled circuit models to a host resource-allocation model and found the coupling is not gentle. Increasing induction buys protein output at the cost of growth, but only up to a point: "cellular capacity reaches a tipping point, beyond which both gene expression and growth rate drop sharply", from ribosomal scarcity. Their conclusion is that burden can "limit, shape and even break down circuit function". A design margin that assumes graceful degradation is the wrong margin.

6. Evolution deletes what it cannot afford. Sleight & Sauro (2013) evolved 192 populations of randomised three-colour circuits and read stability off directly. Circuits expressing below 10% of maximum were significantly more evolutionarily stable, which they read as a fitness threshold for retaining a circuit at all. Note what this does to the design problem: the usable region is bounded above by burden, and burden is set by the rest of the cell, so the specification cannot be written for the circuit alone.

Cardinale & Arkin (2012) is the paper that named this class of problem and tried to classify it. Their opening sentence is the field's mood in one line: despite the effort spent designing biological processes that follow predetermined rules, "their operation remains fundamentally circumstantial."

Six measurements, one assumption

Six results, six methods, one assumption. And they fail it in three different regimes: the dependence is there at steady state (1, 2, 3, 4), it is non-monotonic near the capacity limit (5), and it is enforced by selection over generations (6). "Independent of the details of the system" is not an approximation that holds in a restricted range. There is no range in which it holds.

15Retroactivity, honestly

It is often waved away as trivial. That is half right, and the half that is wrong is the important half.

Retroactivity, from Del Vecchio, Ninfa & Sontag (2008), is the back-action of a downstream module on its upstream driver. A transcription factor that regulates a downstream promoter must bind it, and bound molecules are not free, so the downstream load changes the upstream signal.

The standard objection is that this is obvious. Of course binding titrates the input. Calling it a deviation from modularity dresses up bookkeeping as a discovery.

The objection has force, and the mechanism really is that simple. But two things make it more than bookkeeping, and both are worth conceding before §19 uses them.

First, the effect is measured and it is large. Brewster et al. (2014) and Lee & Maheshri (2012) show that decoy and competing binding sites reshape regulatory input functions substantially, in bacteria and yeast respectively. Whether or not it is conceptually surprising, it is quantitatively decisive.

Second, and this is the part the "it's trivial" reading understates: it is a failure of the input–output abstraction, not an error in a parameter. If a module's transfer function changes when you attach a load, then a module does not have a transfer function. No amount of better parameter estimation repairs that, because the object being estimated does not exist independently of its context. This is exactly why the field's response was engineering rather than measurement: retroactivity attenuation, load drivers, promoters engineered for copy-number invariance. Those are attempts to build the conditions under which the abstraction becomes true, which is the correct response and is precisely what electronics did with resistors.

Where this leaves the objection

Sustained, with an amendment. Retroactivity is a symptom rather than a discovery, and the insulator programme did not deliver scaling. What the dismissal undersells is the diagnostic value of the symptom. Retroactivity is the cleanest available demonstration that the modelling abstraction, not the biology, is what broke, and §19 shows that burden is the same demonstration wearing different clothes.

16The models did not become design tools

Whole-cell models turned out to be excellent at something other than prediction, and the honest version of that is more interesting than "they failed".

Karr et al. (2012) built a whole-cell model of Mycoplasma genitalium accounting for every annotated gene function. Its claimed successes are real and specific: in vivo rates of protein–DNA association, an inverse relationship between the durations of replication initiation and replication, and previously undetected kinetic parameters found by experiments the model directed.

The usual verdict on such models is that they do not really predict much and are not much use for guiding experiments. The full texts support a sharper and less dismissive statement. Look at what Macklin et al. (2020, Science) actually report for E. coli.

What a whole-cell model is really good for
"… the total output of the ribosomes and RNA polymerases described by the data are not sufficient for a cell to reproduce measured doubling times, that measured metabolic parameters are neither fully compatible with each other nor with overall growth, and that essential proteins are absent during the cell cycle — and the cell is robust to this absence. After correcting for these inconsistencies, the model is capable of validatable predictions compared with previously withheld data."
Macklin et al. 2020 PDF

The finding is that the literature's own parameters are mutually inconsistent, and the whole-cell model is the instrument that detects it. That is a consistency checker on the field's measurements, which is genuinely valuable and is not what anyone was promised. Ahn-Horst et al. (2022) continue the programme in the same spirit.

For FBA, the honest summary is the one its own community gives. Lewis, Nagarajan & Palsson (2012) survey the method family, O'Brien, Monk & Palsson (2015) survey what genome-scale models can predict, and Seif & Palsson (2021) write a whole Perspective on improving model quality and lifecycle, which is not a paper you write about a mature design tool.

The generalisable lesson

Across all three modelling programmes the pattern is the same. Each produced a genuine scientific instrument and none produced a design tool. The instrument function is: reveal that your measurements disagree with each other. The design function requires something the instruments do not have, which is a guarantee that the model class contains the real system. §21 is about that gap.

17Where everyone went

Into domains where the cell does the work. And, decisively, into two technologies that succeed by not needing a model at all.

The diaspora is a matter of record: CAR-T and immune engineering, gut microbiome, rhizosphere and soil, bacterial cancer therapy, fermentation and metabolic engineering, antibiotic persistence. What the literature adds is that in each case the cell supplies most of the capability. Roybal et al. (2016) is combinatorial antigen sensing in T cells, where the T cell supplies essentially all of the capability and the engineered logic is an AND gate. Din et al. (2016) is synchronised bacterial lysis for in vivo delivery, where the payload is the point and the circuit is a quorum-triggered lysis timer.

And underneath the diaspora is the sharpest fact in this whole history.

The field routed around the missing theory

Voigt (2020) lists six commercially available products as the decade's deliverables, and lists the technologies that enabled them: metabolic engineering, directed evolution (2018 Nobel), automated strain engineering, metagenomic discovery, gene circuit design, and genome editing (2020 Nobel).

The two Nobel-winning technologies of the era both work by eliminating the need for a predictive dynamical model. Directed evolution replaces design with selection: you do not need to know the transfer function if you can screen a million variants. Genome editing replaces circuit design with genome writing: you do not need a theory of composition if you are editing one thing at a time in a system that already works.

Of the six products, three are purified chemicals from engineered cells or enzymes, and three are the engineered cells themselves. In none of them is a multi-gene dynamical circuit the source of value.

The field's flagship case is worth following to the end, because all three schools' verdicts on it differ and the full texts settle which one is right.

Artemisinin, 2006 to 2016, in four numbers

100 mg/L. Ro et al. (2006) reach "high titres (up to 100 mg l−1) of artemisinic acid" in engineered yeast. This is the paper the field points to as the industry school's proof.

25 g/L. Paddon et al. (2013) report fermentation titres of 25 grams per litre. A 250-fold improvement in seven years. As a feat of strain engineering this is unambiguous success.

150 person-years. Keasling's own estimate of the cost, reported in Kwok (2010), for a pathway of about a dozen genes. Kwok files it under "the complexity is unwieldy".

$250 against $350. Peplow (2016): Sanofi made enough semi-synthetic artemisinin for about 10% of global combination-therapy demand but "has not sold" it. For two years the plant-derived chemical sold for "less than $250 per kilogram, below Sanofi's 'no profit–no loss' margin of around $350–400 per kilogram". Demand stopped rising, and by July 2016 Sanofi was completing the sale of the production site. Keasling's own verdict, in the same piece: "If that price is already very low and there's a bumper crop, there's no reason to fire up a fermenter."

Read the four together and the lesson is not about biology at all. The technical programme worked, spectacularly, by screening and iteration. What it did not do was change the economics, because the competing process is a plant. This matters in a specific way. The industry school's success criterion never required a predictive dynamical model either, and its failure mode was not a modelling failure. Both of the era's headline outcomes, the commercial ones and the circuit ones, are silent on whether a theory of composition would have helped. Only §18's curve speaks to that.

how much it needs a predictive dynamical model → how much it scaled → SCALED WITHOUT THEORY NEEDS THEORY, DID NOT SCALE genome editing (2020 Nobel) directed evolution (2018 Nobel) DNA synthesis and assembly pathway transfer, strain engineering protein structure prediction learned protein design genome-scale metabolic models multi-gene dynamical circuits whole-cell models the frontier the field has not crossed
Figure 6. What scaled, and what it needed. Placement is qualitative and is the argument of this section, not a measurement. The pattern is that everything above the dashed frontier either replaced prediction with selection, replaced composition with one-at-a-time editing, or bought its predictive power from a learned statistical model over a huge dataset rather than from a mechanistic one. Protein structure prediction is the interesting case: it crossed the frontier, but with a learned model class rather than a derived one. §24 returns to what that implies.

18It grew, but it did not scale

Twelve repressors and two activators in one cell, twenty years after three. That is a real gain and it is not scaling, and the field has been keeping score with exactly this metric since 2009.

The claim is that circuit engineering never scaled in the sense that electronics scaled. It is easy to get this wrong in an obvious way, by pointing at the rise from three regulators to fourteen and calling the claim refuted. That applies the wrong test. The claim is about a rate, not about whether any increase occurred, and rates can be measured. Two things make the measurement possible rather than rhetorical. First, there is an agreed metric: Purnick & Weiss (2009) defined circuit complexity as the number of regulatory regions in one designed system, and everyone since has kept score that way. Second, scaling has a definition: a fixed doubling period, sustained.

A  Growth of designed complexity, each field from its own year zero 1 102 104 106 108 1010 1012 0 10 20 30 40 50 60 years since the field’s first artefact components in one designed system doubling every 2 years year 20 29,000 transistors 14 regulators same field age, 2,000× fewer parts repressilator the 2009 plateau Shin 2020 Genetic circuits, regulators per cell Integrated circuits, transistors per chip Metabolic pathway genes per strain DNA synthesis, bp per construct DNA sequencing, bases per run B  Sustained rate of doubling Moore-like Genetic circuits, 2016–2020 8.2 Genetic circuits, 2000–2020 7.1 DNA synthesis 3.1 Integrated circuits 1.7 DNA sequencing 1.4 0 3 6 9 years per doubling
Figure 7. Components per designed system, each field measured from its own year zero. Every genetic-circuit point comes from a paper in literature/: the toggle switch and repressilator (2000), the plateau value implied by Nielsen et al.'s "doubling a plateau first noted in 2009", Cello's largest circuit (2016), and Shin et al.'s decoder (2020). Transistor counts, instrument read lengths and synthesis milestones are public records. Units differ between series, so only the slopes are comparable, with one exception that matters: transistors per chip and regulators per cell are both counts of independently specified components, so those two compare directly. Metabolic pathway genes per strain are shown because they are the same kind of count for work that needs no dynamical model. Caveat: the circuit series tracks the largest single-cell dynamical circuit I could source. Distributing a design across strains, and static pathway or genome edits, both reach higher counts, and §17 argues that is the whole point.

The rates are not close. Integrated circuits doubled every 1.7 years for sixty-four years, DNA sequencing every 1.4 years, chemical DNA synthesis every 3.1 years. Genetic circuits doubled every 7 to 9 years, depending on whether you anchor 2000 at the toggle switch or the repressilator, and the 2016-to-2020 interval is no faster at 8.2 years. Twenty years after its first artefact, the integrated circuit had 29,000 components on one die. Twenty years after its first artefact, synthetic biology had fourteen regulators in one cell.

Two features of that comparison do the real work. The gap is not a story about biology being slow: reading and writing DNA, the substrate the whole programme is built on, scaled at Moore-like rates over the same period. And the circuit curve does not merely sit low, it fails to steepen. A technology that is scaling gets faster as its tools improve. This one did not.

Measured in 2009, by one of the field's founders
"Surprisingly, the actual complexity of synthetic biological circuitry over this time period, as measured by the number of regulatory regions, has only increased slightly; it is possible that existing engineering design principles are too simplistic to efficiently create complex biological systems and have so far limited our ability to exploit the full potential of this field."
Purnick & Weiss 2009, Nat Rev Mol Cell Biol 10:410 PDF. Their Figure 2b plots complexity against year for 2000 to 2008 and reports that it "seems to have reached a plateau".

That is the claim, and the diagnosis, stated in a 2009 review from Ron Weiss's own lab, one of the three groups that built the founding circuits. It is not a retrospective judgement. The plateau was visible while it was happening, it was measured, and the explanation offered was that the design principles were too simple.

Within a year it was on Nature's news pages. Kwok (2010) reports that "although the number of published synthetic biological circuits has risen over the past few years, the complexity of those circuits, or the number of regulatory parts they use, has begun to flatten out", citing Purnick & Weiss directly. And the field kept using the same yardstick. Cello's own discussion measures itself against it:

Seven years later, the same metric
"Our largest circuit has 12 regulated promoters, doubling a plateau first noted in 2009 regarding a limit on the complexity of circuits that could be designed by hand."
Nielsen et al. 2016, Science 352:aac7341 PDF. The largest circuit, Consensus, contains ten regulatory proteins and 55 genetic parts. Of 60 automatically designed circuits, 45 worked in every output state, and 92% of 412 output states matched prediction.

Two doublings in eleven years, announced as an achievement, and correctly so. Nobody in the field claims a Moore's law for circuits. This section is not disputing the field's self-assessment. It is repeating it, with the arithmetic attached.

The 2020 result is the current ceiling. Shin, Zhang, Der, Nielsen & Voigt (2020) encoded a binary-coded-digit to seven-segment decoder in E. coli: seven strains, each carrying a circuit with up to 12 repressors and two activators, and 63 regulators and about 76,000 base pairs across all seven together. Per cell, the number is fourteen. The distinction matters, because a chip's transistor count is per die.

So the arc holds. What the 2020 paper adds is the reason, and the reason is what the rest of this essay is about.

The bottleneck, named by the people who moved it
"Synthetic genetic circuits offer the potential to wield computational control over biology, but their complexity is limited by the accuracy of mathematical models. … Central to this software is the quality of the mathematical model describing the gates, which is used to predict how they will perform when connected. This study introduces a gate model that … captures non-additive effects between input promoters, dynamics, and the utilization of cellular resources that can lead to slow growth and evolutionary breakage."
Shin et al. 2020, Mol Syst Biol 16:e9401 PDF

And the model they replaced, stated in their own equation (1), is a Hill function with additive input composition:

yi = (ymax,i − ymin,i) · Kini / (Kini + xini) + ymin,i,xi = Σ upstream RNAP fluxes

They say so explicitly: "The first model for NOR gates simply treated the input as the sum of the RNAP fluxes of the upstream promoters." That is Hill plus additivity, cited to Kelly et al. 2009 and Nielsen et al. 2016. Going from three regulators to fourteen required abandoning additivity and adding a resource-utilisation term, which is RNAP flux accounting. Those are precisely the two assumptions §19 identifies as one violated inequality.

Why the gains do not compound

Both halves of the picture now have the same explanation. Complexity rose exactly where the model was repaired. It rose slowly because each repair was a specific correction to a specific effect, fitted to a specific gate library, and corrections like that do not compound.

This is the difference that a doubling period encodes. VLSI's doubling compounded because the abstraction held: a designer's gains were reusable by the next designer without revisiting the physics. Here, every increase in circuit size has been bought by re-measuring the parts and patching the model for the effect that broke last time. Sixty-three regulators over seven strains is an achievement of characterisation, not of composition.

So the target is not a better Hill function. It is a model class in which composition is exact and shared resources are represented from the start, so that gains are inherited instead of re-earned. That is the argument of §21 and §23, and Figure 7 is the reason to want it.

Part V

The diagnosis

Two decades of surprises, one violated inequality. This part states it exactly, shows where real circuits sit relative to it, and then argues that the fix is a change of model class rather than a change of parameters.

19Retroactivity and burden are one inequality

Both surprises say the same thing: a pool assumed to be in excess was not in excess.

Here is the central technical claim of this essay. Line up the two headline failures of the engineering programme and ask what each one is, mechanically.

Retroactivity. A transcription factor drives a downstream promoter. To do so it binds the operator. Bound molecules are removed from the free pool. So the effective input to the downstream gene, and the effective level of the upstream signal, both depend on how many downstream sites there are. The assumption that broke is the complex is a negligible fraction of the total transcription factor.

Growth burden and resource competition. A circuit's mRNAs compete with the host's mRNAs for ribosomes, and its genes compete for RNA polymerase. Bound machinery is removed from the free pool. So the expression of each gene depends on how much every other gene is demanding. The assumption that broke is the complex is a negligible fraction of the total ribosome or polymerase pool.

These are not two phenomena. They are one violated inequality, applied to two different shared species. And crucially, it is the same inequality that licenses writing a Hill or Michaelis–Menten function in the first place.

20The inequality, stated properly

It is not "substrate greatly exceeds enzyme". The correct condition is a ratio involving the affinity, and the difference is exactly where gene circuits live.

Stated loosely, the condition is Stot ≫ C, so that Stot ≈ S. That is the right condition. It is worth having the classical sharpened form as well, because the naive version ("enzyme is much less abundant than substrate") is what most people carry around and it is not the criterion.

For the standard scheme E + S ⇌ C → E + P with initial totals E0 and S0, the quasi-steady-state approximation, and therefore the Michaelis–Menten rate law, is valid when

ε = E0 / (KM + S0) ≪ 1Segel & Slemrod 1989, eq. 19

Segel & Slemrod (1989, SIAM Review) derive this as the necessary condition, and note that it is strictly stronger than the timescale-separation condition alone. It reduces to E0 ≪ S0 only when the substrate is well above the affinity. When affinity is tight, KM is small and the criterion becomes harder, not easier.

And here is the review that states the modern consequence in exactly these terms:

The diagnosis, in the literature, verbatim
"Despite the ubiquity of the MM rate law, it accurately captures the dynamics of underlying biochemical reactions only so long as it is applied under the right condition, namely, that the substrate is in large excess over the enzyme-substrate complex. Unfortunately, in circumstances where its validity condition is not satisfied, especially so in protein interaction networks, the MM rate law has frequently been misused. … inappropriate use of the MM rate law distorts the dynamics of the system, provides mistaken estimates of parameter values, and makes false predictions of dynamical features such as ultrasensitivity, bistability, and oscillations."
Kim & Tyson 2020, PLoS Comput Biol 16:e1008258 PDF

Read the list of false predictions against the founding papers of §9. Ultrasensitivity, bistability, oscillations. The toggle switch is a bistability claim resting on a cooperativity exponent. The repressilator is an oscillation claim resting on a Hill coefficient. Neither paper is wrong about what it built, because both were verified experimentally. The point is narrower and more damaging: the model class used to reason about them cannot be trusted to get those three features right, which is why the field could not predict its way from three genes to fourteen.

One number from §9 makes this concrete. The repressilator's reported KM is 40 monomers per cell. A handful of plasmid copies each carrying operator sites puts the site count within an order of magnitude of that. So ε = tsites/(KD + tTF) is not small. The founding model of the field was applied, in its own parameters, at the edge of its own validity.

Interactive

Where the approximation fails. One binding reaction, TF + site ⇌ complex, with everything counted in molecules per cell. Compare the exact occupancy against the Hill-style form that ignores titration of the TF.

−2−10 1234 log₁₀ ( ligand total / KD ) −3−2−1 012 log₁₀ ( binding-site total / KD ) ε < 0.1 Hill / Michaelis–Menten safe: the complex is a negligible slice of the ligand pool ε > 0.1 titration matters; the module has no context-free transfer function ε = tE / (KD + tS) = 0.1 textbook enzyme assay nM enzyme, mM substrate, µM KM repressilator repressor vs its operators KD = 40 / cell, sites on a few plasmid copies TF with decoy sites (Brewster 2014) ribosome pool vs circuit mRNA demand the burden regime a well-insulated part what the insulator programme was buying
Figure 8. The validity region, and where real circuits sit. Axes are the two dimensionless groups that decide whether a Hill or Michaelis–Menten form is legitimate. The shaded region is ε = tE/(KD + tS) < 0.1. Placements are order-of-magnitude estimates from the cited parameters, not measurements. The textbook enzyme assay that motivated the rate law in 1913 sits deep inside the safe region. Gene regulation with a handful of high-affinity sites does not. The insulator and load-driver programme of §15 is best understood as an attempt to drag circuits down and to the right, into the region where the abstraction the field had already adopted becomes true.

21Why this is a missing model class, not a missing parameter

A parameter error shrinks when you measure better. This one does not, because the object being measured does not exist independently of its context.

The diagnosis is that the field lacked a foundational theory of the right kind. Not a law like Newton's, but a characterisation of a model class guaranteed to contain every biological system. This is the part of the case that is an argument rather than a fact, so it deserves the sharpest statement and the strongest available objection.

The argument. Mechanical engineering has the Lagrangian. Every mechanical system you might build is in that class, so system identification is a search inside a family you know contains the answer. Circuit theory has Kirchhoff plus constitutive relations, with the same guarantee. Given such a class, theory and experiment can close a loop: the model proposes, the experiment disposes, and the discrepancy is informative because it must be a discrepancy within the class. That loop is what drove the electrical and mechanical revolutions.

Biology's analogue exists in principle, and that is the pivot. By the universality of biochemistry, the class of all chemical reaction networks certainly contains every cell. Everything mechanical, electrical, thermal or osmotic that interacts with life interacts through that core. So a guaranteed class is available. It was simply not used, because in 2000 it was neither writable nor simulable at scale, and Hill and Michaelis–Menten were.

The consequence is that the field's working class was smaller than biology. Not approximate in a controlled way. Smaller. Real systems sat outside it, and the standard symptoms of that condition are exactly what §13 through §16 document: models that fit in isolation and fail on composition, parameter estimates that shift when context changes, and predicted dynamical features that do not survive.

The strongest objection

A resistor is not a law of nature. Linear passive components are a manufacturing achievement, produced by decades of materials work, and the same is true of clean mechanical bearings. So one could argue the correct reading is not "biology lacks a theory" but "biology lacks manufactured components", and that the insulator programme was therefore right and merely early.

This objection is serious and partly correct. The reply is about which asymmetry binds. In electronics, you can manufacture a component whose behaviour is context-free because you control its material environment and can put a physical boundary around it. Inside a cell, the shared pools are the environment, and there is no boundary to draw. Ribosomes are common to every gene by construction. So component-style insulation in a cell is not a materials problem awaiting effort. It is a request to remove the coupling that makes the cell a cell. That is why insulators buy a factor and not an abstraction, and it is why the class has to change.

What the class has to do

Three requirements, in order of importance. It must contain the real system, so that discrepancy is informative. It must be writable from the kind of information genomes and databases actually supply, which is structural rather than parametric. And it must be computable, or nothing in the loop runs. All-of-chemistry satisfies the first and fails the third. Hill functions satisfy the third and fail the first. The whole problem is the middle.

22Four answers on offer

The field did respond to the missing class. It produced four distinct strategies, and they are still competing.

It is worth being explicit that the diagnosis in §21 is not a complaint that nobody noticed. Everyone noticed. Four coherent responses exist, each with a real success record and a real cost.

Table 2. The four live strategies. Each is a defensible answer to "we lack a model class that contains the system."
StrategyThe moveExemplarsWhat it costs
Patch the rate law Keep Hill functions, add a correction term for each effect you discover: resource demand coefficients, promoter interference, RNAP flux. Qian et al. 2017; Shin et al. 2020; McBride & Del Vecchio 2021; the tQSSA remedy of Kim & Tyson 2020 Works, and demonstrably scales circuits by an order of magnitude. But each patch is discovered by being surprised, so the class is still not known to contain the system, and there is no telling how many patches remain.
Find phenomenological laws Give up on mechanism. Find empirical relations that hold across conditions and reason with those instead. Scott et al. 2010; Klumpp, Zhang & Hwa 2009; Hui et al. 2015; You et al. 2013 Genuinely predictive and genuinely robust, and it explicitly does not need the circuits elucidated. But laws are found, not derived, so you cannot ask a new question until someone finds the law for it.
Stop predicting, screen Build a large library, measure everything, keep what works. Replace design with selection. Kosuri et al. 2013 (explicitly); directed evolution generally; automated strain engineering The most successful strategy of the era by commercial output. Produces no transferable understanding, and cannot reach designs that no library will contain.
Change the class Replace Hill and Michaelis–Menten with a class that makes no excess assumption and is still writable and computable. binding and catalysis; reaction-order geometry; §23 below Requires the theory to be built before the payoff arrives, and has to prove it can do something the other three cannot.
no model class that contains the system patch the rate law keeps mechanism, keeps surprises phenomenological laws keeps prediction, drops mechanism screen, do not predict keeps results, drops understanding change the class keeps both, must be built first 3 → 14 regulators. how many patches left? growth laws. cannot pose a new question alone. six products, two Nobels. no transfer. the bet of §23
Figure 9. Four responses to one gap. The first three are all in active use and all work. The argument here is not that they are wrong, but that only the fourth restores the theory-experiment loop that made the other engineering revolutions cumulative.
Part VI

Forward

The proposed class, the enabling event that has not arrived, and what the record does and does not support.

23Binding and catalysis

A cell cannot change pressure, temperature or pH at will. It can only change whether a catalyst is present. That single restriction is what turns all-of-chemistry into a workable class.

This section states the proposal compactly. The full development is the companion essay, The Biomachine Perspective.

Start from the restriction. A chemist gets to change conditions: heat the flask, change the solvent, swing the pH. A cell does not. Its interior is held within narrow bounds, which is most of what homeostasis is for. So a cell cannot make a reaction go by changing conditions. Its only lever is whether the enzyme for that reaction is present and active. Therefore every load-bearing reaction in a cell is an enzymatic reaction, and every enzymatic reaction has the same two-part structure: molecules bind to form a complex, and the complex is catalysed into products.

The class

Split the network into a fast binding part and a slow catalysis part. Then:

dq/dt = S · v,   v = Kcat x,   N log x = log k,   L x = q

S is the catalysis stoichiometry matrix. L is the subunit composition matrix, recording which complex is made of which parts. N carries the binding equilibria. q are the conserved totals and x the individual species.

Note what has and has not been assumed. Nothing anywhere says a complex is a negligible fraction of its parts. The totals and the free species are kept as separate variables, related by L, which is exactly the distinction whose collapse causes retroactivity and burden.

Three properties make this a candidate rather than just a rewriting.

It contains the systems it needs to. Binding networks are universal in the relevant sense: they can approximate any positive continuous function arbitrarily well. So restricting from all chemistry to binding-plus-catalysis loses no expressive power over the regulatory functions a cell can implement.

Its architecture is fixed and knowable. The catalysis matrix S is what genome annotation and metabolic databases already give you. That is the structural information the genome supplies for free, in Bailey's sense. What is unknown is the binding network L, and that is where the learning problem lives.

The unknown part is at equilibrium, which makes it tractable. Binding is fast, so the binding layer is an equilibrium problem, not a stiff dynamical one. This is the same structural gift that made equilibrium statistical-mechanics models the starting point for machine learning, and it means gradients through the binding layer are well behaved. Mechanism identification then reduces to sparsity in the binding affinities: an affinity of zero means that complex does not form.

What this buys, in the terms of Part IV

Retroactivity is not a correction term in this class. It is what L x = q says. Resource competition is not a patch. It is the same equation applied to the ribosome pool. Reaction orders, the log-derivatives ∂ log x / ∂ log q, are the natural regulatory currency, and their achievable values form polyhedra whose vertices are the dominance regimes. That is the content of the companion tutorial, and the practical consequence is a way to ask, before building, whether a proposed circuit's intended behaviour is even reachable. Liu & Xiao (2026) is that question posed directly as parameter-regime validity for biocircuits.

24The missing shock

Each previous wave needed an enabling event. The honest answer to "what is the third one" is that we do not yet know, and there are three candidates with different implications.

One structural point in the history above is worth naming. Theory alone has never started one of these waves. Each time, a discovery arrived that made a new object visible, and the three schools moved because they could suddenly see the same thing. So the question is not only whether the model class is right. It is what the enabling event will be.

Candidate one: omics at scale. The 1970s gave named mechanisms. The 2000s gave a parts list. Large-scale multi-omics and perturbation data give, for the first time, the parametric half of Bailey's split, at scale and across conditions. If the binding network is what has to be learned, this is the data it would be learned from.

Candidate two: learned model classes. This candidate deserves the most care, because Figure 6 already contains its precedent. Protein structure prediction crossed the frontier that gene circuits did not. It did so with a learned model class, trained on a large curated dataset, rather than a derived one. AlphaFold and then generative design did for one biological question exactly what is argued here to be missing for another. That is encouraging for the ambition and awkward for the method, because it suggests the class need not be derived from first principles at all.

Candidate three: a mechanistic model class that is also learnable. This is the bet behind the foundational virtual-cell programme, and it is an attempt to take both lessons at once. Keep the class mechanistic, so that its interior is interpretable and its discrepancies are informative. Make it learnable at scale, by using the equilibrium structure of the binding layer to get well-behaved gradients. The claim being tested is that binding and catalysis is the class where those two demands are compatible.

The bar this has to clear §17 is the standing challenge. Directed evolution and genome editing won the era's two Nobels by making the missing theory unnecessary. Screening beat prediction on commercial output. So a new model class does not get to justify itself by being more correct. It has to do something the other three strategies structurally cannot: extrapolate to a design no library contains, transfer a mechanism from one context to another, or state in advance that a proposed behaviour is unreachable. Those are the three things only a class that contains the system can do, and they are the deliverables to hold the programme to.

25What the record shows, and what it does not

Three findings in this history are solid enough to argue from. The proposal built on them is not yet, and it should be held to a specific deliverable.

The arc has been: three questions, three moments when a discovery let all three be asked of one object, and a third moment that produced less than the first two. Part V's answer is that a single modelling assumption failed and was never replaced. That answer is an argument, so it is worth separating what the record establishes from what the argument adds to it.

Three findings the record establishes

The diagnosis is in the literature, written by the practitioners. §18: Shin et al. 2020 open by saying circuit complexity "is limited by the accuracy of mathematical models", then replace additive Hill composition with non-additive interference plus resource accounting, and gain an order of magnitude. That is not a critic's reading of the field. It is the field's own methods section.

The field's biggest wins came from not needing the theory. §17: any argument for building the missing model class has to be made against a decade in which screening and editing beat prediction, decisively, twice over. That raises the bar, and it also clarifies the target. The class earns its keep only where selection cannot reach.

The central claim is measurable, and the field keeps the score. "Never scaled like VLSI" reads like rhetoric until you use the field's own metric. Purnick & Weiss defined it in 2009, measured a plateau, and named too-simple design principles as the likely cause. Cello reported itself in 2016 as "doubling a plateau first noted in 2009". Seven to nine years per doubling against 1.7 for transistors is not a rhetorical contrast, and §18 is the arithmetic.

Those three are why the history is worth telling as a history rather than as an opinion. Each one is a number or a sentence from a paper by someone who was building at the time, not a retrospective judgement imposed on them.

What to distrust here

Three places where the evidence is thinner than the prose around it, stated so a reader can discount them appropriately.

The calculation that would settle it

What is missing from the case for §23 is not a paper. It is a calculation: a worked demonstration that the binding-and-catalysis class predicts a composition failure a Hill-function model gets wrong, on a circuit someone has already built and measured. Retroactivity and burden are both available as test cases, because both are documented failures with published numbers, and both are supposed to fall out of L x = q rather than being fitted.

Until that exists, Parts V and VI are a diagnosis with a proposed treatment and no trial. The three deliverables named at the end of §24 are the terms on which the treatment should be judged: extrapolate to a design no library contains, transfer a mechanism from one context to another, or state in advance that a proposed behaviour is unreachable. Nothing in the history above obliges anyone to believe those are achievable. It only shows why they are the right things to ask for.

The one-paragraph version

Physics, system and industry each ask a different question of life, and each is answered with whatever the era can build. Three times a discovery let all three ask their question of the same object. The first time gave proofreading, demand theory and the metabolic plant. The second gave chemotaxis end to end, genome-scale metabolism and the engineered circuit. The third did not arrive, and the reason is not that biology turned out to be too complicated. It is that the model class the field adopted in 2000, Hill functions composed additively, assumes a pool in excess that gene circuits do not have. Retroactivity and burden are that one assumption failing twice. Everything else follows.


References

Every one of the 107 entries below was read in full. PDF marks a verified full text. TXT marks one obtained as complete text rather than PDF, from the PMC open-data archive. Ordered by the section that first uses them.

Part I · the frame

  1. Monod, J. (1971) Chance and Necessity. Knopf. PDF
  2. Nicholson, D.J. (2019) Is the cell really a machine? J Theor Biol 477:108. PDF
  3. Hopfield, J.J. (2014) Whatever happened to solid state physics? Experience at the physics–biology interface. Phys Biol / Annu Rev Condens Matter Phys. PDF

Part II · the 1970s

  1. Hopfield, J.J. (1974) Kinetic proofreading: a new mechanism for reducing errors in biosynthetic processes requiring high specificity. PNAS 71:4135. PDF
  2. Ninio, J. (1975) Kinetic amplification of enzyme discrimination. Biochimie 57:587. PDF
  3. Galstyan, V., Husain, K., Xiao, F., Murugan, A. & Phillips, R. (2020) Proofreading through spatial gradients. eLife 9:e60415. PDF
  4. Xiao, F. & Galstyan, V. (2024) Kinetic proofreading with the leisure of time. PNAS. PDF
  5. Savageau, M.A. (1974) Genetic regulatory mechanisms and the ecological niche of E. coli. PNAS 71:2453. PDF
  6. Savageau, M.A. (1977) Design of molecular control mechanisms and the demand for gene expression. PNAS 74:5647. PDF
  7. Savageau, M.A. (1998) Demand theory of gene regulation, I and II. Genetics 149:1665, 1677. I II
  8. Savageau, M.A. (1976) Biochemical Systems Analysis. Addison-Wesley. PDF
  9. Shinar, G., Dekel, E., Tlusty, T. & Alon, U. (2006) Rules for biological regulation based on error minimization. PNAS 103:3999. PDF
  10. Gerland, U. & Hwa, T. (2009) Evolutionary selection between alternative modes of gene regulation. PNAS / J Mol Evol. PDF
  11. Sasson, V. et al. (2012) Mode of regulation and the insulation of bacterial gene expression. Mol Syst Biol / Mol Cell. PDF
  12. Kacser, H. & Burns, J.A. (1973) The control of flux. Symp Soc Exp Biol 27:65. PDF
  13. Heinrich, R. & Rapoport, T.A. (1974) A linear steady-state treatment of enzymatic chains. Eur J Biochem 42:89. PDF
  14. Varma, A. & Palsson, B.O. (1994) Metabolic flux balancing: basic concepts, scientific and practical use. Bio/Technology 12:994. PDF
  15. Stephanopoulos, G., Aristidou, A. & Nielsen, J. (1998) Metabolic Engineering: Principles and Methodologies. Academic Press. PDF
  16. Bailey, J.E. (1991) Toward a science of metabolic engineering. Science 252:1668. PDF
  17. Bailey, J.E. (2001) Complex biology with no parameters. Nat Biotechnol 19:503. PDF

Part III · the genome

  1. Berg, H.C. & Brown, D.A. (1972) Chemotaxis in E. coli analysed by three-dimensional tracking. Nature 239:500. PDF
  2. Segall, J.E., Block, S.M. & Berg, H.C. (1986) Temporal comparisons in bacterial chemotaxis. PNAS 83:8987. PDF
  3. Barkai, N. & Leibler, S. (1997) Robustness in simple biochemical networks. Nature 387:913. PDF
  4. Alon, U., Surette, M.G., Barkai, N. & Leibler, S. (1999) Robustness in bacterial chemotaxis. Nature 397:168. PDF
  5. Yi, T.-M., Huang, Y., Simon, M.I. & Doyle, J. (2000) Robust perfect adaptation in bacterial chemotaxis through integral feedback control. PNAS 97:4649. PDF
  6. Elowitz, M.B. & Leibler, S. (2000) A synthetic oscillatory network of transcriptional regulators. Nature 403:335. PDF
  7. Gardner, T.S., Cantor, C.R. & Collins, J.J. (2000) Construction of a genetic toggle switch in E. coli. Nature 403:339. PDF
  8. Endy, D. (2005) Foundations for engineering biology. Nature 438:449. PDF
  9. Kelly, J.R. et al. (2009) Measuring the activity of BioBrick promoters using an in vivo reference standard. J Biol Eng 3:4. PDF
  10. Brophy, J.A.N. & Voigt, C.A. (2014) Principles of genetic circuit design. Nat Methods 11:508. PDF
  11. Andrianantoandro, E., Basu, S., Karig, D.K. & Weiss, R. (2006) Synthetic biology: new engineering rules for an emerging discipline. Mol Syst Biol 2:2006.0028. PDF
  12. Kitano, H. (2002) Systems biology: a brief overview. Science 295:1662. PDF
  13. Purnick, P.E.M. & Weiss, R. (2009) The second wave of synthetic biology: from modules to systems. Nat Rev Mol Cell Biol 10:410. PDF The measured plateau, §18.
  14. Alon, U. (2019) An Introduction to Systems Biology, 2nd ed. CRC Press. PDF
  15. Edwards, J.S. & Palsson, B.O. (2000) The E. coli MG1655 in silico metabolic genotype. PNAS 97:5528. PDF
  16. Orth, J.D., Thiele, I. & Palsson, B.O. (2010) What is flux balance analysis? Nat Biotechnol 28:245. PDF
  17. Schaechter, M., Maaløe, O. & Kjeldgaard, N.O. (1958) Dependency on medium and temperature of cell size and chemical composition during balanced growth of Salmonella typhimurium. J Gen Microbiol 19:592. PDF
  18. Kjeldgaard, N.O., Maaløe, O. & Schaechter, M. (1958) The transition between different physiological states during balanced growth. J Gen Microbiol 19:607. PDF
  19. Scott, M., Gunderson, C.W., Mateescu, E.M., Zhang, Z. & Hwa, T. (2010) Interdependence of cell growth and gene expression. Science 330:1099. PDF SI
  20. Klumpp, S., Zhang, Z. & Hwa, T. (2009) Growth rate-dependent global effects on gene expression in bacteria. Cell 139:1366. TXT
  21. Hui, S. et al. (2015) Quantitative proteomic analysis reveals a simple strategy of global resource allocation in bacteria. Mol Syst Biol 11:784. PDF
  22. You, C. et al. (2013) Coordination of bacterial proteome with metabolism by cyclic AMP signalling. Nature 500:301. PDF

Part IV · the cooling

  1. Kwok, R. (2010) Five hard truths for synthetic biology. Nature 463:288. PDF The five truths are tabulated in §12.
  2. Cameron, D.E., Bashor, C.J. & Collins, J.J. (2014) A brief history of synthetic biology. Nat Rev Microbiol 12:381. PDF
  3. Way, J.C., Collins, J.J., Keasling, J.D. & Silver, P.A. (2014) Integrating biological redesign: where synthetic biology came from and where it needs to go. Cell 157:151. PDF
  4. Bashor, C.J. & Collins, J.J. (2018) Understanding biological regulation through synthetic biology. Annu Rev Biophys 47:399. PDF
  5. Elowitz, M. & Lim, W.A. (2010) Build life to understand it. Nature 468:889. TXT
  6. Khalil, A.S. & Collins, J.J. (2010) Synthetic biology: applications come of age. Nat Rev Genet 11:367. TXT
  7. Kosuri, S. et al. (2013) Composability of regulatory sequences controlling transcription and translation in E. coli. PNAS 110:14024. PDF
  8. Mutalik, V.K. et al. (2013) Precise and reliable gene expression via standard transcription and translation initiation elements. Nat Methods 10:354. PDF
  9. Mutalik, V.K. et al. (2013) Quantitative estimation of activity and quality for collections of functional genetic elements. Nat Methods 10:347. PDF
  10. Beal, J. et al. (2016) Reproducibility of fluorescent expression from engineered biological constructs in E. coli. PLoS ONE 11:e0150182. PDF
  11. Cardinale, S. & Arkin, A.P. (2012) Contextualizing context for synthetic biology. Biotechnol J 7:856. PDF
  12. Ceroni, F., Algar, R., Stan, G.-B. & Ellis, T. (2015) Quantifying cellular capacity identifies gene expression designs with reduced burden. Nat Methods 12:415. PDF
  13. Ceroni, F. et al. (2018) Burden-driven feedback control of gene expression. Nat Methods 15:387. PDF
  14. Gyorgy, A. et al. (2015) Isocost lines describe the cellular economy of genetic circuits. Biophys J 109:639. PDF
  15. Qian, Y., Huang, H.-H., Jiménez, J.I. & Del Vecchio, D. (2017) Resource competition shapes the response of genetic circuits. ACS Synth Biol 6:1263. PDF
  16. Tan, C., Marguet, P. & You, L. (2009) Emergent bistability by a growth-modulating positive feedback circuit. Nat Chem Biol 5:842. TXT
  17. Nikolados, E.-M., Weiße, A.Y., Ceroni, F. & Oyarzún, D.A. (2019) Growth defects and loss-of-function in synthetic gene circuits. ACS Synth Biol 8:1231. PDF
  18. Sleight, S.C. & Sauro, H.M. (2013) Visualization of evolutionary stability dynamics and competitive fitness of E. coli engineered with randomized multigene circuits. ACS Synth Biol 2:519. PDF
  19. Borkowski, O. et al. (2016) Overloaded and stressed: whole-cell considerations for bacterial synthetic biology. Curr Opin Microbiol 33:123. PDF
  20. Del Vecchio, D., Ninfa, A.J. & Sontag, E.D. (2008) Modular cell biology: retroactivity and insulation. Mol Syst Biol 4:161. PDF
  21. Jayanthi, S., Nilgiriwala, K.S. & Del Vecchio, D. (2013) Retroactivity controls the temporal dynamics of gene transcription. ACS Synth Biol 2:431. PDF
  22. Mishra, D. et al. (2014) A load driver device for engineering modularity in biological networks. Nat Biotechnol 32:1268. PDF
  23. Brewster, R.C. et al. (2014) The transcription factor titration effect dictates level of gene expression. Cell 156:1312. PDF
  24. Lee, T.-H. & Maheshri, N. (2012) A regulatory role for repeated decoy transcription factor binding sites. Mol Syst Biol 8:576. PDF
  25. Segall-Shapiro, T.H. et al. (2018) Engineered promoters enable constant gene expression at any copy number in bacteria. Nat Biotechnol 36:352. PDF
  26. Karr, J.R. et al. (2012) A whole-cell computational model predicts phenotype from genotype. Cell 150:389. PDF SI
  27. Macklin, D.N. et al. (2020) Simultaneous cross-evaluation of heterogeneous E. coli datasets via mechanistic simulation. Science 369:eaav3751. PDF
  28. Ahn-Horst, T.A. et al. (2022) An expanded whole-cell model of E. coli links cellular physiology with mechanisms of growth rate control. npj Syst Biol Appl 8:30. PDF
  29. Lewis, N.E., Nagarajan, H. & Palsson, B.O. (2012) Constraining the metabolic genotype–phenotype relationship. Nat Rev Microbiol 10:291. TXT
  30. O'Brien, E.J., Monk, J.M. & Palsson, B.O. (2015) Using genome-scale models to predict biological capabilities. Cell 161:971. PDF
  31. Seif, Y. & Palsson, B.Ø. (2021) Path to improving the life cycle and quality of genome-scale models of metabolism. Cell Syst 12:842. PDF
  32. Voigt, C.A. (2020) Synthetic biology 2020–2030: six commercially-available products that are changing our world. Nat Commun 11:6379. PDF
  33. Roybal, K.T. et al. (2016) Precision tumor recognition by T cells with combinatorial antigen-sensing circuits. Cell 164:770. TXT
  34. Din, M.O. et al. (2016) Synchronized cycles of bacterial lysis for in vivo delivery. Nature 536:81. TXT
  35. Paddon, C.J. et al. (2013) High-level semi-synthetic production of the potent antimalarial artemisinin. Nature 496:528. PDF
  36. Ro, D.-K. et al. (2006) Production of the antimalarial drug precursor artemisinic acid in engineered yeast. Nature 440:940. PDF
  37. Peplow, M. (2016) Synthetic biology's first malaria drug meets market resistance. Nature 530:389. PDF DOI is 10.1038/530390a, not 530389a.
  38. Riglar, D.T. & Silver, P.A. (2018) Engineering bacteria for diagnostic and therapeutic applications. Nat Rev Microbiol 16:214. PDF
  39. Nielsen, A.A.K. et al. (2016) Genetic circuit design automation. Science 352:aac7341. PDF
  40. Jones, T.S. et al. (2022) Genetic circuit design automation with Cello 2.0. Nat Protoc 17:1097. PDF
  41. Shin, J., Zhang, S., Der, B.S., Nielsen, A.A.K. & Voigt, C.A. (2020) Programming Escherichia coli to function as a digital display. Mol Syst Biol 16:e9401. PDF
  42. Gorochowski, T.E. et al. (2017) Genetic circuit characterization and debugging using RNA-seq. Mol Syst Biol 13:952. PDF
  43. Hutchison, C.A. et al. (2016) Design and synthesis of a minimal bacterial genome. Science 351:aad6253. PDF
  44. Thornburg, Z.R. et al. (2022) Fundamental behaviors emerge from simulations of a living minimal cell. Cell 185:345. PDF

Part V and VI · diagnosis and proposal

  1. Michaelis, L. & Menten, M.L. (1913), translated in Johnson, K.A. & Goody, R.S. (2011) The original Michaelis constant. Biochemistry 50:8264. PDF
  2. Briggs, G.E. & Haldane, J.B.S. (1925) A note on the kinetics of enzyme action. Biochem J 19:338. PDF
  3. Hill, A.V. (1910) The possible effects of the aggregation of the molecules of haemoglobin on its dissociation curves. J Physiol 40:iv. PDF
  4. Segel, L.A. (1988) On the validity of the steady state assumption of enzyme kinetics. Bull Math Biol 50:579. PDF
  5. Segel, L.A. & Slemrod, M. (1989) The quasi-steady-state assumption: a case study in perturbation. SIAM Review 31:446. PDF
  6. Borghans, J.A.M., de Boer, R.J. & Segel, L.A. (1996) Extending the quasi-steady state approximation by changing variables. Bull Math Biol 58:43. PDF
  7. Kim, J.K. & Tyson, J.J. (2020) Misuse of the Michaelis–Menten rate law for protein interaction networks and its remedy. PLoS Comput Biol 16:e1008258. PDF
  8. Gunawardena, J. (2014) Time-scale separation: Michaelis and Menten's old idea, still bearing fruit. FEBS J 281:473. PDF
  9. Cornish-Bowden, A. (2015) One hundred years of Michaelis–Menten kinetics. Perspect Sci 4:3. PDF
  10. Hill, C.M., Waight, R.D. & Bardsley, W.G. (1977) Does any enzyme follow the Michaelis–Menten equation? Mol Cell Biochem 15:173. PDF
  11. Bintu, L. et al. (2005) Transcriptional regulation by the numbers: models. Curr Opin Genet Dev 15:116. PDF
  12. Buchler, N.E. & Louis, M. (2008) Molecular titration and ultrasensitivity in regulatory networks. J Mol Biol 384:1106. PDF
  13. Gutenkunst, R.N. et al. (2007) Universally sloppy parameter sensitivities in systems biology models. PLoS Comput Biol 3:e189. PDF
  14. Transtrum, M.K. et al. (2015) Perspective: sloppiness and emergent theories in physics, biology, and beyond. J Chem Phys 143:010901. PDF
  15. McBride, C.D. & Del Vecchio, D. (2021) Predicting composition of genetic circuits with resource competition. bioRxiv. PDF
  16. Xiao, F. & Doyle, J.C. (2018) Robust perfect adaptation in biomolecular reaction networks. IEEE CDC. PDF
  17. Xiao, F. et al. (2021) Structure and stability of biomolecular circuits. J R Soc Interface. PDF
  18. Xiao, F., Li, Y. & Doyle, J.C. (2023) Flux exponent control in metabolism. PDF
  19. Liu, Y. & Xiao, F. (2026) Evaluating valid parameter regimes for biocircuits. bioRxiv. PDF
  20. Jumper, J. et al. (2021) Highly accurate protein structure prediction with AlphaFold. Nature 596:583. PDF
  21. Watson, J.L. et al. (2023) De novo design of protein structure and function with RFdiffusion. Nature 620:1089. PDF

Companion essay: The Biomachine Perspective develops the reaction-order and dominance theory referenced in §23.