Note: I used GPT-5.6 Sol to help consolidate my notes and emails into an initial draft of this blog. The underlying work and analysis are my own, and I revised the resulting writing extensively. Some of my original, more organized, AI-free notes are available here.
TLDR: Large-deviation theory can be used to turn a rare-fluctuation problem into an equation of state (EOS) problem. In an interacting 1D chain, a low-order Padé approximation to that equation of state reproduces every limit used to construct it, but becomes nonmonotone at stronger coupling. The resulting rate function is nonconvex, which is thermodynamically forbidden. A Maxwell construction can restore convexity, but in this case it would (probably?) repair an artifact of the approximation rather than highlighting 'interesting physics.'
This note grew out of conversations over a couple of months that I had with Professor Richard Stratt, the Newport Rogers Professor of Chemistry at Brown University. The subject of these conversations started with simplified models of polymer folding, and eventually progressed to using large-deviation theory to infer rare orientational fluctuations. Understanding the motivation for using large-deviation theory is straightforward: important physical events, particularly in glassy liquids, are typically controlled by coordinated rearrangements of molecules. These happen very rarely, and, as such, are hard to observe in simulation or with low-order statistics. Large-deviation theory provides natural tools to study these events using pretty accessible information, eliminating the need for prohibitively long brute-force simulations. The interpolation strategy discussed here is from Mainas and Stratt, J. Chem. Phys. 162, 024501 (2025). I worked through this 1D test case almost entirely, though I consulted Professor Stratt frequently. He also derived these results independently, albeit slightly differently.
The paper’s central idea is neat, and in my opinion, very hard to dislike: combine information that is easy to obtain near equilibrium with analytic information in extreme-field limits, then interpolate between them to reconstruct the full fluctuation distribution.
The method works well in the liquid-crystal systems studied in the paper. The point of this note is to stress-test the generality of the approach by probing simple cases; in the simplest interacting model Professor Stratt and I discussed, the same kind of asymptotic matching can generate a globally impossible EOS even though every imposed local constraint is correct.
the basic statmech picture
A statmech system has an enormous number of microscopic configurations. We usually do not care about which exact configuration the system occupies, but rather about a coarse variable such as the magnetization, density, or, in this note, the extension of a chain. Let’s denote the intensive chain extension by .
Some values of can be produced by many microscopic configurations, while others require the system to organize itself in a very specific way. So the probability distribution contains thermodynamic information. Near its maximum are the ordinary fluctuations seen all the time in a simulation. Far into its tails are rare, highly organized fluctuations. To study these organized fluctuations, it is useful to rewrite the probability as an effective free energy:
For a system made of identical “pieces”, the free energy cost of holding the whole system at an atypical value usually grows approximately in proportion to . Large-deviation theory expresses this as
The rough intuition is that an atypical global value has to be maintained across the whole system: piece 1 must do something weird, piece 2 must also do something weird, etc. For weakly correlated blocks, those probabilities roughly multiply, so the total probability becomes exponentially small in , while the corresponding free energy becomes proportional to .11.The symbol denotes asymptotic equivalence on the exponential scale: . Equivalently, for , as . The argument in the main text is only intuition.
The function is called the rate function. It is essentially the free energy cost per particle, measured in units of . Its minimum is the equilibrium value of . Its curvature near that minimum controls ordinary Gaussian fluctuations, while its shape farther away controls rare events.22.If minimizes , then nearby. Substituting this into gives a Gaussian with variance approximately , which is the usual central-limit scaling of fluctuations about the mean. For the 1D chain examined below, means no net extension and means every link points in the same direction. A large value of means that extension is exponentially unlikely in the chain length.
the Gärtner–Ellis theorem
The difficult object is , because measuring it directly requires sampling the rare tails of . The Gärtner–Ellis theorem gives a complementary description in terms of a field that biases the system toward the value of we want to study.
Let be the corresponding extensive quantity, and define a dimensionless field . Weighting each configuration by favors configurations with large when and small when .33.This reweighting is equivalent to adding the term to the Hamiltonian, since . At finite , we define the scaled cumulant-generating function , where denotes an average in the unbiased ensemble. Its thermodynamic limit form is .
Physically, is minus times the change in free energy per particle caused by the field. Differentiating it gives the average value selected by that field, i.e. . This identity is standard in statmech and easy to see before taking the large- limit. Differentiating the logarithm pulls down a factor of , so
The subscript means an average in the biased ensemble. A second derivative gives , so must be nondecreasing. In particular, increasing the field cannot decrease the equilibrium response. Convexity thus follows directly from the fact that variance is nonnegative.44.I am skipping much of the algebraic manipulation here—it is worthwhile to go through it to convince yourself that these identities do hold. The key point is that the numerator and denominator together are just an expectation value in the field-biased ensemble.
Now for the actual theorem! In its simplest form, the Gärtner–Ellis theorem says that if the large- limit is sufficiently “regular”—most importantly, differentiable over the relevant range—then the rate function is its Legendre–Fenchel transform:
At a smooth optimum, differentiating the expression inside the brackets with respect to gives , and if this relationship can be inverted, the maximizing field is and
There is a relatively simple physical way to understand this. Without a field, realizing costs . Applying rewards that extension by , so the biased system minimizes . The selected value therefore satisfies .
The theorem itself is not the approximation studied here. For the exact 1D model, it gives the correct convex rate function. The approximation enters later, when we do not know the full function , so we try to reconstruct it from limited information.
the EOS shortcut
Now, near , the response of is roughly linear, i.e. , where is the zero field susceptibility. For a symmetric system (as ours will be and many simplified folding models are), . In practice, the weak field slope can be estimated from simulation; in the exactly solvable model below, we will just derive it by expanding around the origin.
At the opposite extreme, a very large field controls the leading approach to saturation, making that limit analytically pretty simple. The interactions can still enter via subleading corrections. In bounded problems, reaching perfect order generally requires a divergent field, but the leading asymptotic behavior is typically known. This is great news for us!
With this in mind, the strategy used by Mainas and Stratt can be summarized as follows:
- Obtain the weak field susceptibility from normal, unbiased simulations.
- Derive the strong field saturation behavior analytically.
- Construct a minimal interpolation between the found limits.
- Integrate that EOS to estimate the rate function.
If it works, information from the center of a distribution predicts its tails without directly sampling them or running a dense set of field-biased simulations. This would be awesome—rare events happen... rarely. The question in this post is whether matching the two endpoint regimes is generally sufficient to provide a valid function between them.
the exact 1D test case
Consider a chain of links, each pointing left or right:
The scaled extension is the total chain displacement per link. Neighboring links interact through the dimensionless coupling , and a dimensionless pulling field couples to the total extension. Positive makes neighboring links prefer the same direction, so the chain develops correlated domains. Positive favors right-pointing links. At zero field the two directions remain symmetric and the mean extension is zero, but increasing makes the chain much easier to polarize; once a small field tips one link, its neighbors prefer to align.
This is the nearest-neighbor 1D Ising model in a field.55.Calling this a polymer model is slightly silly, but it has the variable we care about: a bounded extension produced by locally interacting links. More importantly, it is exactly solvable. As many readers likely know, this model is analytically solvable, and a common solution is via the transfer matrix method. A transfer matrix packages the Boltzmann weights of each neighboring pair into a matrix. Multiplying this matrix along the chain sums over all link configurations, and in the thermodynamic limit its largest eigenvalue determines the free energy. That eigenvalue is
so differentiating the corresponding free energy with respect to the field gives the mean extension per link:
Inverting this relation gives the exact EOS
which we’ll use to test our approximate forms. Its derivative is
Every factor is positive for . Therefore, is strictly increasing and is strictly convex at every finite . Another way to see this is by using the fact that is a variance, since , so must be nonnegative. This condition is important, and we’ll revisit it as such!
constructing the Padé approximation
Now, we use the large-deviation approach of Mainas and Stratt. The field diverges logarithmically as , whereas its derivative has a simple pole. It is therefore convenient to interpolate instead of .66.In the liquid-crystal problem studied by Mainas and Stratt, the field itself has the pole used in their Padé form. Here diverges logarithmically, so differentiating turns the endpoint divergence into a simple pole. This is therefore a closely related adaptation of their interpolation strategy, rather than literally the same ansatz.
The symmetry makes odd and even. The minimal ansatz with the endpoint poles and enough coefficients to match one weak field condition and two strong field conditions is77.A Padé approximant is a rational-function approximation. Here the denominator enforces poles at , the numerator is even so that is even, and the three coefficients are fixed by one weak field and two strong field conditions.88.It would also be reasonable to use two conditions from the weak field expansion and only the leading strong field pole. Mainas and Stratt discuss the analogous alternative in the supplementary material and explain why it is less practical.
where the P subscript indicates the Padé approximant. I will also usually suppress the explicit dependence below.
weak field
Expanding the exact response around gives with . Thus , so , which fixes .
strong field
The expansions here are a little tricky and took me a while to get right.99.I am skipping a lot of algebra here. To match the strong field behavior, define and . Both and approach zero as . Expanding the exact force-extension curve in small gives
Inverting this series to write in terms of gives
Since and , this implies
Expanding the Padé ansatz about the same endpoint gives
Matching the pole and the constant term gives the constraints
Together with , we have 3 equations with 3 unknowns. These conditions give
Integrating from the origin, with , gives
The construction is exact at , has the exact weak field slope for every , and has the exact first two strong field terms. At small positive coupling it is also numerically very accurate. However, in between, the results are a bit disappointing.
the failure
At , the model is noninteracting and the approximation is exact:
The approximation remains extremely accurate at weak positive coupling:
This is encouraging! The approximation has the correct slope at the origin, the correct divergence near saturation, and seems to interpolate smoothly between the two.
By , the two curves are still close, but differences begin to appear in the interior:
At , the Padé approximation develops a large backward-bending region:
This cannot be the exact equilibrium EOS.1010.Thermodynamic stability fixes the sign of the response between conjugate variables. Here the field couples as , so increasing must increase the equilibrium extension: . For the familiar pressure-volume pair, the sign convention is reversed because . Stability then requires : increasing pressure should compress the system. A branch where increasing pressure increases volume is thermodynamically unstable in an ordinary equilibrium system, though related behavior can occur in constrained, metastable, or otherwise unusual systems. The exact remains monotone, while the approximation predicts a region where increasing the extension requires a smaller field. Equivalently, over part of the interval, so the inferred rate function satisfies there. In other words, the approximate rate function is nonconvex.
The breakdown becomes more dramatic at stronger coupling:
The interesting part is that the approximation is still correct in every limit used to build it. It has the exact weak field slope and the exact strong field pole and constant term. The failure is entirely in the intermediate region, which the asymptotic constraints do not control.
This is the main result. Matching the local physics at both ends does not guarantee a globally valid EOS between them.
a brief note on the Maxwell construction
A standard response to a nonconvex approximate free energy is to replace it with its convex envelope. In the EOS picture, this is the Maxwell construction: the backward-bending branch is replaced by a constant-field plateau chosen using an equal-area condition.
Equivalently,1111.I don't think it's obvious that this is actually equivalent, it's pretty cool. taking the Legendre–Fenchel supremum globally rather than following every local solution of discards the backward-bending branch and returns this convex envelope. That would make the approximate rate function convex again. However, I do not think it resolves the underlying issue here. The exact finite- 1D Ising model has no phase transition, so the Maxwell construction would (I think) just be repairing the Padé approximation rather than recovering the true model behavior.
I did not probe much further into whether there is some interesting physical interpretation of the nonconvex branch. My strong suspicion is that there is not. The exact result remains well behaved, while the pathology appears only after imposing a low-order rational interpolation.
reference
E. Mainas and R. M. Stratt, “Exceptionally large fluctuations in orientational order: The lessons of large-deviation theory for liquid crystalline systems”, J. Chem. Phys. 162, 024501 (2025)
Thanks to Professor Richard Stratt for helpful discussions and guidance with this work. Thanks to Corin Wagen, Moe Zhang, and Raphael Stone for reading drafts of this. Last updated 07/21/2026.