Unmodeled


An explanatory statement on
prediction, experiment, and the
science not yet in the model

August 2026

What does it mean when a calculation and a measurement disagree? Which disagreements deserve a scientist's attention, and which are housekeeping? If the useful ones are rare, can they be found on purpose instead of by accident?

This statement gives brief and introductory answers to those questions.

What a model contains

A model is a record of what someone chose to represent. It holds the variables they thought mattered, at the scales they thought mattered, under the conditions they had in mind. Everything else about the material is absent from it, and the model has no way to report what is absent. This is true of a fitted equation, a density functional calculation, a machine-learned interatomic potential, and whatever replaces them.

That absence is the subject here.

A model that predicts well is not thereby complete. It can be right for a reason it does not contain, across the range where someone happened to check it. The check tells you the model agrees with the world in that range. It does not tell you the represented mechanism is the operating one.

Five reasons a calculation and a measurement disagree

Most disagreement is dull. It is useful to distinguish among five explanations, in the order a working scientist should exhaust them.

The first is the instrument. It drifted, or it was calibrated against the wrong standard, or the reading was taken before the sample reached equilibrium.

The second is the specimen. What went into the furnace is not always what came out of it. Two samples with the same nominal composition can differ in ways nobody wrote down, and processing history is the variable most often lost.

The third is the range. Every model was built and tested somewhere. Asking it about conditions outside that region is asking for an extrapolation it was never qualified to make, and it will answer anyway.

The fourth is the comparison. A computed migration barrier and a measured conductivity are different quantities. Several assumptions sit between them, and a disagreement can belong to any of those assumptions rather than to the physics at either end. This one is the easiest to miss, because both numbers are real and both are correct for what they describe.

The fifth is that the material is doing something the model does not contain.

The first four account for nearly all of it. A scientist who reaches for the fifth before exhausting the others will spend a career explaining artifacts.

What ordinary already looks like

Before calling a gap interesting, it helps to know what a boring gap looks like. In my own work I use a tolerance of one to two percent for density functional theory, depending on the functional, and two to three percent for classical molecular dynamics with a good force field. Generalized-gradient functionals systematically overestimate lattice constants by about one percent. That is expected. It is not an error, and it is not a discovery.

A five percent error in a lattice constant is a different matter, because it cascades. Every downstream prediction inherits it, and the failure shows up somewhere far from its cause.

So the threshold is not "the numbers differ." The numbers always differ. The threshold is that the difference is larger than the known behavior of the method, and that it does not go away when the ordinary explanations are removed one at a time.

What earns the word mechanism

A disagreement that survives the audits is still only a lead. Turning it into a mechanism takes evidence, in a definite order.

It has to repeat. Measured again, on a fresh sample, by someone else if possible, the difference is still there.

It has to have a shape. The disagreement varies systematically with something under deliberate control, such as composition, temperature, atmosphere, grain size, or processing history. Scatter is noise. Structure is a signal that something is responding to a variable.

It has to survive the audits honestly performed. That means someone tried to find the mistake and failed, rather than someone deciding early that the mistake was not there.

And then it has to predict. A proposed mechanism must imply a second effect nobody has measured yet, in a direction that would embarrass the proposal if it came out wrong. Then that measurement gets made.

Only the last step earns the word. Everything before it is an anomaly, which is a respectable thing to have and a poor thing to publish.

The reason for the order is that a flexible correction can fit any residual. A model with enough free parameters will absorb a discrepancy and report an improved error, and nothing has been learned. Prospective prediction is what separates an explanation from a curve that has been made to pass through the points.

A case from my own work

Working on the shear response of a layered material, I expected water at the interface to weaken it, monotonically, more water giving lower strength. The intuition is ordinary lubrication and it is what a simple picture predicts.

The answer was non-monotonic. A small amount of water lubricates the interface and drops the shear strength by an order of magnitude, from 103 MPa dry to 8 MPa at 25 percent hydration. More water then rebuilds strength, through hydrogen-bond networks that bridge the layers.

The simple picture was not wrong about lubrication. It was missing a mechanism that only appears at higher water content, and that mechanism reverses the trend. Had I averaged across hydration states, or sampled only the dry and wet ends, the reversal would have been invisible and the average would have been meaningless.

That is the shape of the thing. A model gave a smooth expectation, the measurement disagreed in a structured way, and the structure was the location of the missing physics.

What I got wrong

I underestimated validation. For a long time I treated agreement as the finish line, which meant that when a calculation matched, I stopped looking. Looks right is not is right, and a model that matches in the range you checked will tell you nothing about the range you did not.

The correction was tedious rather than clever. Predictions registered before the measurement. Scored blind. Failures written down with the condition attached, so that a failure is a bounded statement about where the model stops rather than a number in a table.

I would also say plainly that I have not yet run this loop at the scale the rest of this statement describes. The individual pieces are ordinary practice. Assembling them into something that searches for missing mechanisms deliberately, rather than noticing them by luck, is what I am working on now.

Order of work

The steps below are ordered by what they cost and what they can settle. Nothing in the earlier steps requires new apparatus.

Repeat the measurement first. Recalibrate the instrument. Remake the sample. Check that the computed quantity and the measured quantity are actually the same quantity.

Then vary one controlled thing and watch the shape of the disagreement. A residual with a shape is worth more than a residual with a magnitude.

Later, design the experiment that separates two candidate explanations, rather than the experiment that confirms the preferred one.

Later still, predict something nobody has measured, and then measure it.

After that, put the mechanism back into the model and start again. The updated model has a new boundary, and the boundary is where the next one is.

Doing this on purpose

Most of the loop above is currently manual, and that is the constraint. A researcher notices an odd residual, remembers a similar case, designs a follow-up, waits, and interprets. The rate is set by human attention, and human attention is the scarcest input in a laboratory.

The parts that can be delegated are the parts that are bookkeeping: tracking which comparisons have been made, which explanations have been eliminated, which assumptions each model carries, and which experiment would most efficiently distinguish two candidates. Automated synthesis and characterization change the rate at which the loop can turn. Models trained across many materials change the breadth of the first guess.

None of that changes the standard of evidence. A system that runs a thousand comparisons and flags the largest residuals will mostly find broken instruments and bad samples, at speed. The value is in the discrimination, not the throughput.

The useful version searches for the experiment that settles a question, rather than the experiment that produces the largest disagreement.

Two observations

The first is that this does not displace ordinary science. Most of the work remains building models that are right, and a program built only on failure would have nothing to fail against. A model has to be good before its errors are worth reading.

The second is that this method is the easiest to abuse. Any sufficiently flexible correction will fit a residual, and calling that a discovery is how a field embarrasses itself. The order of evidence above is not bureaucracy. It is the only thing standing between a research program and a collection of fitted artifacts.

I am working on this now, at small scale. The claim in this statement is narrow: a disagreement that survives every ordinary explanation is evidence of a mechanism the model does not contain, and that is worth looking for deliberately.