We hand a patient a Randot book, they call out a few circles, we write "40" arcsec in the chart, and we move on. That number gets treated like a fact about the patient — a fixed property of their visual system, the way axial length or corneal curvature is a fixed property. The problem is that stereoacuity isn't a fixed measurement. Instead, it's a single sample pulled from a distribution, and the distribution is the part that we're not writing down.

Start with what the cortex is actually doing when it computes disparity. It is not comparing two retinal images pixel by pixel and outputting a clean number. Disparity-tuned neurons in V1, V2, and MT respond to a given disparity with real trial-to-trial variability — present the same stimulus twice and you don't get the same firing pattern twice. For a long time this variability was treated as noise to be averaged away. Ma, Beck, Latham, and Pouget's 2006 paper in Nature Neuroscience made the more useful argument: that variability isn't just noise the system tolerates, it's the substrate the system computes with. A population of neurons with Poisson-like variability automatically represents a probability distribution over the stimulus, not a point estimate. The width of that distribution — how much the population's response varies — carries information about certainty. A narrow, high-gain response means the system is confident. A broad, noisy one means it isn't, even if the two situations produce the same "best guess."

Depth perception, under this framework, isn't a readout. It's an inference — a posterior (using Bayesian vocabulary) built from noisy, ambiguous input, updated by whatever else the visual system has to work with: vergence state, accommodative effort, prior expectations about the scene, other depth cues competing or agreeing with disparity. Two visual systems can arrive at the same central estimate — the same "40 arcsec" — while carrying very different amounts of confidence around that estimate. Clinically, the number written in the chart is the mean; what isn't getting measured is the variance.

This is where the clinical tools show their seams. O'Connor and Tidbury's 2018 review in Clinical and Experimental Optometry lays out plainly what most of us already sense but rarely say out loud: current stereoacuity tests measure static disparity at a fixed number of coarse levels, and they can misclassify a patient as stereoblind who has real, usable binocular potential that the test simply isn't sensitive enough to catch. Add to that the test-retest data. Adler, Scally, and Barrett tested 139 UK schoolchildren twice on Randot and found stereoacuity improved by an average of a full plate on repeat, with some children shifting by up to three test levels between visits. A single score isn't a stable trait measurement. It's one draw from a noisy process, and the process itself is what should be discussed and studied.

Two frameworks already circulating within our own field describe the same phenomenon. Paul Harris's binocular continuum makes the point that having two eyes doesn't mean a patient is continuously fusing with them — binocularity is a state the system can occupy when needed, not a switch that's simply on versus off. A person can move through the spectrum of binocularity levels depending on their visual skills or what the task is asking of them. Eric Hussey's language of intermittent central suppression pushes on the same idea from the suppression side: the visual system doesn't suppress in a fixed, all-or-nothing way, it suppresses dynamically depending on target, attention, and demand. Both are describing a system that fluctuates around a probabilistic center rather than sitting at a fixed operating point — exactly what you'd predict if disparity processing is a distribution being sampled rather than a switch being flipped.

Sue Barry is the case that should have ended the debate about whether a low or absent score on a booklet test means a patient has no usable stereopsis. Barry — a neurobiologist, strabismic from infancy who underwent three strabismus surgeries as a child and who did vision therapy as an adult in her 40s, who came to be known as "Stereo Sue" after Oliver Sacks profiled her in the New Yorker — did not walk into developmental optometry offices with a clean 20-arcsec Randot score. What people in our field have described watching her do after undergoing vision therapy is localize the depth and float of a Vectogram target off a wall with a level of precision that had no business coexisting with her formal test results. That's not a contradiction the field should be comfortable filing away. It's the same distribution problem again, just at the far end: a patient whose real-world disparity processing was demonstrably functional, in certain conditions, in ways the standardized booklet was never built to catch, because the booklet samples one point on a distribution and calls it the whole answer.

And the distribution itself isn't fixed. Ding and Levi's 2011 PNAS paper trained adults who had been stereoblind or stereoanomalous their entire lives — failing the clinical stereo test outright, thresholds above 400 arcsec — on a repetitive stereoscopic discrimination task. After training, several of them tested at 40 to 140 arcsec on the same clinical test that had previously called them stereoblind. Nothing changed anatomically in that window. What changed was what the system had learned to do with the same noisy input — its top-down search and attentional weighting toward the disparity signal improved, and stereopsis that had been statistically buried became recoverable. Once the system knows what it's looking for, it can search for it more efficiently, and the same input that used to fail to reach conscious threshold starts to pop out. That's not measurement noise being cleaned up. That's the posterior shifting.

The clearest clinical demonstration of this sits in divergence excess exotropia. Stathacopoulos and colleagues, in their 1993 Ophthalmology paper, compared near and distance stereoacuity in 44 patients with intermittent exotropia against 50 normal controls. They found that distance stereoacuity was significantly and specifically degraded relative to controls, while near stereoacuity in many of these patients looked entirely unremarkable. You can have a patient walk out of your exam with a perfect near stereo score and functionally no usable stereopsis past arm's reach. That's the same probabilistic system sampled under two different demands — and if you only ever sample it at near, you'll never know the distance number was failing. Declaring a patient "has stereopsis" off one test, at one distance, isn't a finding. It's an incomplete sample mistaken for the whole picture.

That's the real lesson sitting inside the divergence excess data: a patient's near score and their distance score aren't two facts about two separate systems. They're two draws from the same probabilistic system under two different demands, and the gap between them is diagnostic information most of us throw away the moment the near number looks fine. Taken seriously, that gap — near versus distance, this visit versus the last one, booklet versus real-world report — is the data worth understanding, not noise around a single "true" score.

Here's my opinion: optometry treats stereoacuity the way it treats axial length — a number taken once at intake, logged, and never revisited, because it feels like a fixed structural fact rather than something that moves. It isn't. Ding and Levi's stereoblind adults didn't have a different eye by the end of training; they had a system that had learned to search the same noisy input more efficiently. That makes the number in the chart a treatment target, not just a diagnostic snapshot — something we can shift the odds on the same way we already shift the odds on convergence insufficiency or accommodative dysfunction, if we test the conditions under which it holds up and build treatments around widening the range where it doesn't.

The reason most of us aren't doing that is simpler: most of us aren't testing what each eye is actually contributing to the binocular scene in the first place. Polarized refraction has essentially disappeared from routine practice. Red/green refraction survives mainly as a comparison of chromatic sharpness through a bisected target, not as the randomized-letter procedure that actually isolates each eye's contribution — and almost no one is using it as an anti-suppressive coefficient inside the refractive equation itself.

If stereopsis is a probability and not a measurement, then checking it once at intake and never revisiting it is the wrong model of care — the equivalent of taking a single blood pressure reading in a healthy-looking patient and closing the chart for good. The right model treats that first number as a baseline, tests it under enough different conditions to find where it collapses, and works to narrow and stabilize the distribution over the course of treatment. The patients most worth worrying about aren't the ones with a bad number. They're the ones with a good number and no idea how fragile it is.