Tactical Vision
By Jessica Lake
For months I have been trying to simplify vision.
Every simplification removed something I had assumed was essential.
Eventually I realized I had been asking the wrong question.
The question is not:
“How do we build an internal picture of the world?”
The question is:
“What is the minimum computation required to navigate it?”
That realization divided vision into two very different problems.
The first I call strategic vision.
It is the rich, seamless experience we consciously enjoy:
Color.
Depth.
Texture.
Shadows.
Surfaces.
Motion.
Beauty.
Whether that presentation exists primarily for communication, planning, conscious thought, or something else does not concern me here.
The second I call tactical vision.
Tactical vision has no need for pictures.
It needs only enough information to recognize what matters and act upon it.
That single decision simplified everything.
Then Constraint Convergence simplified it again.
VISION IS A REALM
I no longer think tactical vision requires its own cognitive architecture.
Vision is an application of a more general one.
Constraint Convergence.
Visual recognition operates within a geometric realm.
A realm is a space in which information can be expressed in compatible terms.
Geometry has its own lingo.
Position.
Distance.
Orientation.
Adjacency.
Containment.
Connectivity.
Proportion.
Topology.
Shape.
Continuity.
Perhaps many others.
Within that realm is an arsenal of fundamental principles.
A fundamental principle is a statistical invariant expressed in the lingo of the realm.
Something that tends to remain true.
Something sufficiently stable to be useful.
Tactical vision does not need to understand any of this.
It merely operates upon it.
OBJECTS
An object is not fundamentally a picture.
Nor is it a stored canonical image.
It is an accepted identity.
During learning, Constraint Convergence eliminates geometric principles that cannot belong until what remains specifies the identity sufficiently for the need.
That accepted convergence can be retained.
I call the retained result a Nexus.
CLOCK can be a Nexus.
CHAIR can be a Nexus.
DOG can be a Nexus.
Once retained, the system does not have to rediscover that identity from scratch every time reality presents another instance.
The expensive work has already been done.
Recognition is retrieval.
SETS AND MULTISETS
This is where my earlier object model survives.
I had noticed that object recognition could often proceed from collections of parts.
A clock may contain:
A circular form.
Numbers.
An hour hand.
A minute hand.
Perhaps a second hand.
Originally I treated these as components in a recursive object hierarchy.
That was useful, but too procedural.
The important property was simpler.
The collection has no inherent order.
It is a set or multiset.
That matters.
A multiset can say that some relationship or component occurs several times without inventing a sequence in which those occurrences must be processed.
Recognition therefore need not walk through an object description.
Many geometric declarations can participate concurrently.
If the multiset alone sufficiently identifies the object, wonderful.
If not, additional geometric constraints can be introduced.
Relative position.
Containment.
Orientation.
Connectivity.
Proportion.
Whatever distinction becomes critical.
Use only as much information as necessary.
Sufficiency is enough.
DECLARATIVE VISION
This is the most important change.
Tactical vision is declarative.
A declaration says what is so.
It does not say what must happen next.
Suppose Molly is before me.
Many declarations may be simultaneously true.
Molly is Living.
Molly is Animal.
Molly is Dog.
Molly is Labradoodle.
Molly is Molly.
There is no requirement to traverse these identities in sequence.
They are her Heritage.
All are true at once.
Criticality determines which resolution matters.
If I shout:
“Watch out for the dog!”
DOG is sufficient.
If someone asks:
“What breed is she?”
LABRADOODLE becomes critical.
If I am looking for Molly specifically, MOLLY becomes critical.
The identity has not changed.
The need has.
This matters for vision because recognition does not have to resolve everything to maximum specificity.
It resolves only as far as the current need requires.
CONCURRENCY
Declarative representation naturally permits concurrency.
There is no prescribed path:
First test this.
Then test that.
Then inspect this component.
Then traverse to that category.
Many geometric relationships can participate at once.
Many retained identities can be tested at once.
Many constraints can operate at once.
This gives Constraint Convergence enormous computational leverage.
Recognition does not need to serially inspect every possible object.
It allows incompatible possibilities to disappear concurrently.
What remains sufficiently stable becomes the accepted answer.
This is not procedural search.
It is convergence.
THE NOTCH FILTER
My earlier theories still carried one unnecessary piece of machinery.
Prediction.
I knew from the beginning that tactical vision needed the difference between reality and what was already understood.
The delta.
My first solutions therefore imagined some version of:
Construct expectation.
Represent reality.
Compare the two.
Report the difference.
Conceptually, two screens.
One showing what I expected.
One showing what reality supplied.
Then an automated comparator looked for disagreement.
There was nothing impossible about this.
It was merely unnecessary.
A retained nexus already embodies what has been accepted about an identity.
So do not express that knowledge as a simulation.
Make the nexus receptive.
Reality enters.
The nexus behaves something like a notch filter.
What conforms to the retained identity is suppressed.
What does not conform survives.
Reality → Nexus → Delta.No generated picture.
No canonical simulation.
No second screen.
No explicit comparison between two complete representations.
The constitution of the Nexus performs the comparison.
What I already understand disappears into the notch.
What remains is news.
RECOGNITION
Consider a clock.
Reality supplies geometric information.
Circularity.
Relative positions.
Repeated symbols.
Hands.
Centering.
Orientation.
Other relationships.
Those observations encounter retained geometric Nexuses.
Most candidate identities fail some of the active constraints.
They disappear.
Perhaps CLOCK survives.
If CLOCK is sufficient for the current need:
Done.
Bing!
Nothing had to generate what a clock should look like.
Nothing had to rotate a canonical clock.
Nothing had to construct all possible appearances of clocks under every distance, orientation, lighting condition, and partial occlusion.
The retained identity already contains what matters sufficiently for recognition.
Reality meets it directly.
TRANSFORMATIONS
My earlier model gave enormous importance to transformations.
Lighting.
Perspective.
Distance.
Scale.
Orientation.
Color casts.
Occlusion.
I treated these as things that had to be estimated and removed before an observation could be normalized into canonical form.
I no longer think that must be the central operation.
The important distinction is simpler:
Does the transformation alter identity?
A chair remains a chair when viewed from another angle.
A clock remains a clock under different illumination.
A dog remains a dog when partially hidden by a table.
Therefore the defining geometric principles of the identity must tolerate the variation that reality normally supplies.
Transformation does not necessarily have to be explicitly inverted.
It may simply fail to violate the nexus.
That is much cheaper.
If orientation is irrelevant to the identity, orientation need not be normalized away.
If scale is irrelevant within useful bounds, scale need not be corrected first.
If partial occlusion leaves sufficient constraints intact, recognition can still converge.
Only when some transformation makes the identity ambiguous does additional information become critical.
OCCLUSION
Occlusion becomes particularly simple.
Objects continue when they are not seen.
That is a fundamental principle within the relevant realm because reality supports it overwhelmingly well.
A tree passing in front of a house does not ordinarily cause the house to cease to exist.
A table hiding the lower half of a person does not imply that the person terminates at the tabletop.
Continuity survives.
So tactical vision does not have to invent every possible completion of a partially hidden object.
The retained identity plus the principle of continuity already constrains what remains plausible.
Again:
Do not reconstruct what is unnecessary.
Use what survives.
THE ROOM
Now expand from an object to a room.
A room is not fundamentally different from a clock.
It is simply an identity at another scale.
Its geometry contains relationships among objects.
Walls.
Floor.
Ceiling.
Doors.
Furniture.
Windows.
Whatever matters.
Sparse observations can be sufficient to recognize ROOM.
Once ROOM is retrieved as a Nexus, I gain access to associated knowledge.
Typical contents.
Spatial relationships.
Affordances.
Maps.
Expected relationships among objects.
I may therefore know considerably more than I have directly sampled.
That does not require constructing a visual replica of the entire room.
Recognition gives me an address into semantic knowledge.
THE MAP
Maps remain useful.
But they no longer need to generate predictions.
A map is retained relational knowledge associated with a nexus.
Suppose two familiar objects occur in a familiar spatial relationship.
That relationship may be sufficient to retrieve the room or wall with which they are associated.
Once that nexus is active, its map supplies other known relationships.
Reality can then encounter those relationships directly.
What conforms disappears.
What fails survives.
Perhaps something moved.
Perhaps something is missing.
Perhaps something new appeared.
Perhaps the retained map is no longer sufficient.
Delta.
The map does not need to paint an expected scene.
It only needs to participate in the receptive constraint structure.
FIXATION
This changes my interpretation of fixation.
A fixation is not fundamentally a step in a procedural recognition algorithm.
It is simply another sample of reality.
The useful question becomes:
Where would another sample provide information critical to the unresolved identity?
If the current constraints already identify CLOCK sufficiently, another fixation is unnecessary.
If CLOCK and COMPASS both remain admissible, the unresolved distinction determines where additional information is useful.
The next fixation therefore serves criticality.
It supplies information where the current convergence is insufficient.
That is different from constructing rival predictions and deliberately sampling where those generated pictures disagree.
No picture is required.
The current unresolved constraints themselves identify what information is missing.
STATISTICAL FIXATIONS
Some observations have greater eliminative power than others.
That remains true.
A fixation that resolves a broad environmental constraint may affect many identities at once.
Where is vertical?
What is the dominant geometry?
Is this an indoor or outdoor environment?
What large-scale spatial relationships are present?
Such information can be extraordinarily useful because it constrains many simultaneous possibilities.
But there is no mandatory first fixation.
There is no prescribed sequence.
The architecture remains declarative.
If some observation is already available, use it.
If another distinction becomes critical, sample for it.
The statistical advantage determines usefulness, not procedural order.
THE ENVIRONMENT
I once thought the first few fixations should estimate global transformations before object recognition began.
That now seems too strong.
Environment and object identity can constrain one another concurrently.
Recognizing a stove contributes evidence for KITCHEN.
Recognizing KITCHEN makes certain other identities more plausible.
Recognizing a tree contributes evidence for OUTDOORS.
Recognizing OUTDOORS alters which geometric and functional relationships are likely to matter.
No direction has to come first.
This is precisely what declarative representation buys us.
Constraints can flow both ways because there is no prescribed execution order.
The system converges.
WHERE IS WALDO?
Consider “Where’s Waldo?”
The goal is:
Find Waldo.
WALDO already exists as a retained Nexus.
Associated declarations identify him sufficiently.
Glasses.
Hat.
Striped shirt.
Face.
Other characteristic relationships.
The visual field contains enormous amounts of irrelevant information.
Most of it fails the constraints associated with WALDO.
It disappears.
Potential matches remain.
If several remain admissible, additional information becomes critical.
Sample.
Constrain.
Eliminate.
Eventually one identity is sufficiently resolved.
Waldo.
Done.
No special Waldo algorithm.
No full reconstruction.
No exhaustive visual search in the strong procedural sense.
The same Constraint Convergence framework simply operates within the geometric realm under a very specific criticality.
APPARENT INTELLIGENCE
This is the part that fascinates me.
Where exactly is the intelligence?
Reality supplies observations.
A realm supplies compatible lingo.
Fundamental principles supply stable relationships.
Retained Nexuses supply accepted identities.
Maps supply associated relationships.
Criticality determines required resolution.
Constraints eliminate what cannot remain.
Declarations participate concurrently.
A discrepancy survives when reality no longer fits what is retained.
And something useful happens.
At no point do I need to insert a tiny Jessica who understands the whole problem.
Every component can remain comparatively stupid.
2
The apparent intelligence emerges from their interaction.
Perhaps recognition is not intelligence in the sense we normally imagine.
Perhaps recognition is what happens when enough independent constraints operate concurrently upon reality until a sufficiently useful identity remains.
WHAT REMAINS
I began Tactical Vision believing I needed a special visual architecture.
I no longer think I do.
The useful pieces survived.
Sparse sampling.
Geometric relationships.
Sets and multisets.
Maps.
Continuity.
Criticality.
Sufficiency.
But the mechanism underneath them is no longer peculiar to vision.
There is a geometric realm.
It has an algebra.
Within it are fundamental principles.
Learning uses constraints to establish sufficiently useful identities.
Those identities are retained as Nexuses.
Reality encounters those Nexuses receptively.
Conformity disappears.
Discrepancy survives.
Criticality determines how much resolution is required.
Declarative representation permits concurrency.
Constraint Convergence does the rest.
There is no internal picture.
There is no requirement for a generated prediction.
There is no canonical simulation.
There is no exhaustive procedural search.
There is only reality meeting retained knowledge within a geometric realm.
What fits is recognized.
What does not becomes information.
Tactical Vision, it turns out, is not really a theory of vision.
It is Constraint Convergence wearing geometric clothes.