Professional Documents
Culture Documents
Sensation and Perception
Sensation and Perception
2
Seeing: What Is It?
3
Chapter 1
4
Seeing: What Is It?
5 --- ..
Right eye
Optic tract
Superior colliculus
(left lobe)
Optic radiations
surrounds the question of what cells in the lateral two-way up-down traffic going on in this visual
geniculate nuclei do. They receive inputs not only pathway and we discuss its possible functions in
from the eyes but also from other sense organs, so later chapters.
some think that they might be involved in filter- Before we go on to discuss the way fiber ter-
ing messages from the eyes according to what is minations are laid out in the striate cortex, note
happening in other senses. The lateral geniculate that the optic nerves provide visual information to
nuclei also receive a lot of fibers sending messages two other structures shown in 1.6-the left and
from the cortex. Hence there is an intriguing right halves of the superior colliculus. This is a
5
Chapter 1
brain structure which lies underneath the cortex, First, note that a face is shown mapped out on
1.5, so it is said to be sub-cortical. Its function is the cortical surface (cortical means "of the cortex").
different from that performed by regions of the This is the face that the eyes are inspecting.
cortex devoted to vision. The weight of evidence Second, the representation is upside-down .
at present suggests that the superior colliculus The retinal images (not shown in 1.6) are also
is concerned with gu iding visual attention. For upside-down due to the way the optics of the eyes
example, if an object suddenly appears in the field work, 1.1. Notice that the sketch of the "inner
of view, mechanisms within the superior colliculus screen" in 1.1 showed a right-way-up image, so it
detect its presence, work out its location, and then is different in that respect from the mapping found
guide eye movements so that the novel object can in the striate cortex.
be observed directly. Third, the mapping is such that the representa-
It is important to realize that other visual path- tion of the scene is cut right down the middle, and
ways exist apart from the two main ones shown in each half of the cortex (technically, each cerebral
1.6. In fact, in monkeys and most probably also hemisphere) deals initially with just one half of the
in man , optic nerve fibers directly feed at least scene.
six different brain sites. This is testimony to the Fourth, and perhaps most oddly, the cut in the
enormously important role of vision for ourselves representation places adjacent regions of the central
and similar species. Indeed, it has been estimated part of the scene farthest apart in the brain!
that rough ly 60% of the brain is involved in vision Fifth, the mapping is spatially distorted in that
in one way or another. a greater area of cortex is devoted to central vision
Returning now to the issue of pictures-in-the- than to peripheral: hence the relatively swollen
brain, the striking thing in 1.6 is the orderly, albeit face and the diminutive arm and hand, 1.7. This
curious, layout of fiber terminations in the striate doesn't mean of course that we actually see people
cortex. in this distorted way-obviously we don't. But it
6
Seeing: What Is It?
reveals that a much larger area in our brain is as- an "inner screen" and to examine in detail serious
signed to foveal vision (i.e., analyzing what we are logical problems with the "inner screen" idea as a
directly looking at) than is devoted to peripheral theory of seeing. We begin this task by considering
vision. This dedication of most cortical tissue to man-made systems for seeing.
foveal vision is why we are much better at seeing
details in the region of the scene we are looking at Machines for Seeing
than we are at seeing details which fall out toward A great deal of research has been done on building
the edges of our field of view. computer vision systems that can do visual tasks.
All in all, the cortical mapping of incoming These take in images of a scene as input, analyze
visual fibers is curious but orderly. That is, adjacent the visual information in these images, and then
regions of cortex deal with adjacent regions of the use that information for some purpose or other,
scene (with the exception of the mid-line cut). The such as guiding a robot hand or stating what ob-
technical term for this so rt of mapping is topo- jects the scene contains and where they are. In our
graphic. In this instance it is called retinotopic as terminology, a machine of this type is deriving a
the mapping preserves the neighborhood relation- scene description from input images.
ships that exist between cells in the retina (except Whether one should call such a device a
for the split down the middle). The general orderli- "perceiver," a "seeing machine," or more humbly
ness of the mapping is reminiscent of the "inner an "image processor" or "pattern recognizer," is a
screen" proposed in 1.1. But the oddities of the moot point which may hinge on whether con-
mapping should give any "inner screen" theorist sciousness can ever be associated with non-biologi-
pause for thought. The first "screen," if such it is, cal brains. In any event, scientists who work on the
we meet in the brain is a very strange one indeed. problem of devising automatic image-processing
The striate cortex is not the only region of machines would call the activity appearing on the
cortex to be concerned with vision-far from it. "inner screen" of 1.1 a kind of gray level descrip-
Fibers travel from the striate cortex to adjacent tion of the painting. This is because the "i nner
regions, called the pre-striate cortex because they screen" is a representation signaling the various
lie just in front of the striate region. These fibers shades of gray all over the picture, 1.8. (We ignore
preserve the orderliness of the mapping found in color in the present discussion, and also many
the striate region. There are in fact topographically intricacies in the perception of gray: see Ch 16).
organized visual regions in the pre-striate zone and Each individual brain cell in the screen is describ-
we describe these maps in Ch 10. For the present, ing the gray level at one particular point of the
we just note that each one seems to be special- picture in terms of an activity code. The code is
ized for a particular kind of visual analysis, such simple: the lighter or more brightly illuminated the
as color, motion, etc. One big mystery is how the point in the painting, the more active the cell.
visual world can appear to us as such a well-inte- The familiar desktop image scanner is an ex-
grated whole if its analysis is actually conducted at ample of a human-made device that delivers gray
very many different sites, each one serving a differ- level descriptions. Its optical sensor sweeps over the
ent analytic function . image laid face down on its glass surface, thereby
To summarize this section, brain maps exist measuring gray levels directly rather than from a
which bear some resemblance to the kind of "inner lens-focused image. Their scanning is technically
screen" idea hesitantly advanced by our fictional described as a serial operation as it deals with dif-
"ordinary person" who was pressed to hazard a ferent regions of the image in sequence.
guess at what goes on the brain when we see. Digital cameras measure the point by point
However, the map shown in 1.6-1.7 is not much intensities of images focused on their light sens i-
like th e one envisaged in 1.1, being both distorted, tive surfaces, so in this regard they are similar to
upside-down and cut into two. biological eyes. They are said to operate in parallel
These oddities seriously undermine the pho- because they take their intensity measurements
tographic metaphor for seeing. But it is timely to everywhere over the image at the same instant.
change tack now from looking inside the brain for Hence they can deliver their gray levels quickly.
7
Chapter 1
Input image
The intensity measurements taken by both scan- The answer is simple: as long as there is a
ners and digital cameras are recorded as numbers systematic correspondence between the outside
stored in a digital memory. To call this collection scene and the retinal image, the processes of image
of numbers a "gray level description" is apt because interpretation can rely on this correspondence, and
this is exactly what the numbers are providing, as build up the required scene description according-
in 1.8. ly. Upside-down in the image is simply interpreted
The term "gray level" arises from the black- as right-way-up in the world, and that's all there is
and-white nature of the system, with black being to it.
regarded as a very dark gray (and recorded with a
If an observer is equipped with special spectacles
small number) and white as a very light gray (and
which optically invert the retinal images so that
recorded with a large number) .
they become the "right-way-up," then the world
The numbers are a description in the sense
appears upside-down until the observer learns to
defined earlier: they make explicit the gray levels
cope with the new correspondence between image
in the input image. That is, they make these gray
levels immediately usable (which means there is and scene. This adjustment process can take weeks,
no need for further processing to recover them) by but it is possible. The exact nature of the adjust-
subsequent stages of image processing. ment process is not yet clear: does the upside-down
Retinal images are upside-down, due to the world really begin to "look" right-way-up again, or
optics of the eyes (Ch 2) and many people are is it simply that the observer learns new patterns of
worried by this. "Why doesn't the world therefore adjusted movement to cope with the strange new
appear upside down?", they ask. perceptual world he finds himself in?
8
Seeing: What is it?
9
Chapter 1
In sharp contrast to this, the computer registers a number within a spreadsheet, a few moments
that perform the same job as the hypothetical brain later it might be holding the code for a letter in
cells would not be arranged in the computer in a document being edited using a word processor,
way which physically matches the input image. or whatever. Indeed, the capacity for the arbitrary
That is, the "anatomical" locations of the registers assignment of computer registers, to hold differ-
in the computer memory would not necessarily ent contents that mean different things at different
be arranged as the hypothetical brain cells are, in times according to the particular computatio n
a grid- like topographical form that preserves the being run , is held by some to be the true hallmark
neighbor-to-neighbor spatial relationships of the of symbolic computation.
image points. But this capacity for arbitary and changing as-
Instead, the computer registers might be ar- signment is quite unlike the brain cel ls supporting
ranged in a variety of different ways, depending on vision , which, as far as we presently understand
many different factors, some of them to do with things, are more or less permanently committed to
how the memory was manufactured, others stem- serve a particuLar visual function (but see the caveat
ming from the way the computer was programmed below on learning). That is, if a brain cell is used
to store information. The computer keeps track of to represent a scene property, such as the orienta-
each pixel measurement in a very precise manner tion of an edge, then that is the job that cell always
by using a system of labels (technically, pointers) does. It isn't quickly reassigned to represent, say, a
for each of its registers, to show which part of the dog, or a sound, etc., under the control of other
image each one encodes. The details of how this brain processes.
is done do not concern us: it is sufficient to note It may be that other brain regions do contain
that the labels ensure that each pixel value can be cells whose functional role changes from moment
retrieved for later processing as and when required. to moment (perhaps cells supporting language?). If
Consequently, it is true to say that the hypotheti- so, they would satisfY the arbitrary ass ignment defi-
cal brain cells of 1.1 and the receptors of 1.2 are nition of a symbolic computational device given
servi ng the same representationaL function as the above. However, some have doubted whether the
computer memory registers of 1.8-recordi ng the brain's wiring really can support the highly accurate
gray level of each pixel-even though the physicaL cell-to-cell connections that this would require. In
nature of the representation in each case differs any event, visual neurons do not appear to have
radically. It differs in both the nature of the pixel this property and we will not use this definition of
code (cell activity versus size of stored number) symbolic computation in this book.
and the anatomical arrangement of the entities that A caveat that needs to be posted here is to do
represent the pixels. with various phenomena in perceptual learning:
The idea that different physical entities can we get better at various visual tasks as we practice
mediate the same information processing tasks is them, and this must reflect changes in vision brain
the fundamental assumption underlying the field cells. Also, plasticity exists in brain cell circuits in
of artificial intelligence, which can be defined ea rly development, Ch 4, and parts of the brain to
as the enterprise of making computers do things do with vision may even be taken over for other
which would be described as intelligent if done by functions following blindness caused by losing the
human s. eyes (or vice versa: the visual brain may encroach
Before leaving this topic we note another major on other brain areas).
difference between the putative brain cells coding But this caveat is abo ut slowly acting forms of
the gray level description and computer memory learning and plasticity. It does not alter the basic
registers. Computers are built with an extremely point being made here. When we say the visual
precise organization of their components. As stated brain is a symboli c processor we are not saying that
above, each memory register has a label and its its brain cells serve as symbols in the way that pro-
contents can be set to represent different things grammable computer components serve symbolic
according to the program being run on the com- functions using different symbolic assignments
puter. One moment the register might be holding from moment to moment.
10
Seeing: What is it?
11
Chapter 1
around us that we are understandably misled into images but it is precisely because it cannot decide
taking its effortless scene descriptions for granted. what is in the scene from whence the images came
Perhaps it is because vision is so effortless for us that we would not call it a "visual perceiver." De-
that is tempting to suppose that the scene we are vising a seeing machine that can receive a light im-
looking at is "immediately given" by a photograph- age of a scene and use it to describe what is in the
ic type of representation. But the truth is the exact scene is much more complicated, a problem which
opposite. Arriving at a scene description as good is as yet unsolved for complex natural scenes.
as that provided by the visual system is an im- The conclusion is inescapable: whatever the
mensely complicated process requiring a great deal correct theory of seeing turns out to be, it must
of interpretation of the often limited information include processes quite different from the sim-
contained in gray level images. This will become ple mirroring of the input image by simple
clear as we proceed through the book. Achieving a point-by-point brain pictures. Mere physical
gray level description of images formed in our eyes resemblance to an input image is not an adequate
is only the first and easiest task. It is served by the basis for the brain's powers of symbolic visual scene
very first stage of the visual pathway, the light-sen- description. This point is sometimes emphasized by
sitive receptors in the eyes, 1.2. saying that whereas the task of the eyes is forming
This point is so important it is worth reiterating. images of a scene, the task of vision is the opposite:
The intuitive appeal of the "inner screen" theory image inversion. This means getting a description
lies in its proposal that the visual system builds of the scene from images.
up a photographic-type of brain picture of the
observed scene, and its suggestion that this brain Images Are Not Static
picture is the basis of our conscious visual experi- For simplicity, our discussion so far has assumed
ence. The main trouble with the theory is that al- that the eye is stationary and that it is viewing
though it proposes a symbolic basis for vision, the a stationary scene. This has been a convenient
symbols it offers correspond to points in the scene. simplification but in fact nothing could be further
Everything else in the scene is left unanalyzed and from the truth for normal viewing. Our eyes are
not represented explicitly. constantly shifted around as we move them within
It is not much good having the visual system their sockets, and as we move our heads and bod-
build a photographic-type copy of the scene if, ies. And very often things in the scene are mov-
when that task is done, the system is no nearer to ing. Hence vision really begins with a stream of
using information in the retinal image to decide time-varying images.
what is present in the scene and to act appropri- Indeed, it is interesting to ask what happens if
ately, e.g., avoiding obstacles or picking up objects. the eyes are presented with an unchanging image.
After all , we started the business of seeing in the This has been studied by projecting an image from
retina with a kind of picture, the activity in the re- a small mirror mounted on a tightly fitting contact
ceptor mosaic. It doesn't take us any further toward lens so that whatever eye movement is made, the
using vision to guide action to propose a brain image remains stationary on the retina. When this
picture more or less mirroring the retinal one. is done, normal vision fades away: the scene disap-
The inner screen theory is thus vulnerable to what pears into something rather like a fog. Most visual
philosophers call an infinite regress: the problems processes just seem to stop working if they are not
with the theory cannot be solved by positing an- fed with moving images.
other picture, and so on ad infinitum. In fact, some visual scientists have claimed that
So the main point being argued here is that the vision is really the study of motion perception; all
inner screen theory shown in 1.1 totally fails to else is secondary. This is a useful slogan (even if an
explain how we can recognize the various objects exaggeration) to bear in mind, particularly as we
and properties of objects in the visual scene. And will generally consider, as a simplifYing strategy for
the ability to achieve such recognition is anything our debate, only single-shot stationary images.
but an immediate consequence of having a pho- Why do we perceive a stable visual world
tographic representation. A television set has pixel despite our eyes being constantly shifted around?
12
Seeing: What Is It?
13
Chapter 1
14
Seeing : What is it?
1.13 A real-life version of the kind of situation depicted in the computer graphic in 1.12
Remarkably, the square in shadow labeled A has a lower intensity than the square labeled B, as shown by the copies at
the side of the board. The visual system makes allowance for the shadow and to see what is "really there." [Black has
due cause to appear distressed . After: 36 Bf5 Rxf5 37 Rxc8+ Kh7 38 Rh1 , Black resigned. Following the forced exchange
of queens that comes next, White wins easily with his passed pawn. Topalov vs. Adams, San Luis 2005.] Photograph by
Len Hetherington .
strange though it might seem at first sight. Rather, The outcome in 1.12 is not some quirk of com-
it is an example of the visual system making due al- puter graphics, as illustrated in 1.13 in which the
lowance for the shading to report on the true state same thing happens from a photograph of a shaded
of affairs (veridical). scene.
15
Chapter 1
a b
Other brightness "illusions" are shown in 1.14. on the nature of the retinal image. Irs task is to use
These also illustrate the slippery nature of what retinal images to deliver a representation of what is
is to be understood by a "visual illusion." Figure out there in the world. The idea that vision is about
1.14a could well be a case of making allowance for seeing what is in retinal images of the world rather
shading but 1.14b doesn't fit that kind of interpre- than in the world itself is at the root of the delu-
tation because we do not see these figures as lying sion that seeing is somehow aki n to photography.
in shade. However, 1.14b might be a case of the That said, if illusions are defined as the visual
visual system applying, unconsciously, a strategy system getting it seriously wrong when judged
that copes with shading in natural scenes but when against physically measured scene realities, then
applied to certain sorts of pictures produces an human vision is certainly prone to some illu-
outcome that surprises us because we don't see a sions in this sense. These can arise for ordinary
shaded scene. This is not a case of special pleading scenes, but go unnoticed by the casual observer.
because most visual processes are unconscious, so The teacup illusion shown in LISa is an example.
why not this one? On this argument, the varied The photograph is of a perfectly normal teacup,
perceptions of the identical-in-the-image gray together with a normal saucer and spoon. Try
triangles is "illusory" only if we are expecting the judging which mark on the spoon would be level
visual system to report the intensities in the image. with the rim of the teacup if the spoon was stood
But, when you think about it, that doesn't make upright in the cup.
sense. The visual system isn't interested in reporting Now turn the page and look at 1.1Sb (p . 18).
The illusory difference in the apparent lengths
of the two spoons, one lying horizontally in the
saucer and one standi ng vertically in the cup, is
remarkable. Convince yourself that this percep-
a b c
16
Seeing: What Is It?
Power supply
Switch closed if
photocell falls below a
set level
Photocell's
Photocell measures
light intensity in the
corridor
tual effect is not a trick dependent on some subtle the vertical element bisects the horizontal one.
photography by investigating it in a real-life setting The simplest version of this effect, illustrated in
with a real teacup and spoon. It works just as well 1.16a, is known as the vertical-horizontal illu-
there as in the photograph. Real-world illusions sion. The effect is weaker if the vertical line does
like these are more commonplace than is often not bisect the horizontal as in 1.16b but it is still
realized. Artists and craftsmen know this fact well, present. It is easy to draw many realistic pictures
and learn in their apprenticeships, often the hard containing the basic effect. The brim in 1.16c is as
way by trial and error, that the eye is by no means wide as the hat is tall , but it does not appear that
always to be trusted. Seeing is not always believ- way. The perceptual mechanisms responsible for
ing-or shouldn't be. the vertical-horizontal illusion are not understood,
This teacup illusion nicely illustrates the usual though various theories have been proposed since
general definition of visual illusions-as percep- its first published report in 1851 by A. Fick.
tions which depart from measurements obtained The illusions just considered are instances
with such devices as rulers, protractors, and light of spatial distortions: vertical extents can be
meters (the latter are called veridical measure- stretched, horizontal ones shortened, and so
ments). Specifically, this illusion demonstrates on. They are eloquent testimony to the fact that
that we tend to over-estimate vertical extents in perceptions cannot be thought of as "photographic
comparison with horizontal ones, particularly if copies" of the world, even when it comes to a
17
Chapter 1
18
Seeing: What Is It?
19
Chapter 1
a preted accordingly-bumps become hollows and
. ..
vICe versa on inversIOn.
This is a fine example of how a visual effect can
reveal a design feature of the visual system, that
is, a principle it uses, an assumption it makes,
in interpreting images to recover explicit scene
descriptions. Such principles or assumptions are
technically often called constraints. IdentifYing
the constraints used by human vision is a critically
important goal of visual science and we will have
much to say about them in later chapters.
Figure/Ground Effects
20
Seeing : What Is It?
21
Chapter 1
22
Seeing: What Is It?
track of this global aspect; otherwise, we would another object, such as one's arm, through the gap
never notice that impossible figures are indeed while an observer is viewi ng the model co rrectly
impossible. But the overall impossibili ty is a rather aligned, and so seeing the impossible arrangement.
"cogn itive" effect-a realization in thought rather As the arm passes through the gap, it seems to the
than in experience that the figures do not "make observer that it slices through a solid object!
sense." An important point illustrated by 1.25 is the
If the visual system insisted on the global aspect inherent ambiguity of flat illustrations of 3D
as "making sense" then it could in principle have scenes. The real object drawn in 1.25, left, has rwo
dealt with the figures differently. For example, it limbs at very different depths: but viewing with
could have "broken up" one corner of the impos- one eye from the correct position can make this
sib le triangle and led us to see part of it as com ing real object cast just the same image on the retina
out toward us and part of it as receding. This is as one in which the rwo limbs meet in space at the
illustrated by the object at the left of 1.25. same point.
But the visual system emphatically does not do This inherent am biguity, difficult to com-
this, not from a lin e drawing nor from a physical prehend fully because we are so accustomed to
embodiment of the impossible triangle devised by interpreting the 20 retinal image in just one way,
Gregory. He made a 3D model of 1.25, left. When is revealed clearly in a set famous demonstrations
viewed from just the right position, so that it by Ames, shown in 1.26. The observer peers with
presents the same retinal image as the line-drawing, one eye through a peephole into a dark room and
then our visual apparatus still gets it wrong, and sees a chair, 1.26a. However, when the observer
delivers a scene description which is impossible is shown the room from above it becomes appar-
globally, albeit sensible locally. V iewi ng this "real" ent that the real object in the room is not the chair
impossible triangle has to be one-eyed; otherwise, seen through the peephole. In the example shown
other clues to depth come into play and produce in 1.26b, the object is a distorted chair suspended
the physically correct global perception. (Two-eyed in space by invisible wires, and in 1.26c the room
depth processing is discussed in detail in eh 18.) contains an odd assemblage of luminous lines,
One interesting game that can be played with also suspended in space by wires. The co llection of
the trick model of the impossible triangle is to pass parts is cunningly arranged in each case to produce
23
Chapter 1
••
a What the observer sees when
he looks into the rooms in b •
and c through their respective ••
peepholes.
Peepholes
b Distorted chair whose parts
are held in space by thin invis- for looking •• ••
ible wires. The chair is posi- into each •• ••
tioned in space such that it is
darkened •• ••
seen as a normal, undistorted , room
chair through the peephole,
without distortion.
c Scattered parts of a chair
that still look like a normal chair
through the peephole, due to the
clever way that Ames arranged
the distorted parts in space so
that they cast the same retinal
image as a.
a retinal image which mimics that produced by system (along with a typical computer vision sys-
the chair when viewed from the intended vantage tem) utilizes what is called the general viewpoint
point. In the most dramatic example, the lines are constraint. A normal chair appears as a chair from
not formed inro a single distorted object, but lie in all viewpoints, whereas the special cases used by
space in quite different locations-and the "chair Ames can be seen as chairs from just one special
seat" is white patch painted on the wall. vantage point. The general viewpoint constraint
The point is that the two rooms have things justifies visual processes that would yield stable
within them which result in a chair-like retinal structural interpretations as vantage point changes.
image being cast in the eye. The fact that we see The general viewpoint constraint can be embed-
them as the same-as chairs-is because the visual ded in bottom up processing. It is not necessary
system's design exploits the assumption that is "rea- to invoke top down processing in explaining the
sonable" to interpret retinal information in the way Ames chair demonstrations-that is, knowing the
which normally yields perceptions that would be shape of normal chairs, and using this knowledge
valid from diverse viewpoints. It is "blind" to other to guide the interpretation of the retinal image.
possibilities, but that should not deceive us-those Normal scenes are usually interpreted in one
possibilities do in fact exist. way and one way only, despite the retinal image
Another way of putting this is to say that the information ambiguity just referred to. But it
Ames's chair demonstrations reveal that the visual is possible to catch the visual system arriving at
24
Seeing: What Is It?
Initial perception
25
Chapter 1
within psychology, neuroscience, and machine im- Even so, most people would ptobably be more
age-processing, and many samples of this progress impressed with a world-class chess-playing compu-
will be reviewed in this book. But we are still a ter than they would be with a good image-proces-
long way from being able to build a machine that sor, despite the fact that the former has been real-
can match the human ability to read handwriting, ized whereas the latter remains elusive. It is one of
let alone one capable of analyzing and describing our prime objectives to bring home to you why the
co mplex natural scenes. problem of seeing remains so baffiing. Perhaps by
This is so despite multi-million dollar invest- the end of the book you will have a greater respect
ments in the problem because of the immense for your magnificent visual apparatus.
industrial potential for good processing systems. Meanwhile, we have said enough in this open-
Think of al l the handwritten forms, letters, etc. ing chapter to make abundantly clear that any at-
that still have to be read by humans even though tempt to explain seeing by building representations
their contents are routine and mundane, and all which simp ly mirror the outside world by some
the equally mundane object handling operations in sort of physical equi valence akin to photography is
industry and retailing. bound to be insufficient. We do not see our retinal
Whether we will witness a successful outcome images. We use them, together with prior knowl-
to the quest to build a highly competent visual edge, to build the visual world that is our represen-
robot in the current century is debatable, as is the tation of what is "out there." We can now finally
question of whether a solution would impress the dispatch the "inner screen" theory to its grave and
ordinary person. concentrate henceforth on theories which make
A curious fact that highlights both the difficul- explicit scene representations their objective.
ties inherent in understanding seeing and the way In tackling this task, the underlying theme of
we take it so much for granted is that computers this book will be the need to keep clearly distinct
ca n already be made which are sufficiently "clever" three different levels of analysis of seeing. eh 2
to beat the human world champion at chess. But explains what they are and subsequent chapters
computers cannot yet be programmed to match will illustrate their nature using numerous exam-
the visual capacities even of quite primitive ani- ples. We hope that by the time you have finished
mals. Moves are fed into chess playing computers the book that we will have convi nced you of the
in non-visual ways. A computer vision system has importance of distinguishing between these levels
not yet been made that can "see" the chessboard, when studying seeing, and that you will have a
from differing angles in variable lighting condi- good grasp of many fundamental attributes of hu-
tions for differing kinds of chess pieces-even man and, to a lesser extent animal, vision .
though the computer can be made to play chess
brilliantly.
26
Seeing: What Is It?
27
Chapter 1
days of the field artificial intelligence: deciding Cambridge University Press 335-344.
which edges in images of a jumble of toy blocks Gregory RL (1980) Perceptions as hypotheses.
belong to which block. This culminates in a list of Philosophical Transactions ofthe Royal Society of
hard-won general lessons from the blocks world London: Series B: 290 181-197.
on how to study seeing. We hope the beginner will
Gregory RL (1971) The Intelligent Eye. Wiedenfeld
find here valuable pointers about how research on
and Nicolson.
seeing should and should not be studied.
Ch 22 discusses the vexed topic of seeing and Gregory RL (2007) Eye and Brain: The Psychology
consciousness. This is currently much-debated, ofSeeing. Oxford University Press, Oxford.
but our conclusion is that little can be said with Gregory RL (2009) Seeing Through Illusions: Mak-
any certainty. Visual awareness is as mysterious ing Sense ofthe Senses. Oxford University Press.
today as it always has been. We hope this chapter Marr 0 (1982) Vision: A computational investiga-
reveals just why making progress in understanding tion into the human representation and processing of
consciousness is so hard. visual information. San Francisco, CA: Freeman.
Ch 23 ends the book with a review of our
Shepard RN (1990) Mindsights. WH.Freeman
linking theme, the compurational approach to
and Co.: New York. Comment A lovely book with
seeing. Amongst other topics, it sets this review
many amazing illusions.
within a discussion of the strengths and hazards of
different empirical approaches to studying seeing: Thompson P (1980) Margaret Thatcher: a new il-
computer modelling, neuroscience methods, and lusion. Perception 9 483-484. Comment See Percep-
psychophysics. tion (2009) 38 921-932 for commentaries on this
classic paper.
We hope that by the time you have finished
the book you will have a firm conceptual basis
Collections of Illusions, Well Worth Visiting
for tackling the vast literature on seeing. We have
necessarily presented only a small part of that http://www.michaelbach.de/ot/
literature.
http://www.grand-illusions.com/opticalill usions
We also hope that this book will have convinced
you of the importance of distinguishing between www.telegraph.co.uk!news/
the computational theory, algorithm, and hardware newstopics/howaboutthat/3520448/
levels when studying seeing. Most of all, we hope Optical-Illusions---the-top-20.html
that this book leaves you with a sense of the pro- www.interestingillusions.com/en/subcat/
found achievement of the brain in making seeing
color-illusions/
possible.
http://lite.bu.edu/vision-flash 1O/applets/lite/lite/
Further Reading lite.html
28