Showing posts with label computer vision depth perception stereo vision. Show all posts
Showing posts with label computer vision depth perception stereo vision. Show all posts

Thursday, March 29, 2007

Computer Vision (10)

Let’s come back to depth, which is our main topic. To perceive depth we know that both our eyes need to capture some common region. Even in this common region you cannot see two different objects at once even though you have two eyes. Try it out right now! Take up a long word and try to see the first and the last characters at the same time. You will not be able to do it because whenever you look somewhere, the same object will be placed on the macula of both the eyes. Now isn’t that redundant? No, that’s exactly what is responsible for the perception of depth. But, how does one eye know where the other is seeing? What if we have many similar objects placed around us, will our brain be fooled? This is exactly what the computer vision scientists are trying to crack from several decades. The concept is called stereo correspondence. In order to mock what our eyes are doing we use two cameras, place them at an offset similar to how our eyes are placed and take an image from both of them. When you look at such a photograph (I have a sample below) a lot of objects would have appeared in both the images, which are redundant for 2D perception but required for 3D viewing. So these are the objects we are interested in, and need to match them in both the images to get the relative depth.

In case of our eye since the entire surrounding cannot be captured on the macula, we have to move them relative to each other to see different objects. Our retina is not a uniform sensor. In order to see something clearly we have to place it on the macula and hence the need for this movement. On the other side, a camera sensor is uniform in density and hence the entire surrounding can be analyzed just with a single shot from both the cameras, no movement required.

Your brain can actually perceive depth from these two 2D images, if viewed properly. You will need some practice for that. Here’s how, if you are interested in it. To appreciate how our brain creates 3D out of these two 2D images and why we are so keen in copying from it, it’s better you learn and then only proceed.
In my next post I will explain about triangulation which is the central idea behind the calculation of depth using two 2D images.

Thursday, March 22, 2007

Computer Vision (8)

The reader, by this time would have understood the problem at hand and also how we are looking forward to solve it. If not, just keep a few things in mind. With just one eye, it is not possible to perceive depth. Without depth, it is not possible to segment the objects around us so effectively. Without segmentation it is not possible for us to learn or update our knowledge. If you have got the essence of it, some of the questions that would definitely pop in your mind are:
  1. Is solving this problem so difficult?
  2. Why would we want to solve it the way our brain does, isn’t there a better way?
  3. When a camera auto focus system can estimate the depth using IR, why can’t we use say LASER to get the exact depth?

To explain why we would want to solve it in the same way as our brain does, I would like to quote these lines taken from the introduction section of one of the related papers from MIT. It states,
“The challenge of interacting with humans constrains how our robots appear physically, how they move, how they perceive the world, and how their behaviors are organized. We want it to interact with us as easily as any another human. We want it to do things that are assigned to it with a minimum of our interaction. In other words we can never predict how it is going to react to a stimulus and what decision it is going to take.

For robots and humans to interact meaningfully, it is important that they understand each other enough to be able to shape each other’s behavior. This has several implications. One of the most basic is that robots and humans should have at least some overlapping perceptual abilities. Otherwise, they can have little idea of what the other is sensing and responding to. Vision is one important sensory modality for human interaction, and the one in focus here. We have to endow our robots with visual perception that is human-like in its physical implementation. Similarity of perception requires more than similarity of sensors. Not all sensed stimuli are equally behaviorally relevant. It is important that both human and robot find the same types of stimuli salient in similar conditions. Our robots have a set of perceptual biases based on the human pre-attentive visual system. Computational steps are applied much more selectively, so that behaviorally relevant parts of the visual field can be processed in greater detail.”

I think that this completely justifies the claim made above. For us what is important is how useful it will be for us humans. Take for example, the compression algorithms used in audio and image processing. Audio compression is based on our ability to perceive or reject certain frequencies and intensities. It is compressed such that there won’t be any perceptual difference between the original and compressed data for our system. For a dog it might really play out weird! Image compression also works on the same basic concepts.

As you go on reading my posts you will get to know whether the problem is difficult or not (that is the main reason why I started writing). I can’t answer this question in one or two lines here. Coming to the third question, LASER will always give you an exact depth or distance of an object, but our brain doesn’t work on exactness. Even though your brain perceives depth it doesn’t really measure it. Secondly getting intelligence out of a LASER based system is a tough one. If you use a single ray to measure the depth of your surrounding, what if your LASER is always pointing on an object moving in unison with the LASER? We need a kind of parallel processing mechanism here, like the one that we get from an image. The entire surrounding is captured at one shot and analyzed, which a LASER fails to do. You cannot use multiple LASERs, because in that case, how would you distinguish the received signals from the ones that left out. The ray that leaves the transmitter a particular point need not comeback to the same point (due to deflections). In that case what will be resolution of the transmitters and receivers or how densely should we pack them? What if there was something we wanted to perceive in between this left out space? This is neither the best way to design a recognition system nor a competitor to our brain, so let’s just throw it away.

Assuming that evolution has designed the best system for us, which has been tried and tested continuously for millions of years, we don’t want to think of something else. We have a working model in front of us, so why not replicate it? And this is not something new for us; we have designed planes based on birds, boats based on marine animals, robots based on us and other creatures, etc, etc.

Monday, March 19, 2007

Computer Vision (6)

In my earlier post I was saying that even though we have two eyes we cannot use them independently. If our eyes cannot move independent of each other, what is it that is holding them? For both the eyes to see the same object either our brain has to be doing some kind of correlation between the images and providing feedback to the eyes asking them to position on a common point or the eyes themselves know where they have to be pointing. I mean, either it is a process of learning or it comes along when our system (life) is booted.
This is actually a debatable topic. I tried to find an answer to this by observing it in small babies, but haven’t been successful enough to conclude. Anyway I have some other observations to share. Depth perception does not produce an interrupt in the brain like the way sound, motion or color do. During the initial learning stages it is interrupt that matters because you need to draw the attention of a baby’s brain to observe something, so depth takes a back seat. I term it is an interrupt because it immediately brings your brain into action. In order to achieve this you generally tend to get some colorful toys that make interesting sounds and wafture in front of a baby. So how does it work?
Sound, as you know definitely produces interrupt in your brain, which is why you use an alarm to wake up in the morning. Colorful objects produce high contrast images in your brain which are like step and impulse functions; strong signals that your brain becomes interested in. Now you know what kind of dress to wear to draw the attention of everyone around you!
If you remember awakening a day dreamer by wavering you hand in front of him, you know how motion produces interrupt in your brain. This is actually because of the way our visual processor and retina are designed, which I will come to shortly. So next time you are buying a toy think about these.Secondly why interrupt matters is because the new born baby’s brain is like a formatted hard disk, ready to accept data, but has nothing. When it doesn’t understand anything around it, there is absolutely no meaning in perceiving depth. Whether it perceives or not, it is just going to be a colored patch and nothing else. Again it wouldn’t know which color it is! So interrupts help it to make sense of its surrounding, and when that is done depth and motion help it to segment the objects from one another to form its database.

Monday, March 12, 2007

Computer Vision(1)

I have been doing research in the field of vision since the last few years but a universal solution is nowhere to be seen; not only in my pockets but in any of the research institutes as well. Sometimes I feel it is impossible to find one universal solution to solving depth through stereo. Researchers have been thinking very deeply about just the stereo correspondence problem from many decades, but in vain. Either the problem is very difficult or we have missed the right track at the very beginning. I believe in either of these. The problem might be very difficult because it has taken its present shape through evolution, over millions of years. When I say evolution I refer to Darwin’s theory of “The survival of the fittest”. All these complex biological systems have taken rigorous stress test from nature and have survived to date, and to crack it might be a very challenging task. On the other side of the court we have man who has been able to build high speed miniature components and complex systems which should easily be able to replicate these relatively low speed systems. This is because biological vision is seen in small organisms like the insects, which hardly have say a few thousands of neurons dedicated for visual processing. Aren’t our GHz processors able to achieve what has evolved in these small creatures? After so much of research and thought at least I believe that it is impossible to achieve depth perception through the currently tried out image processing techniques. Towards the end of my discussion I will try to take you on a walk along what path I believe can solve this problem. It might not be practical as of today, but who knows what technology is waiting for us at the door. If the current image representation and processing techniques are so poor at achieving recognition, why are all those researchers glued to it even now? That’s because we always take our brain as reference for developing any recognition system and our brain is still able to perceive depth given a 2D stereo image pair. I believe the reader understands what a stereo image pair is. If not Google it right now! Or just skip it for now; I am going to take you all on a long journey covering each and every topic related to stereo images, depth perception, etc, etc. A lot of questions and answers were brought about during my research period and I have tried to give my best possible solutions to all of them. The whole idea behind writing these articles is to share these Q&A’s so that a person fascinated about vision today will be able to start off from a much thoughtful point rather than repeating the experiments again and again. To race against nature we have to make sure we compress those millions of years of evolution into a much smaller duration.