Wednesday, November 26, 2008

Computer Vision(40): 3D Physical to 2D Projection of Depths

When our brain is not sure of what is being sensed it tries to use knowledge to come to a conclusion. When there is some audio playing with a lot of foreground noise your brain will find it difficult to recognize it only till it uncovers the language in which the background audio is playing. If the language is the one it knows it decodes it else everything will be noise. When it comes to illusions also you can tune your brain to recognize the image in certain way only and not the others. I guess one of the things that might have happened to your friend is this. The other is in the way he might have seen the image. I have clearly mentioned in the footer of the image that this applies only when the image is viewed as cross stereo. I will rule out the later case as even you tend to see it either way.
Sense is something that is fuzzy in humans; you know something is hot but cannot tell its exact temperature, you know something is farther away from the other but cannot tell the exact distance between the two, you know some sound is louder than the other but not by how much, etc… So long as you are interested in only the relative information our brain performs exactly but if you question its exactness it becomes relative and fuzzy.
Let me explain relative information taking the top view of the scenario as shown below.
Currently the objects are placed one beside the other (images not to scale). The relative depth between them will be zero as the distance between the objects in both the eyes will be the same. To observe relative depth between them we will have to first create physical depth between the two by pushing either one of them back.
CASE 1: Let me push the rectangle back initially as shown below. ‘L’ and ‘R’ are the projected distances between the objects on the left and the right eye respectively. Whenever the object on the right is in front compared to the one on the left, the projected distance on the left eye will be greater than that on the right.

CASE 2: Now consider the image below in which the circle is pushed back. For our brain to perceive this physical relative depth the projected distance on the right eye has to greater than that on the left.
Let me bring into picture the image that actually created this doubt. If you cross view this stereogram the 2D distance between the objects that the left eye would perceive will be greater than that on the right which is CASE 1.

Saturday, August 30, 2008

Computer Vision (39): SDF continued

S1 and S2 - Two sensor positions.
C1 and C2 - Cones of two real points whose focused images are P1 and P2 respectively.
L - Lens.

To solve for a cone, one needs the base diameter and the height, if it is right circular. But for points away from the optical axis the cone is no longer right circular, so in order to solve for it, one needs to know it at two cross sections. Joining these two at all points around it would give you the cones, as shown above. If one can solve this for all points on the sensor, you can then control the focus point through software, post capture. But there is a problem.
C3 - Newly introduced cone.

Wednesday, August 27, 2008

Computer Vision (38): Software Defined Focus

A software solution is always rated over the corresponding hardware one. Some of the reasons being that a software solution is more portable, does not rot or burn and is easily scalable. Macro photographers have a bad time trying to get a perfectly focused image in every one of their shots. In post processing you can only sharpen the entire image a bit but not focus the region you want. Video surveillance systems though would have captured the criminal’s footage might be blur and hence difficult to recognize. What if we could find a solution to correct the focus even after the live video/photo was captured?

I had already written a detailed article on what focus actually means to a camera through simple ray diagrams. Now I will take it beyond and will try to find a solution and define focus through software. I will again use simple ray diagrams to explain my observations and not the complex QED.


To revise the concepts a bit:

1. Light diverges in all possible directions from a point after getting reflected or emitted.
2. When it intersects the camera aperture, only a section of this spherical region enters to reach the camera sensor.
3. A cross section at the face of the lens wrt the point of commencement of light would give us a cone which in turn gets converged by the lens to fall on the sensor.
4. The image this cone creates on the sensor (circle or a point) depends on where along this convergence path it intersects the sensor.

Pr - real point of commencement of light.
Po - some other real point.
L - Lens.
Puu - Unfocused Uncrossed Image of the point Pr with sensor at S1.
Pf - Focused Image of the point Pr with sensor at S2.
Puc - Unfocused Crossed Image of the point Pr with sensor at S3.
S1, S2 and S3 - Different positions of the sensor.

Unfocused image is the one obtained when the sensor is in a position other than S2 for point Pr. Points at different distances from the lens would have different positions of the sensor where they would converge. So, though point Pr would produce a focused image at S2 point Po would register a circle. Focusing a point in one plane wrt the lens would out of focus the points in other planes.

One way to get most or all the points to appear in focus is to increase the depth of focus by decreasing the aperture to as minimum (in size) as possible. This would make the light cone very narrow in width and the digital sensor would not capture it as a circle in spite of not being focused.


Now a days due to the increase in the sensor pixel density even a slight movement in the sensor from the position S2 would give an out of focused image of Pr. Though auto focus systems might be very accurate even a small movement in the camera position would out of focus the desired point in the image especially for macro shots in wide aperture. Once the image is frozen there is no way you can correct the focus except for a slight sharpening. This is because a circle is not uniquely defined by a cone. For example in the figure shown below image of the circle Ic on sensor S can be formed by any of the cones getting focused at F-1 (before the sensor), F1, F2 or F3 (after the sensor). In this format there is no way we can bring the focus back once the image is captured since we do not know the cone whose cross section this is.


Assuming even intensity distribution some people would argue that irrespective of where this point would converge we could simply take the sum of the spread intensity and put it at the location of focus. This can only be true for points on the optical axis and a single lens system. Camera lenses generally contain many groups of lenses and hence would be complex to analyze. If there are any experts who can solve this problem do let me know.

Secondly, in the software solution that I am talking about as one point gets focused others at different planes should change correspondingly which is difficult in the above method. Moreover it is not possible to isolate a circle of a particular point in a natural scene where millions of points would get mixed up at the sensor. So what is the way out?

Monday, August 18, 2008

Photography and travel: Melkote

One day trip from Bangalore and around 130 km. Take Mysore road initially and take a deviation to the right after reaching Mandya. There is not much to see other than the place where the movie GURU was shot, so not a recommended place to plan an exclusive trip.

This is the place where Ashwaria Rai runs/dances with the ducks for some song that I don't remember currently.

Sunday, August 17, 2008

Photography and Travel: Masinagudi

A weekend trip to wildlife and nature enthusiasts. Masinagudi is 7 km from Mudhumalai elephant reserve in Tamil Nadu. Masinagudi is just a small town and a resting place while the safaris are mainly conducted in Mudhumalai itself. One can also opt for night safari. If you can leave early in the morning Gopalswamy betta would also be a good place to be on the way nera Bandipur. Ooty is just 30 km from here and if there is a day to spare you can drive to the top.

Mudhumalai Morning Safari

Ooty landscape

Gopalaswamy Betta

Wednesday, June 11, 2008

Photography and Travel: Savandurga

Savandurga has the largest monolithic rock in Asia. It is around 45 kms from Bangalore and is good for one day drives.

Road is pretty good at some places along the stretch as you can see below.

One can also visit the nearby dam and backwaters, which is a good spot to tent at night. On the way back you can also visit Dodda Alada Mara (The big baniyan tree). This is where one of the fight scenes of Khalnayak was shot if you remember.

Wednesday, April 2, 2008

Stargazing Olympics 2008:

Being close to the beta version of my MSMM software, I find Olympics 2008 the rightest place to demonstrate it. Here is a list of events to be held in this event: http://en.beijing2008.cn/cptvenues/schedule/. MSMM can be used to depict motion events like; Athletics, Badminton, Basketball, Canoe/Kayak -- Slalom, Artistic Gymnastics, Gymnastics -- Trampoline, Rhythmic Gymnastics, Aquatics -- Diving, Taekwondo, Volleyball and Beach Volleyball in an effective way. In fact from the video transmission perspective once a MSMM image is created there would be no need to transmit the replay of the entire sequence, as this image(MSMM) would depict all of this in a single frame. So it would be a kind of compression in a way.

Any company willing to promote its camera through this software??? Check out the online demo here: http://www.multishotimaging.com/

Friday, March 7, 2008

Voice munching

Lip reading in computer vision tries to uncover the conversation through an audio less video sequence. It tries to exploit the movement of the lips and the jaw which are assumed to have a unique correspondence to what we speak. When it comes to broadcasting of media content for example say a news show, ideally it would only be required to transmit the video of the person and the software residing locally should be able to give you the news through lip reading. But would it really be worth the effort? How much bandwidth would the audio signal after all take? I would rather be impressed if the whole concept was reversed; bring in the lip movement by looking at the voice. Of course this would not replicate the exact video of the person talking, but no other better example could be found to support this concept. Transmit only the first frame of video and guess the next frames, I mean lip movement through the transmitted voice. The bandwidth to transmit a news channel will just be equal to the bandwidth of voice which means I will be able to use my landline phone with no special modem to make a video call. Crazy stuff! But all depends on how best you can make the person’s lips dance to the tunes of his voice.

Thursday, December 27, 2007

Computer Vision (37): Sensing through Seismics, The Golden Mole

Nature has always outwitted humans in its creativity and optimization. Humans are one of the few creatures bestowed with a complex and highly developed visual sensitivity. Even though we ourselves haven't been able to crack the algorithms of our visual cortex, researchers are trying hard to replicate the behavior in robots. I myself have strived for years to unravel the enigma, but in vain. I then started to look out for other suitable ways to allow robots to become autonomous in one way or the other and came across a category of owls that could pin point their prey through hearing and have already blogged about it.
Some pythons have the ability to sense the infra red radiation from creatures and can even use it to hunt down their prey. Usually these are called pit snakes. Though not very well developed they still have eyes for vision, which leave these creatures not that special compared to the golden moles that I came across recently.


These creatures do not have eyes at all. They have extremely sensitive hearing and vibration detection, and can navigate underground with unerring accuracy. Morphological analysis of the middle ear has revealed a massive malleus which likely enables it to detect seismic cues. The make use of this seismic sensitivity to detect prey as well as to navigate when burrowing through sand. While vibrations are used over long distances to detect prey, smell is possibly used over shorter distances.

FORSAKEN FANFARE

Travel and Photography
Had been to kerala recently and wanted to have this post under my usual Travel and photography theme, but the message I wanted to convey was much more than just this, so made that a sub heading. I questioned myself; What does it take to be a celebrity? Fame doesn't shake hands with those who are just talented. There is something else missing in these people which I yet need to discover or probably find out from you. As they say, your family name most of the times would do all the magic in the film industry. Our country with such a bursting population would have a lot of such cases that fail to exhibit their endowment at the right place and time. I met one such case in Fort Kochi.
This person could embellish the algae clad wall with just a few colored chalks, and of course a lot of his esteemed abilities. People gathered to watch him chalk his imagination, but pretended not to recognize that it was not a charity show. He stood there smiling at the audience waiting to at least settle his accounts on the money he had spent for the chalks. It was shocking to see everyone disperse from there without even a single penny flying to his side. Seriously I feel that my Canon 350D failed to reproduce the shades (In fact I borrowed this snap from my friend) that he could create on such a dirty wall. With a canvas I think he will touch the skies.
Here are some of the glimpses of kerala (Cochin, Attirapalli and Alleppey backwaters) through my camera:
http://www.flickr.com/photos/57078108@N00/.

Tuesday, October 30, 2007

Computer Vision (36): Mechanical or Knowledge based CORRESPONDENCE

In spite of expending sleepless nights giving deep thoughts on what could be the technique behind our brain solving the problem of depth perception, my brain only gave me a drowsier day ahead. So I started to filter out the possibilities to narrow down to the solution. The question I asked to myself was; is our brain using knowledge to correspond the left and the right images, or is it something that happens more mechanically? I had tried out a lot of knowledge based approaches, but only in vain and even the discussion that we had in the earlier post concluded to nothing. I wanted to take a different route by thinking of a more mechanical and less of a knowledge based approach. My brain then pointed me to the age old theory proposed by Thomas Young to explain the wave nature (interference) of light, “The double Slit Experiment”. How could this be of use to solve a seemingly unrelated problem of depth perception? On comparing you will find a few things in common between the two setups. Both are trying to deal with light and both of them are trying to pass the surrounding light through two openings and combine them later. I excitedly thought, have I unlocked the puzzle?

Let’s analyze and understand it better to know if I really did! I am neither an expert in physics nor biology, so I can only build a wrapper around this concept and not verify its complete path.

Young’s experiment used a single source of monochromatic light to pass through the two slits and got an interference pattern on the rear screen, the two slits being equidistant from the source. The approximate formula for this experiment is given below.

Where,


λ is the wavelength of the light

s is the separation of the (slits/eyes)

x is the distance between the bands of light (also called fringe distance)

D is the distance from the (slits to the screen/eye and retina)

In this experiment the distance between the light source and the slits was kept constant and the separation between the slits varied. In our case the distance between the eyes remains same object to be seen varies over depth in the 3D space. Assuming D and λ to be constants, decreasing s will increase the frequency of the pattern on the screen. Conceptually decreasing the distance between the slits is equivalent to increasing the distance between the source and the slits. So the frequency of the pattern on the screen can be said to depend on the depth of the source from the slits in the 3D space. If the source is placed along a line bisecting the two slits, the pattern would be symmetric on either sides of this line on the screen. Every point along this line in the 3D space would have a unique frequency and hence pattern. The diagrams here are only for conceptual understanding and not the exact experimental outcome.
As the source starts to move away from this bisecting line the symmetry in the pattern should start to degrade.

If a light source is placed at 3 different locations equidistant from the center of the slits, the one at red would produce a symmetric pattern and the other two I guess would not. I have not experimented this and hence the letters NX (Not eXperimented). If my guess is right, a light source placed anywhere in the 3D space would produce a unique pattern on the screen!!! This means an analysis of this pattern would tell us the exact location of the source in the 3D space.

This concept can be applied to our vision, by replacing the light source by any point in space that reflects light acting as a passive source.

Tuesday, October 16, 2007

Photography and Travel: Kudremukha

This is the season of the year when nature embraces the dry mountains of the western ghats with a green velvety floral carpet. The season after the rains brings with it a fine spray of mist over these hillocks. The drive along the mountain veers amidst fresh vegetation and serene valleys makes your experience even more invigorating.

Wednesday, October 3, 2007

Computer Vision (35): Segmentation Verses Stereo Correspondence

One question that always keeps rambling in my mind is if segmentation is a 2D or a stereo phenomenon. A 2D image on analysis can get us no more than a set of colors, edges and intensities. A segmentation algorithm even though would at its best try to exploit one or more of these image lineaments, would still fall short of what is expected from it. This is because when we define segmentation as a process of segregating the different objects in our surrounding or a given input image, not all them can be extracted using the combination of the cited features. A lot of times we might have to coalesce more than one segment to form an object and the rule to do this has kept our brains cerebrating for decades. Below image is a simple illustration of this.

Hope it doesn’t get difficult for your brain at least to get the contents in the image. On observing keenly, it shows a dog testing its olfactory system to find something good for its stomach. You can almost recognize the dog as a Dalmatian. Now I bet if anyone can get me a generalized segmentation algorithm that can extract the dog from this image!!!

Some people might argue that it’s almost impossible to achieve this from a 2D image, since there is no way to distinguish the plain of the dog from that of the ground. Remember, your brain has already done it! In a real scenario even if we come across such a view our stereo vision would ensure that the dog forms an image separate to the plain of the ground and hence would get segmented due to the variation in depth. Our brain can still do it from the 2D image here due to the tremendous amount of knowledge it has gathered over the years. In short, stereo vision helped us build this knowledge over the years and this knowledge is now helping us to segment objects even from a 2D image. The BIG question is, how do we do it in a computer?

Whenever I start to think about a solution to the problem of stereo correspondence the problem of segmentation would barricade it. This is why. The first step to understand or solve for stereo correspondence is to experiment with two cameras taking images at an offset. Below is a sample image.

It is very obvious that we cannot correspond the images pixel by pixel. Which blue pixel of the sky in the left image would you use to pair with a particular blue pixel in the right? Some pixels in the right would not correspond with the left and vice versa, but how do you know where these pixels are? This again loops us back to use some segmentation techniques to match similar objects in the two images, but I think we had just now concluded that segmentation was due to stereo!!!

Thursday, September 20, 2007

Computer Vision and Photography (34): The Focus Story Continues…

Was thinking of ways to solve the auto focus problem perfectly even for surfaces without any contrast and came up with a few proposals which I do not know are practical or not. Light that is converged by the lens we all know forms a 3D cone between the point from where it takes off and the surface of the lens towards the point. This cone will be right circular only for points on the optical axis. The lens only sees the 2D projection of the light points existing in the 3D space around it and so there can only be one point on the optical axis that the lens will be able to see at any instant of time. This is the point that I am trying to FOCUS on. I need to somehow come up with a technique to detect the rays emerging from a point on the optical axis.
One difference between a point on the optical axis and any other is that, rays emanating from the former point meet a circle of radius R drawn from the center of the lens, at the same angle.

But any point on the lens would receive light from all points visible around it. So at any point on the lens light rays will be converging from every possible angle, which leaves us with no way to pinpoint the ray that started from the optical axis.There are many more problems with this very way of thinking to solve the problem. From the perspective of the lens we never know where the real point is located on the optical axis. Different points from the surrounding space can create the same effect as though there was a real point at a different location on the optical axis. This indeed can happen continuously all along the axis! Assuming that the frequency of light reflected from a real point will almost be the same when it meets the circle and the probability of such a thing happening for a virtual point zero, the problem could be solved. But if you recall, the very reason why I started to think about this, was to get a solution to cases where there is zero contrast.
After a while I came across a theory called QED that solved a lot of these problems but kept the hardware required to achieve it out of our current technology’s reach. According to QED, a photon represents the “particle” of light, and its instantaneous phase the “wave” counterpart. This phase depends on the frequency of the light under consideration. A lens focuses light because the probability that the photons reach the focus point with the same phase is high and zero anywhere else. For more details refer to the book “QED: The Strange Theory of Light and Matter”. Putting the same theory into action for our current scenario, this would hold good only for a real particle. Since phase is something that repeats as the photon travels through space, the random points that form the virtual particle should be present at exact locations (again that can repeat in space) to meet the point “a”, all with the same phase!, which is highly improbable in a practical scenario. Now this should work for ZERO contrast!

Saturday, August 25, 2007

Computer Vision and Photography (33): Capturing stereo images using a single camera

Chak De INDIA

When I started my work on stereovision I used to download stereo images from the internet for my experiments. I didn’t have a camera then. By the time I could afford one, the statement “taking images at an offset” had become as synonymic as stereo images. Since I couldn’t afford two similar cameras for both cost and reasons of inutility I started to take my own stereo images by moving the camera a bit to the side (to create that offset) for shooting the second image. I even went to the extent of thinking to developing an attachment for my tripod to create a flat base for this movement! I always thought that research would make the mind sharper to thinking of new ways to deal with the subject, but seems like it didn’t work out in my case at that instant. Instead of developing an attachment for my tripod, why couldn’t I think of developing an attachment for my camera so that I could totally eliminate this requirement of moving it for shooting the second image? It didn’t take much time for this innovation to bloom in me and I immediately rushed to the nearby glass vendor to prepare this arrangement.

It is simple (figure only for conceptual understanding). Like our eyes, I keep two plane glasses at an offset and direct light to the lens of my camera through a pair of pair of 45 degree mirrors as shown in the figure. So in a single sensor I capture both the images, each on one half of it.

Given a camera and a requirement to take stereo images, this arrangement was anyone’s mind game. I thought simple things like these need not be documented. A few months after this finding I saw a paper on exactly the same concept from a university in Switzerland. Do we really need PhDs to build this school level optics? Now I know the reason behind the very poor numbers behind India’s contribution towards the world’s papers and patents. Are we too fainéant to put our findings on paper or do we think we are not up to the mark in creating new things when compared to others? That too when others are confident of such simple findings.

CHAK DE INDIA: FLAUNT YOUR FINDINGS

Saturday, August 11, 2007

Computer Vision (32): Monocular Cues

Under this topic, focus was the only one I wanted to drag a lot since it is very much required for my techniques on depth perception. I will mention a few other cues here that are even though very obvious to anyone, are required for the completeness of the topic. Monocular cues are better understood from photographs, so here’s one to explain all the three I will be mentioning:

Taking any one of the cars in this picture as reference we can very easily guess the relative position of the others in the image. This cue is called “Familiar size”. It works not only for similar objects but anything around you. Where this cue takes a beating sometimes there is another that drops in to resolve this issue; “Interposition”. On the left we have a red and a silver car projecting the same size even though both are not at the same depth. How do I know? My brain tells me that some portion of the red car is occluded by the silver one which means that the latter should be in front of it.

Bringing in the rest of the image you can see that the road in the above figure appears to get narrower farther it is considered from the camera. Taking this cue as reference you can almost separate the different regions in this image into their depth categories. The small hilly region on the right is farther away from the lake on the left. The fountain is definitely closer to the camera than the lake, etc. This is called “Linear Perspective”; the convergence of parallel lines as they move away from you.
All these cues supplemented with our knowledge will always give us if not accurate a misty information about depth even in a 2D scenario.

Thursday, July 26, 2007

Ideas and Technology: More Intelligent Alarms

From the time I started traveling by bus to office, I am wasting a minimum 2hrs every day. Minimum because, the time it takes to travel depends on various factors like the condition of the traffic and the driver. Some of them make it in just 50 min, while some take me on a long 90 min journey. With more traffic it only gets worse. Reading is not something I can do, due to low light in the evenings, so the best thing could be to take a nap. I tried this option a lot of times, but my mind always tries to be over careful so that I just don't miss my stop. The result; I only end up resting my eyes. Probably I could set off an alarm at 50 min starting from the departure time, but then on quite a few occasions it wakes me up midway whenever a slow driver blends with a bad traffic. Why can't alarms be more intelligent? How do I make sure my alarm gets active at almost the same time that my bus reaches a specific PLACE. There you are, my alarm should better track the location than time! A lot of people think of GPS whenever it comes to keeping track of the location. But then it would require that a GPS receiver be integrated into your most commonly used, all in one device; your mobile. If it is too complex then forget it, I don't want it to be very accurate for this application. I am OK with my alarm waking me up on reaching the nearest tower to my house, which can be done with conventional mobiles as they will know the location code of the towers. I think, if this was so simple, mobile companies would have implemented it way back, or probably it has done it but I am unaware of one? I found quite a few number of literatures on this topic on the net, but most of them just think of GPS when it comes to location tracking. Hope this is a much simpler way out. Well, for my problem at least!

Sunday, July 22, 2007

Computer Vision (31): "Seeing" through ears

Till a few days back even I wasn’t aware of the existence of such creatures in Nature. I had not even thought of trying out something like this, even though it has been years getting into researching in this field. Nature again outwitted us in its design and complexity. I am actually talking about creatures having ears at a vertical offset to extract yet another dimension; depth that our ears/brain fail to solve through hearing. The Great Horned Owl (Bubo Virginianus), the Barn Owl (Tyto Alba) and the Barred Owl are some of such Nature’s selected gifted creatures. This offset helps them to hone on a creature with more sensitivity and helps them hunt down creatures even in complete darkness. With this ability they don’t even spare creatures like mice that usually hide under snow and manage to escape from their sight. Evolution has created wonders in Nature. These predators usually live in regions with long and dark winters and hence have developed the ability to “see” through their ears.

But how does it all work? With just horizontal offset our ears manage to tell us the direction of sound in the 3D space. Imagine it to be an arrow being hit in that particular direction. You don’t know the distance of the target but just fire it in that direction. The arrow actually leaves from a point which is the horizontal bisector of your ears. Applying the same concept on vertical offset there will be another arrow leaving from a point which is the vertical bisector of the ears (in the case of these specially gifted creatures). From primary school mathematics we all know that two straight non parallel lines can only meet at one point in space, which in this case happens to be the target.
Even Nature can only produce best designs and not perfect ones and the Owls will definitely have to starve if their prey manages to remain silent. To make its design more reliable and worthy, Nature has never allowed a prey to have this very thought in its mind.

Saturday, July 21, 2007

Computer Vision (30): Why wasn't our face designed like this?

I have already touched upon the reasons behind having two sensors and their advantages (1). We now know why we were born with two ears, two nostrils and two eyes but only one mouth. What I failed to discuss at that point of time is their placement. I will concentrate more on the hearing and vision (placement of ears and eyes) which are better developed in electronics than the smell.
Our eyes as any layman will be aware of, is a 2D sensor similar to the sensors present in a camera (not design wise of course!). They capture the 2D projection of the 3D environment that we live in. In order to perceive depth (the 3rd dimension), our designers came up with the concept of two eyes (to determine depth through triangulation). For triangulation (2) to work the eyes only needed to be at an offset; horizontal, vertical, crossed anything would do. So probably the so called creator of humans (GOD for some and Nature for others) decided that they would place them horizontally to give a more symmetric look wrt our body. But why wasn’t it placed at the side of our head in place of our ears? (Hmmm, then where would our ears sit?).
Light, we know from our high school physics does not bend along the corners (I mean, a bulb in one room cannot light the other beside it), and from my earlier posts we know that to perceive depth our eyes need to capture some common region (the common region is where we perceive depth), they have to be placed on the same plane and parallel to it. This plane happened to be the plane of our frontal face and so our eyes came at the front where they are today.
Let’s move on to hearing now! When we hear some sound, our brain will be able to determine the direction of it (listen to some stereo sound), but will not be able to pin point the exact location (in terms of depth or distance from us) of the source. That is because our ears unlike our eyes are a single dimensional sensor able to detect only the intensity of sound (of course they can separate out the frequencies, but that is no way related to direction of sound) at any point in time. In order to derive the direction from which it came our creators/designers probably thought of reusing the same concept that they had come up for our sight and so gave us two ears (to derive the direction of sound through the difference in timing when it arrives at each of the ears). To derive the direction, our ears only needed to be at an offset; horizontal, vertical or crossed, so to give a symmetric look they probably thought of placing it at a horizontal offset. But why was it placed at the side of our face and not at the top like a rabbit, dog or any other creature?
Again from high school physics we know that sound can bend along the corners and pass through tiny gaps and spread out. So you can enjoy your music irrespective of where you are and where you are playing it in your house (well! if you could only adjust with its differing intensity). So our ears never demanded to be on the same plane and parallel to it! The common signals required to perceive the direction would anyway reach it irrespective of its placement since sound can bend. Secondly our ears were required to isolate the direction in the 360deg space unlike our eyes that only projects the frontal 180deg. Probably the best place to keep it was at the side of our face.
Our 2D vision was now capable of perceiving depth and our 1D hearing could locate the direction. Since our visual processing is one of the most complex design elements in our body and very well developed to perceive depth the designers never thought of giving any extra feature for any of our other sensors. But Nature has in fact produced creatures that have a different design to what we have, with ears at a vertical offset, on top of their head, etc, which I will be discussing in my next post.

References:
1. http://puneethbc.blogspot.com/2007/03/computer-vision-4.html
2. http://puneethbc.blogspot.com/2007/04/computer-vision-13.html