Thursday, June 26, 2014

EPSRC 3.5 year PhD studentship: Eye guidance in real-world scenes (UK students only)

Can't recommend this opportunity more highly. This is a chance to work with two pioneers of active vision (Nuthmann) and computer vision (Fisher) in one of my favourite cities on the planet!

.....
Dear students,

A (last-minute) fully funded(!) PhD studentship is available to work with Dr. Antje Nuthmann (principal supervisor) and Prof. Bob Fisher (co-supervisor) at the University of Edinburgh. 

Details can be found here: http://www.ppls.ed.ac.uk/students/postgraduate/PGFunding.php#EPSRC

Please don't hesitate to contact me if you have any questions.

Best wishes,
Antje Nuthmann
http://nuthmann.de/antje/Site/Welcome.html

Sunday, November 03, 2013

BIMI study day: Cognition at the Movies




I am proud to announce the first 'Cognition at the Movies' study day hosted by myself and Birkbeck Institute for the Moving Image (BIMI).

Saturday November 9th 10am to 6:30pm (B35, Birkbeck, Malet Street, London, WC1E 7HX; number 1 on the map http://www.bbk.ac.uk/downloads/centrallondon.pdf)

Since cinema’s inception filmmakers and theorists have been interested in the relationship between film and its audience. How do directorial decisions influence what we see on the screen and how does a viewer’s prior beliefs and interests influence how they experience a film? Cognitive Science, the interdisciplinary investigation of mental phenomena using theories and techniques from neuroscience, psychology and philosophy has recently begun to be applied to these questions of film cognition. This workshop will bring together film theorists, cognitive psychologists and philosophers in an exploration of the relationship between film and its audience.

Keynote presentation ‘by Prof Torben Grodal (Copenhagen) author of Moving Pictures and Embodied Visions.


Comic Entertainment, Film, and the Embodied Brain
The lecture will first provide a short description of how muscles and action is important for the embodied brain and for our experience of narratives.  The basis for the standard narrative reflects the Brain’s PECMA flow: Perception, Emotion, Cognition, and Motor Action. Characters and viewers want to modify some states of the world by motor action, including verbal actions. The lecture will then discuss the embodied brain’s three ’bail out’ mechanisms where the modification of the world by action is supplanted with self-modification: Crying, as in sad melodramas, laughter, as in comedies, and freeze reactions as effects of sublime submission to the exterior world. The lecture will especially focus on comic entertainment and discuss the processes that allows the brain to evaluate something as ’not real’, as ’not a cause for action’ and redirect the arousal from a given scene from tense world-directedness to laughter. It will finally discuss the social nature of comic entertainment and those mammalian play-functions that serve as facilitators for the reality status evaluations in comic entertainment that makes it possible to experience shame, failure and other negative events with a strongly positive hedonic tone.

Event is free but please register as spaces are limited: https://cognitivism.eventbrite.com/

Schedule:


9:30-10 Registration

10-10:15 Welcome and Introduction
10:15-11 Prof. Ian Christie (Birkbeck) - Psychology in the dark: just what is it we want to know?
11-11:45 Dr. William Brown (Roehampton) - He(u)retical Film Theory: Cinema and the Brain
11:45-12:30 Prof. Sheena Rogers (James Mason U.) - Towards Transcendence: Cognitive Components of the Sublime in Art

12:30-1:30 Lunch break

1:30-2:30 KEYNOTE: Prof. Torben Grodal (Copenhagen) - Comic Entertainment, Film, and the Embodied Brain
2:30-3:15 Dr. Paul Taberham  (Kent-Canterbury) - Avant-Garde Film in an Evolutionary Context
3:15-4:00 Prof. Murray Smith (Kent-Canterbury) - Murder Ballads

4-4:15   Coffee break

4:15-35 Steve Hinde  (Bristol) - A Study of Attention  while People Watch Movies.
4:35-5 Parag K. Mital (Goldsmiths) - Resynthesizing Perception
5-5:45 Dr. Tim J. Smith  (Birkbeck) - Cinematic Universality: Do I see the same movie you see?
 
5:45 - 6:30 Discussion


6:30 + Wine reception

For further information email tj.smith@bbk.ac.uk 

Tuesday, October 01, 2013

Where’s Walter? How the finale of Breaking Bad used your eye movements to build suspense.


Where’s Walter? How the finale of Breaking Bad used your eye movements to build suspense.
By Dr. Tim J. Smith and Rebecca Nako

WARNING: If you are a fan of Breaking Bad and have not yet watched the finale do not read on. Spoilers ahead!

So it is all over. The season finale of Breaking Bad has been eagerly devoured by fans in America and across the world. We now know the fate of the anti-hero, meth-cook extraordinaire, Walter White and the extended group of characters both good and evil. The final episode is a masterful end to an exceptionally crafted series that has always found the perfect balance of intensity, humour and nuanced plot. It also serves as a wonderful demonstration of how the loyal Breaking Bad viewers have often been complicit in the creation of the tension. As Walt’s character develops across the series he becomes a menacing figure who’s actions are less and less predictable. Unlike some of the gangster characters he encounters (and usually defeats) he is not prone to irrational outbursts or  sudden violence. Walt’s menace comes from his intellect and cool planning. We know his actions are morally wrong but throughout the series we continue to empathise for Walt and see the action from his perspective, rooting for him to succeed. This is perfectly exemplified in the season finale in which the action builds slowly to an ultra-violent crescendo in which Walt’s ingenuity triumphs over those who have wronged him and his family. The slow scenes that build to this climax show us how he first gets his own back on the Schwartz’s, the couple who he believes profited off his early ideas, Lydia, Walt’s meth distributor, and then says goodbye to his wife, Skyler. Each scene is incredibly tense to watch but, unlike most mainstream TV or Cinema the tension is created by the viewers, not typical cinematic devices.  The action is often framed in shots that linger on the screen for an uncharacteristically long amount of time for TV. Hidden within the shot, unbeknownst to either the viewer, the characters or both, is the menacing figure of Walt. The director is creating tension by playing a game of ‘Where’s Walter?’ with the viewer.

For example, in the sequence in which Walt confronts Lydia, Walt seems to appear from nowhere in the café after Todd, the new meth cook arrives. Walt’s appearance behind Lydia and Todd shocks the audience but not through a traditional use of sudden cut to close-up or dramatic change in accompanying music. The shock comes from the viewer surprise that they didn’t notice him in the scene earlier. However, if we look back at two earlier shots of the café, Walt can clearly be seen sitting off to the side.


When we first watch this sequence our interest is in Lydia and her conversation with Todd. Due to physiological limits in what we can attend to and see at any one moment we have to choose where to fixate in the scene. By predicting our interest in Lydia, the director (Vince Gilligan, the creator of the series) ensures that our eyes linger near her and do not locate Walt in the periphery. There are many techniques for influencing where people fixate in a film, as I have discussed in length elsewhere (Smith, Psychocinematics, 2013) but one of the strongest is to make the viewer want to look somewhere. If the viewer is complicit in their choice of fixation location they will be even more surprised when it is revealed that they failed to see something. This is a technique of subtle misdirection that magicians have used for centuries and we have recently shown can operate in magic tricks even when visual cues are used to try and force the viewer to look at the source of the trick (Smith, Lamont, & Henderson, Perception, in press). In this scene from Breaking Bad, the director uses this inattentional blindness to play with the viewer and reward the active viewer who discovers Walt before Lydia and Todd do. Along with this episodes dense use of back-references to earlier plot points, subtle character cues and symbols (such as Jesse’s box), this game of hide-and-seek with Walt serves to reward the committed viewer with a sense of discovery and enrichment of what will be their final glimpse of this world.

The impact of this knowledge on how people watch this scene is evident if you record their eye movements. Using a Tobii TX-60 eyetracker, I recorded the eye movements of two participants watching the café scene. One participant had never seen Breaking Bad before (Yes, I ruined the whole of Breaking Bad for her by showing her the finale!). The other participant was an avid fan who had already watched the finale the night before. If we visualise their eye movements as red dots on top of the video (see below) we can see how their eyes and their attention shift across the screen. Each red dot signifies the location of the viewer’s gaze during one sample of the eyetracker (1/60th of a second). When the gaze clusters together in one place their eyes are in a fixation. When the gaze suddenly jumps to a new location they are performing a saccade.




At the beginning of the clip we can see how both the experienced (bottom video) and novice viewer (top video) track Lydia’s bag as she drags it through the café and then saccade up to her body and face once she sits down in the next shot. If we were to plot the two gaze patterns on top of each other we would see a striking degree of coordination between the two viewers. This synchronisation of attention across viewers is characteristic of how we watch most TV and film. Although we think we are highly idiosyncratic in how we watch a program most of the time the director is ensuring we all look in the same place at the same time as I demonstrated by eyetracking multiple viewers watching There Will Be Blood here and here.

After Lydia is seated at the table the camera then cuts across the table to reveal the rest of the café to Lydia’s left. Immediately following the cut, the experienced viewer saccades directly to Walt seated in the background. The novice viewer only looks at that part of the frame once the waiter enters the shot and blocks our view of Walt. After watching this clip the experienced viewer stated that he had not known Walt was in this shot until watching the scene during the experiment. His direct saccade to Walt suggests that knowing Walt would appear in the scene at some point had primed his attention and made it easier for him to find him earlier than the novice viewer.

The camera then cuts to Lydia and a series of close-ups of her face and the Stevia sweetener which will later play a critical role in the scene. We then cut back to a longer shot of the café in which Walt is now lurking discretely. At this point both viewers saccade directly to him  even though he hasn’t yet moved and the main action is still taking place in the rear of the shot. Both viewers now know Walt is present in the scene and about to approach Lydia and Todd who are still unaware of this presence. This mismatch between what the viewer knows and what the characters know creates tension about what will happen next.

 An even more impressive use of camera positioning to create tension occurs later in the episode when we overhear a phone conversation between Skyler and her sister, Marie. The sequence begins with a slow camera pan across Skyler’s new apartment as the phone rings and the answerphone picks it up. We see Skyler smoking at the kitchen table. Our view of the kitchen is complete except for a small patch occluded by a column at screen centre (image above). The slow pan of the room and this final long shot suggests that Walt is not in the scene. However, after several cuts back and forth between Skyler and Marie as Marie informs Skyler of Walt’s presence in town the viewer begins to get the impression that Walt may be either coming for Skyler or already be hiding somewhere in the scene. After Skyler hangs up, the camera cuts back to the earlier long shot and we again see that the scene is empty. This belief is trashed as the camera slowly moves into the scene and Walt is revealed behind the column. We were denied knowledge of Walt’s presence in the scene by the director due to the clever choice of camera position. The tension we feel after discovering Walt’s presence is due to the mismatch between what we have known up until that point and what Skyler must have known all along: Walt is present in the room. What will Walt do next and why is Skyler hiding his presence from Marie? Our sudden awareness of Walt creates a flurry of questions and an interest in how the scene develops.


Watching the eye movements of our two participants viewing this scene reveals the strong influence knowledge of Walt’s presence has on our experienced viewer’s eye movements. Both viewers begin the scene by saccading around the apartment, fixating and tracking objects as they are revealed by the panning camera. As soon as the column behind which Walt is hiding comes into view the experienced viewer (bottom video) becomes obsessed with this boring piece of architecture. His gaze dwells on the column and saccades around its edges, trying to find some evidence of Walt. By comparison, the novice viewer saccades directly to Skyler and focuses on the phone conversation.


After the conversation finishes and the camera cuts back to the long-shot (02:19) the experienced viewer immediately saccades to the column and then saccades back-and-forth between the column and Skyler, waiting for Walt’s reveal. The novice viewer glances briefly at the column but mostly concentrates on Skyler. It is only once the camera begins moving in that her attention to the column increases and finally peaks once she catches a glimpse of Walt’s jacket sticking out behind the column. The novice viewer is actively viewing the scene; trying to check that Walt isn’t present given the suspicion Marie has just created but she can only see what is visibly present in front of her. The experienced viewer perceives Walt behind the column in the very first shot due to his memory from previously watching the scene. The experienced viewer’s gaze interrogates the column seeking out confirmation of Walt’s presence even though for the majority of the scene all  you can see is a bland wood column.

These example scenes demonstrate how film and TV can create suspense by withholding information from either the viewer (e.g. Walt’s presence in the kitchen with Skyler), the characters (e.g. Lydia and Todd’s knowledge of Walt’s presence in the café) or both. Often such suspense is created by not cutting to a detail that we desperately want to see. However, such techniques can often appear heavy handed and position the director at odds with the viewer. In the scenes discussed above the director cleverly plays around with what the viewer can and cannot see whilst always giving us the impression that we have access to the full scene. This false belief makes Walt’s eventual reveal all the more powerful. This is further evidence for why Breaking Bad was such exquisite TV.

*If you are interested in seeing more examples of how our expectations about a dynamic scene can influence where we look check out my recent study published in Perception. This study used a simple card trick to bias participant’s gaze towards one part of the screen whilst the trick occurred in plain sight elsewhere on the screen.  Eye tracking revealed that participants completely fail to look at the location of the trick during the first viewing due to their own belief about what is relevant. During a second viewing all participants look in the right place and see how the trick worked. These findings (and the Breaking Bad examples above) are completely at odds with most current theories of how attention is guided in dynamic scenes which state that basic visual features such as motion guide attention (see my article on the topic here; Smith & Mital, JoV, 2013).

** CAVEAT: The two participants tested above may be extreme examples of how a novice and experienced viewer might watch these sequences and a full empirical study would require a larger sample of participants in each group. Effects such as the bias of the experienced viewer’s gaze to the column are unlikely to be absolute but may prove to be statistically significant if the gaze was quantified  across more participants. The precision of the eyetracking, the synch of the audio during playback and the image quality is also not adequate for a full empirical study (hence why the gaze sometimes seems to be offset from objects in the scene). However, this quick and dirty demonstration allows us to quickly analyse the scenes whilst the episode is still fresh in people’s minds.

Wednesday, July 17, 2013

Melbourne's Eye Tracking and Moving Image Research Group

The application of eye tracking technology to questions of moving image spectatorship has risen in popularity over the last couple of years. Aided by the dropping cost of eyetracking hardware and the ease of use of presentation and analysis software eyetracking is becoming practical for researchers across a broad range of disciplines.

For example, the recently formed Eye Tracking and Moving Image Research Group lead by Sean Redmond and Jodi Sita in Melbourne has two central goals:

" bringing the group together; we wanted to utilise eye tracking technology more centrally in the analysis and examination of the moving image; and we wanted to draw together scholars and practitioners from the Sciences, and the (Creative) Arts and Humanities so that different modes of enquiry and theoretical and methodological apparatus were placed in the same analytical arena."



I very much welcome groups like this and hope their research proves fruitful and informative. Only by spreading the research questions across a broad range of researchers can we hope to tackle the complex questions of spectatorship and by sharing research methods/techniques we can avoid each of us reinventing the 'eyetracking and film' wheel. To that end I hope my recent publications on the topic can serve as a useful starting point for researchers beginning to apply eye tracking to these questions:

  • Smith, T. J. (2013) Watching you watch movies: Using eye tracking to inform cognitive film theory. In A. P. Shimamura (Ed.),Psychocinematics: Exploring Cognition at the Movies. New York: Oxford University Press. (pdf)
  • Smith, T. J. (2012) The Attentional Theory of Cinematic Continuity, Projections: The Journal for Movies and the Mind. 6(1), 1-27. (pdf)

Thursday, June 06, 2013

Reaction videos capture the heart (and twisted soul) of Film/TV

Videos of viewer reactions to film and TV have been cropping up on Youtube for several years and I have always been fascinated with how they capture what, for me is the heart of the the cinematic experience: the power to manipulate our minds and emotions. Watching a grown adult shriek, cry, and laugh in response to artificial patterns of light and sound on a screen demonstrates the power film and TV have over us and why we seek it out as a source of entertainment, escapism, and emotion that in everyday life would be considered too extreme or dangerous.

I've been meaning to post on this topic for a while but had to act this week in response to the TV event that was the Game Of Thrones 'Red Wedding' (series 3, episode 9). For fans of the books, the events of this episode came as no surprise as they had already been traumatised by the shocking and brutal events when reading the third book in the series, 'The Storm of Swords'. *spoilers* For the fans of the TV series, subtle manipulation of plot events and characters by the producers of the series meant that fans of the show had no way of knowing the duplicitous nature of Walder Frey and what was about to happen in the episode. As the Stark's celebrate the wedding of Edmure Tully to one of Walder Frey's daughters they are brutally slaughtered by the Frey's and the Bolton's, swarn bannermen of the Starks who have secretly made an allegiance with their enemies, the Lannisters. Three of the main characters of the series are brutally killed in a matter of minutes: Robb Stark (the King In the North), his mother, Catelyn and Robb's new bride, Talisa who was pregnant with their first child. The brutality and graphic nature of the murders is shocking even to readers of the books who knew what was going to happen. For viewers of the series without prior knowledge of the events, the murders were..... well judge for yourself *end of spoilers*




The facial reactions, gasps, screams, and comments directed to the screen and to the person filming (who had obviously read the books and knew what was going to happen, filming their friend's/partner's/parent's reactions with sadistic glee) show how involved they are in the series and the characters. For a brief moment at least, they are as moved by the deaths of these fictitious characters as they are to the news of a friend's death. Emotionally, they are across the fourth wall.

Some other classic reaction videos are the reaction of children to the reveal at the end of Empire Strikes Back that Darth Vader is Luke Skywalker's father:


Not surprisingly most of the videos that are used to elicit these shocking reactions are either horrific, shocking, pornographic or sometimes all three at the same time. This brings me to the reaction videos that started the whole craze: People's reactions to '2 Girls 1 Cup'. I'm not going to say anything about the original video as it is too vile to mention but there are several reaction videos on-line from which you can infer the general events of the stimulus. Here are selection:




The whole gamut of human emotions are right there in those videos.

Thursday, March 21, 2013

Film reflected in the human eye



Simply beautiful!

I could spend days watching the human eye move in response to film.....Oh, that's right, I do!

Friday, March 08, 2013

Park Chan-Wook on crosscutting




Interesting interview with Park Chan-Wook about STOKER on aintitcoolnews.com. here are some excerpts on his distinctive visual style.

"But it's not always the case that you can explain everything with words. For instance, "Why do I want to use the color red here? Why do I want the camera to move forward here?" Sometimes I make those decisions because I just feel that that would be the best thing to do here. But after I'm finished making the film and watch it later on, I realize why I made those choices. Take the crosscut for instance. I always thought that decision made no sense because of some musical explanation or reason. That is to say, there is a rhythm to the way scenes are crosscut, and I just like the rhythmic nature of using crosscut. However, later I realized, having seen the film again, what I really wanted to express by using these crosscuts is the concept of fate. In other words, crosscut mixes different individuals' past and present, reality and fantasy. Crosscutting is an effective way to weave these together, and the result that I was looking for in doing so is saying, "It is all a fabric, part of a bigger fate."

Discussing a beautiful transition between Nicole Kidman's hair and a field of grass:

"That particular transition didn't really take much in the way of deliberation. It's something that easily came out, so much so that I can't even remember the thought process that brought me there. It was almost instinctive. All throughout the film, I decided to use crosscutting technique. Once I decided on that, I promised myself that I will have one principle that I will abide by, and that principal was each shot, and how they move on to the next - whether the cuts would crash or whether the cuts would continue on smoothly - how each shot transitions into the other shot in these crosscut sequences, it needs to be something very well designed. So I applied many techniques to achieve this, be it match cut, dissolve, what have you. It's born out this base principle. But I did think that this particular transition was particularly important because it was a transition from a mother moment going into a father moment."

STOKER is an audiovisual masterpiece that perfectly continues Chan-Wook's style into Hollywood. At times the aloofness of the characters, lack of explanation and depth of character can make it seem cold and distancing but I personally marvelled in the film's style and this more than made up for its problems of pacing and failure to thrill (which is ironic considering how Hitchockian it feels!).

The style is opitimised by this wonderful trailer recut to DJ Shadow


Thursday, January 31, 2013

Using Superpowers in Virtual Reality to Encourage Prosocial Behavior



Rosenberg RS, Baughman SL, Bailenson JN (2013) Virtual Superheroes: Using Superpowers in Virtual Reality to Encourage Prosocial Behavior. PLoS ONE 8(1): e55003. doi:10.1371/journal.pone.0055003 (link)

A new study from Stanford University shows that being given the superhero power of flight in a virtual environment immediately changes your likeihood to help another person in the real-world:

From DiscoveryNews

"For the study, 30 female participants and 30 male participants were immersed in a foggy virtual reality city and given the power of flight — like Superman — or the experience of riding as a passenger in a helicopter. Those groups were then assigned one of two tasks: help find a missing diabetic child in desperate need of an insulin injection or leisurely tour their virtual environment. Therefore, the study was a two-by-two design, with participants assigned to one of four groups.

After their VR experience, participants were taken out of their head-mounted-display masks and asked to have a seat. While the experimenter fumbled with the VR equipment, she “accidentally” knocked over a cup of 15 pens sitting on a table near the participant’s chair.

Researchers found that participants who experienced the power of flight in virtual reality were not only quicker to help pick up the pens than their helicopter-riding counterparts, they also picked up more pens. Of the six participants that didn’t help, all were in the helicopter condition. The task of ‘helping the diabetic child’ showed no main effect; only the superpower of flight did."

This is a very nice controlled, empirical design that for once discusses the positive potential of playing computer games.

Monday, January 28, 2013

Walter Murch on the collaborative creation of continuity


"Your job [as an editor] is to anticipate, partly to control the thought processes of the audience. To give them what they want and/or what they need just before they have to "ask" for it- to be surprising yet self-evident at the same time. If you are too far behind or ahead of them, you create problems, but if you are right with them, leading them ever so slightly, the flow of events feels natural and exciting at the same time." 

Walter Murch, In the Blink of an Eye (2001; 2nd edition; page 69)

Friday, January 25, 2013

Steven Spielberg on attentional synchrony



Tom Shone interviewing Steven Spielberg about Lincoln in Sunday Times Culture magazine 20/01/13

"He still goes to see movies - picks an out-of-the-way cinema, sneaks in with his wife or kids after the lights have gone down, then disappears again as the credits roll. He always takes an aisle seat and buys no food or drinks for himself. He's just there for the film, or, more specifically, the film and its audience. He loves feeling the heat rise in the cinema during an especially exciting action sequence, or after a gag has rocked everyone back in their seat.

Spielberg- "You walk into an air-conditioned, freezing theatre and, about 20 minutes in, it starts to get really hot. People start making noise and having a good time. You're lifted by it. The first thing that happens is, people stop eating. They even stop swallowing."

At this point, the third-person plural drops away. "And all of us go into a kind of lock step where, if we were watching a tennis match, you'd see that perfect synchronicity of heads going left-right, left-right. The same thing in a movie theatre, when the movie is working and the audience is galvanised, almost hypnotised, all watching the same things, all knowing where to look at the exact same time...it's a wonderful thing. There is nothing greater than that."



sport wimbledon baltacha 1280x704 web from TheDIEMProject on Vimeo.


Monday, January 14, 2013

1+3 yr MRC PhD studentship on Autism, home eyetracking and cultural differences (Japan/UK) available








I am very pleased to be able to announce that Dr Atsushi Senju and myself have a fully funded PhD studentship starting October 2013. Details below.

The project will be utilising similar home eyetracking technology to that recently demonstrated by Tobii at CES. See the video above for a sneak peek.

>>>>
We are pleased to offer a full 1+3 year MRC Industry CASE PhD studentship entitled "Going Global: Application of Portable Eye-tracking Technology to Study the Effect of Cultural Norms on the Development of Social Cognition". The studentship will be based at the Centre for Brain and Cognitive Development, Department of Psychological Sciences, Birkbeck, University of London, and be conducted in conjunction with Acuity ETS Limited & the Institute of Psychiatry. The studentship will cover course fees at the usual level for UK and EU studentships and a stipend in accord with research council rates.
Much of what we currently know about the developmental disorders comes from Western cultures, and few multicultural studies have been conducted. A major barrier is that the equipment for neurocognitive assessment is often expensive, heavy and requires dedicated lab space, which prevents the assessments being practicable to run in many countries, areas and communities. To overcome this challenge, we will develop a software suite with a portable and affordable eye-tracker, and use it to conduct a series of cross-cultural eye-tracking studies on social cognition in typically developing children and children with autism spectrum disorders (ASD). The successful PhD candidate will take a leading role in this project, including (1) identifying a suitable eye-tracker, developing a software suite, and testing it in the UK, (2) taking this eye-tracker suite to Japan and running the same experiment with Japanese children, and (3) testing children with ASD in both the UK and Japan.
Graduates in experimental psychology or related subjects with a good first degree are encouraged to apply. Experience in some of the relevant research areas and/or methodology (e.g. developmental psychology, autism research, eye-tracking methodology, software development) will be an advantage. Programming experience (e.g. Matlab, Java, C++) or willingness to learn is an advantage. We also expect the candidates to have a high motivation and enthusiasm to the project, good communication and person skills.

The student will receive four year training (1-year MSc and 3-year PhD) in theoretical, methodological, practical and commercial aspects of eye-tracking system. Both the academic supervisors (Dr Atsushi Senju and Dr Tim Smith) have strong track record in eye-tracking research, which will complement the industrial supervisor (Mr Scott Hodgins) from dedicated developers and distributors of eye-tracking system and from the clinical perspective (Prof. Tony Charman). Academic supervisors will also provide training of theoretical background in developmental cognitive neuroscience, autism research and cross-cultural study, development of original research design, programming of stimulus presentation and data acquisition, data recording from infants, children and clinical population, data analyses, and writing-up scientific papers and dissemination to non-academic user communities. The industry supervisor will train the student on the theory & use of eye-tracking in the first instance, and supervise the development of cognitive assessment software suite and the integration of the software suite to the portable eye-tracker.

The Centre for Brain & Cognitive Development (CBCD) at Birkbeck, University of London, has an outstanding track record in training phd students.  Our excellence in training has just been rewarded with the designation “Marie Curie Centre of Excellence for doctoral training” which places us in the top 5% of life science training centres in the EU.  Further, our national training record is reflected in the recent award of the Queen’s Anniversary Prize for Higher Education 2005 for “Neuropsychological work with the very young”.  Acuity ETS is the leading independent eyetracking systems vendor in the world. Acuity is the biggest customer of two of the leading eyetracking manufacturers. Acuity actively strives to encourage collaboration between clients, and to share best practice across the client base.
Further details about the project may be obtained from:

Dr Atsushi Senju
Dr Tim Smith

Further information about PhDs at Birkbeck, University of London is available from:

Application forms and details about how to apply are available from:

Francesca Carter (f.gumbs@bbk.ac.uk)

Candidates must supply a CV, full transcripts of their qualifications and a statement of no more than 500 words indicating what skills and academic and professional experience you can bring to this project and why you consider you would be the best person to undertake this research.  If possible, this should include evidence of your knowledge of the relevant literature in the field.

The deadline of application is 1 March 2013. Shortlisted candidates will be interviewed in late March.

Thursday, December 20, 2012

The vision science of 48fps












Given the recent release of Peter Jackson's The Hobbit in High Frame Rate (HFR) 3D there has been a lot of discussion about the pros and cons of moving from the entrenched 24fps to 48fps (or even the 60fps proposed by James Cameron). I recently weighed in on the vision science behind the perception of higher frame rates for Tested here:

http://www.tested.com/art/movies/452387-48-fps-and-beyond-how-high-frame-rates-affect-perception/

The article is a nice summary of the topics the journalist and I discussed but his personal dislike for HFR overshadows several of my points about why I think the move to 48fps or higher is necessary and will become the standard in cinema. To understand the problems with the current 24fps filming and projection process you need to understand how we are able to see a rapidly presented series of still images as a continuous moving sequence. Here is a passage from an encyclopaedia entry I wrote on film perception a few years ago:

 Smith, T.J. (2010) Film (Cinema) Perception. PDF icon In E.B. Goldstein (ed.)The Sage Encyclopedia of Perception.

"Movies consist of a series of still images, known as frames projected on to a screen at a rate of 24 frames per second. Even though the frames are stationary on the screen and are momentarily blanked as a new frame replaces the old we experience film as a continuous image containing real motion. The two perceptual phenomena contributing to this experience are persistence of vision and apparent motion. Persistence of vision refers to the continued activation of visual neurons after visual stimulation has been removed. During film projection the light is obscured as the frame is changed. If this only happened 24 times a second (Hz) there would be a noticeable flicker. To avoid this flicker each frame is blanked three times by a shutter. This creates a presentation rate above the critical flicker fusion rate of 60Hz. Above this rate persistence of vision ensures that the blank is masked by continued activation of visual neurons and we perceive the projected image as continuous.

The motion we perceive in film is apparent because it is based on static visual information not real motion. It is commonly believed that the apparent motion perceived in films is beta movement. Beta movement is perceived when a simple object such as a line is alternately presented at two different locations around 10 times a second. The two lines are perceived as a single line moving smoothly between the two locations. Due to the slow rate of presentation and the large distances covered, long-range apparent motions such as beta movement are thought to be processed late in the visual system and require inferences based on knowledge of real motion and the most likely correspondences between objects in the image sequence.

Beta movement, along with other long-range motion phenomena such as apparent rotations and transformations may occur during film perception but they cannot account for the majority of motion perceived in film. The 24Hz presentation rate used in film is too fast for long-range motion and film frames are too complex, making the task of identifying corresponding objects in subsequent frames very difficult. Instead, apparent motion in film is due to the same short-range motion system used to detect real motion. Motion detectors in the early visual system respond in the same way to the retinal stimulation caused by real motion and by rapidly presented (>13Hz) static images that depict only slight differences in object location. This processing occurs very early in our visual system and does not require perceptual inferences. The directness with which film is processed results in an experience of motion that is indiscernible from real-motion."

As you can see there are two processes involved that allow us to see a series of frames as motion: persistence of vision and apparent motion. Frames need to alternate faster than ~60 times per second (i.e. Hz)  if we are going to perceive constant luminance, i.e. not perceive a flicker. Old film projectors reached this threshold by using a shutter to present each frame twice (=48Hz) or three times (=72Hz). Modern digital projectors don't have a shutter as the images is constantly present and doesn't need to accommodate the next frame being registered in front of the lens so instead they present each frame 3 times ("triple flash"=72Hz). This is sufficient to remove the flicker but when we move to stereoscopic 3D digital projection we encounter a problem with the amount of light presented during each frame. Most 3D projectors (such as RealD) alternate the left/right eye images, with each being presented at 24fps (24fps x 2 eyes = 48Hz). Each of these left or right images are subsequently flashed 3 times creating a total flicker rate for stereo 3D movie of 144Hz (72Hz per eye)! This ensures that we don't see the flicker in either eye even though they are alternately blind to the image.

Unfortunately, due to the radial polarisation needed to ensure only the left image is seen by the left eye and the right image by the right eye the amount of light reaching the viewer's eyes is significantly less than a traditional 2D presentation. This creates a murkier image and makes it harder to perceive apparent motion as our eyes cannot create the correspondence between moving objects in each frame. This problem is exaggerated by the film being photographed at 24fps per eye. Moments of high camera or object motion create motion blur in the image as the camera's shutter is open too long. This motion blur makes the edges of objects hard to locate and decreases our perception of apparent motion, making the image appear to jump across the screen instead of flowing smoothly. Given that this motion stuttering is happening alternately between the two eyes it makes it difficult for our visual system to fuse the 3D image, resulting in a loss of depth perception and eye strain.

The solution to both problems of light loss and motion stutter in a stereo 3D movie is to increase the frame rate. I'm not sure whether the new HFR/48fps projectors use a double or triple flash but whichever they use the rate of presentation per eye will exceed the critical flicker fusion rate (double flash = 96 Hz; triple = 144Hz per eye). Because each frame is a sharper image with less motion blur the left and right images registered by our eyes will be brighter, clearer and easier to fuse in depth to perceive 3D. Camera and object motion will be clearer as we are better able to perceive apparent motion between the crisper edges of objects and the overall effect should be less cognitive load on the viewer and less eye strain.

The bizarre irony of Peter Jackson's decision to move to 48fps in an attempt to get 3D cinema closer to reality is that it has revealed the artificiality of the Hobbit. As I say in the Tested article, like the move from SD to HD the increased in information on the screen makes the imperfections of the image easier to see. The move to 48fps may not be increasing the spatial resolution of the image but by increasing the temporal resolution (i.e. frame rate) it makes each pixel easier to see and each face prosthetic and matte backdrop easier to notice. Suspension of disbelief is harder in the quiet sequences at the beginning of the Hobbit and it is only when the action picks up in the final act when the higher frame rate and 3D really gel. Many reviewers have reported growing used to the 48fps as the movie progresses and have noted that the chase sequences at the end of the movie are easier to see, more fluid and result in less eyestrain than typically experienced in 3D movies. It is only when Jackson presents a combination of filmed live-action, sets and digital characters or backdrops together on the screen at the same time and gives the viewer time to interrogate the image that viewers seem to have issue with the higher frame rate. We would only really know the impact of the 48fps on filmgoer experience by performing a controlled psychological test on audiences. Viewers would have to be naive to which frame rate presentation they were seeing and various aspects of their experience of the film monitored. Only then could we see if it actually had an impact on their experience without any pre-existing bias against it or resistance to new technologies getting in the way.

Personally I believe the creative potentials of stereo 3D is massive and only starting to be tapped with movies like Scorsese's Hugo and (apparently, although I'm yet to see if) Ang Lee's Life of Pi. If higher frame rates encourage more filmmakers to experiment with 3D without having to worry about viewer eye strain and discomfort I think it is a great step forward.

p.s. Merry Christmas :)

Tuesday, December 11, 2012

Sight & Sound video essay

Kevin B. Lee (@alsolifelike) posted a video essay for Sight & Sound on the evolution of Paul Thomas Anderson's steadicam work which discusses my eyetracking work on There Will Be Blood and the DIEM project here. You can view the video here:

http://www.youtube.com/watch?v=3GGI5mVH6pg&feature=player_embedded

The video essay discusses the careful use of staging and choreography to introduce the viewer to spaces and characters critical to several of Anderson's films including Hard Eight  and Boogie Nights. The use of steadicam in all of these sequences varies from the bravura long and complex sequences of Boogie Nights and Magnolia to subtle uses in There Will be Blood and Punchdrunk Love. This analysis reveals the many ways in which steadicam can create affinity or conflict between what the viewer wants to see and how the camera moves relative to the characters in the scene. This affinity was clearly demonstrated in my analysis of the eye movement behaviour during the table-top sequence of There Will be Blood (as posted on David Bordwell's blog here). By choreographing the camera moves to natural attentional cues such as dialogue switches, character movements and the introduction of characters in from the side of the frame the filmmaker can make a reliable prediction about where most viewers are likely to be attending.

As discussed in the video essay, eyetracking film viewers gives us a direct line to the viewer experience of a film and can be used to validate filmmaker intentions for such sequences. It also provides us with ways to test hypotheses about how production decisions can influence the resulting viewer experience. With the decreasing cost of steadicam equipment and digital production in general the use of such techniques is becoming more and more common. But what this video essay, and my eyetracking research shows is that such sequences will result in viewer disorientation and confusion unless they are carefully designed with viewer sequential attention in mind.  For example, our recent paper in Journal of Vision (http://www.journalofvision.org/content/12/13/3.full) shows how viewer attention can be altered for the same sequence of close-up shots just by excluding audio. I have recently reviewed the influence of such factors and compositional decisions in general in a journal article and book chapter:

 Smith, T. J. (in press) Watching you watch movies: Using eye tracking to inform cognitive film theory. In A. P. Shimamura (Ed.),Psychocinematics: Exploring Cognition at the Movies. New York: Oxford University Press.

Smith, T. J. (2012) The Attentional Theory of Cinematic Continuity, Projections: The Journal for Movies and the Mind. 6(1), 1-27. (pdf)

 There have been very few empirical studies looking specifically at the influence of steadicam shots on gaze behaviour but one recent study by Wang and colleagues (http://www.journalofvision.org/content/12/1/16.full)  showed how powerful such shots could be for creating similarity in gaze across viewers. Using long steadicam clips from Russian Ark and Children of Men, the authors showed that introducing cuts and scrambling the order of frames within the steadicam sequences disrupted gaze behaviour but the control of each shot over viewer attention was so strong that viewers were able to very quickly reorient to the disordered sequences and re-attend to the centre of interest. This study shows that by using a carefully choreographed steadicam shot, the director can give the viewer the illusion of freedom to roam a continuous shot whilst actually constraining where they look when, creating continuity of attention within the frame and across viewers.


Tuesday, October 16, 2012

Guest post: Camera Views of Candidates’ Debates Could Play Key Role in Winning Style


This week our blog hosts a guest post from Lester Loschky (Cognitive Psychologist) on the recent US presidential and VP debates and how subtle directorial decisions may impact our impressions of the candidates. (Tim J. Smith)

------------- 
Camera Views of Candidates’ Debates Could Play Key Role in Winning Style
by Lester Loschky

I will make a claim that many people may find counter-intuitive: The camera views of the US Presidential candidates in their debates could prove important in determining who “wins” those debates.  But before you close your browser window on this seemingly crazy idea, read on, and see if you don’t find it more persuasive.  There is a lot of research, and a lot of punditry that backs it up.

Those following the current US Presidential election campaign know that the impact of the Presidential debates has assumed a greater importance than any in recent memory.  President Obama’s poor performance relative to Governor Mitt Romney in their first debate apparently led to his losing a commanding 5 point lead in the general election polls in the period of a week

In addition, most of the commentary on that debate has shown that it was particularly the “style” of each of the candidates that was particularly important.  The importance of style is consistent with what has been said about other important US Presidential debates of the past.  For example Richard Nixon’s sweating and five o’clock shadow compared to JFK’s cool demeanor in the 1960 debates, and Al Gore’s superior seeming sighs compared to W’s folksy manner, have both been credited with influencing the outcomes of their respective elections. 

I would like to point to one particular point of style that was very apparent in the first Obama/Romney debate—namely eye contact with the camera.  Howard Kurtz noted “stylistically, Romney came on strong, showing a confident command of facts and figures even as he tried to moderate or distance himself from some of his proposals. He also made direct eye contact with the camera while Obama often seemed to be looking down [emphasis added], never adjusting his intensity and acting like he was at a garden-variety news conference” Howard Kurtz, Oct 3, 2012 10:35 PM EDT).  Thus, Obama’s lack of eye contact with the camera during his debate may have been a factor is his losing of the debate.

Importantly, this issue also came up in the Vice Presidential debate between Vice President Biden and Congressman Ryan.  However, in that debate, it can be argued that it was due to the ABC Debate Director’s decision as to which views of each candidate to show to the TV audience.  Specifically, the camera views in the Biden versus Ryan debate Closing Statements favored Ryan.  Biden was shown looking at the wrong camera for the entire 1:19 of his final remarks but Ryan was show looking at the right one (see the Youtube embedded video clip below and a couple of screen captures from the clip). 




This is odd, because there is a red light on the camera you are supposed to look at, and Biden must know this very well.  So why was Biden looking at the wrong camera?  It seems implausible that Biden could not see the red light on the camera he was supposed to look at, or that he intentionally looked at the wrong camera, or he chose to address his comments to the chair of the debate.  Most importantly, the ABC Debate Director in the control room was the person who ultimately chose how Biden was presented to the national TV audience.  If it was argued to have been due to a lapse of attention by the Director, and the person below the Director who was in charge of pressing the button that selects the camera view to show the TV audience, then it was an extremely long lapse of attention, since the camera shot on Biden lasted for 79 seconds (i.e., 1:19), at the single most important (final) portion of the debate.  However, we can assume that the Director of the debate in the control room was a consummate professional, since s/he was chosen as Director for this very high stakes debate.  Thus, we can also assume that it was not a simple mistake due to a lapse of attention.  This means that it had to have been a conscious decision.  If so, it is a big problem.

Specifically, research has shown that failure to make eye contact reduces the likeability of a person (Mason, Tatkow et al. 2005), and makes a speaker less persuasive (Yokoyama & Daibo, 2012).  Thus, the Debate Director's choice of camera view for Biden's closing statement made him less likeable and persuasive (he wouldn't look you in the eyes), and made Ryan more likeable and persuasive (he looked you in the eyes).  Again, assuming this was not a simple mistake, for the reasons give above, put it into the realm of a “plausibly deniable” political “dirty trick” of the sort that Richard Nixon’s staff was famous for in the Watergate scandal.

Of course, one could argue that the camera view choice was a small thing, for only 1:19 of the Vice Presidential debate, which common wisdom says will not change the course of an election.  The counter argument to that is that the V.P. debate was argued to be critical in determining the momentum of the Presidential election campaign, and that the Closing Statement is the last thing that viewers see in the debate, and should therefore be most memorable.  This is based on the extremely well-known phenomenon of the “recency effect” which research has shown also affects long-term memory for things such as memory for US presidents (e.g., name all the US presidents you can remember in reverse chronological memory—most people’s memory is best for the most recent Presidents)(Roediger & Crowder, 1976). 

More importantly, what if the same "mistake" happens tonight in President Obama's or Governor Romney’s closing statement?  These simple directorial decisions may impact our perception of each candidate in subtle ways that cumulatively effect our overall confidence in them and their politics.


Lester Loschky
Cognitive Psychologist