Sonically Induced Fear: Experimenting With Non-linear Binaural Audio Environments
This experiment was done as part of the Experimental research in the human-computer interaction course. It has been supervised by professor Veikko Surakka and Jani Lylykangas. Below you can find the full scientific report of the experiment.
ABSTRACT
This study utilizes nonlinear binaural sounds in 3 immersive audio-only environments to determine which sounds can incite the most fear, discomfort, stress, and arousal in humans. For this purpose, the researchers have produced different audio stimuli to investigate the intensity of the aforementioned emotions. 18 participants took part in the study, 9 of which were conducted in a lab environment. The experiment was a within-subjects design with sound as a factor and utilized a specially developed native android mobile app for playing the audio stimuli and collecting participants’ responses. Non-parametric analysis methods were used; the Friedman test was used for statistical analysis. Bonferroni-corrected Wilcoxon signed-rank tests were used for post hoc tests. Statistical analysis showed that the third environment, utilizing humans in distress sounds, induced the most fear, discomfort, and stress in comparison to the other two environments. There was no significant difference between other environments.
INTRODUCTION
This study utilized nonlinear binaural sounds in 3 immersive audio-only environments to determine which sounds can incite the most fear in humans. The experiment also aimed to test the intensity of other negative emotions that may accompany fear, namely discomfort, stress, and arousal. This work can provide insights for sound designers, game developers, filmmakers, practitioners, and researchers in the fields of immersive technologies aiming to evoke or even avoid various levels of distress and negative emotions as part of their developed experiences.
Non-Linear sounds, defined as sounds produced outside the normal range of an instrument or voice, such as human/animal screams (Tiwari, 2015), are a staple of the horror genre and have always been used to create an unsettling atmosphere in various forms of media (film, theatre, video games, sound design, etc.). According to Yan et al. (2019), antipredator responses in mammals and birds can be evoked through non-linear sounds, which explains their constant use in horror films. Reymore (2018) argued that certain acoustic cues can activate defensive circuits in humans, which create tension in the body and even result in fear by providing information on the real and present threats.
In former studies, researchers have experimented with unimodal stimuli (visual stimuli) in comparison with bimodal stimuli (simultaneous audio-visual stimuli) to test fear levels (Wheeler et al., 2014; Bowie, 2016; Chełkowska-Zacharewicz & Paliga, 2019). Blumstein et al. (2012) manipulated neutral stimuli by adding noise, simulated distortion, and sudden frequency up and downshifts to increase levels of arousal and of negative valence. Toprac & Abdel-Meguid (2010) studied sound effects in game design that can promote fear, suspense, and anxiety in players and tap into game immersion and resulting emotions. However, little attention is given to binaural sounds and immersive audio-only environments, specifically in experiments utilizing unimodal stimuli (audio only) to incite fear, discomfort, stress, and arousal in humans.
Binaural audio refers to audio that would be heard exactly the way it is in real life and is recorded in such a way that the brain can track the spatial origin of the sound (Lalwani, 2015). With the rise of extended reality experiences and the accelerated shift to the digital world, sound has proven to be of utmost importance in increasing believability and immersion in virtual environments (High Fidelity, 2021). According to Hancock (2018), Although binaural sound offers a unique acoustic virtual reality and harnesses sound’s properties of invisibility, intimacy, and invasiveness, it has largely been overlooked in the entertainment industry.
Although non-linear sounds are notorious for their reputation for inducing negative emotions and binaural audio can offer an immersive form of horror experiences, there is a clear gap in the literature regarding the several types of non-linear audio, the context of the recordings, the use of binaural audio and immersive audio-only environments, and a clear comparison of which sounds can induce more fear, anxiety, discomfort, and stress than others.
METHODS
Design
This experiment tested the ability of sound-only environments in evoking fear, discomfort, stress, and arousal in humans. The independent variable is sound, which is manipulated to create 3 test conditions (levels) in the form of audio-only environments. The first level presents unsettling sounds in empty environments. The second level presents unfamiliar non-human sounds in a semi-occupied environment. The last level presents the sound of humans in distress in an occupied environment. The dependent variable is the level of fear, along with accompanying negative emotions, namely discomfort, stress, and arousal.
For measuring the dependent variable, self-reporting surveys utilizing the Subjective Units of Distress Scale (SUDS) were used. SUDS is a scale ranging from 0 to 10, measuring the subjective intensity of distress experienced by an individual (Benjamin, 2010). To avoid the order effect, counterbalancing was used. The audio environments were arranged in three scenarios according to a Latin Square (Richardson, 2018). Participants received different ID numbers that determined which scenario would be played, for instance, ID 1 got order A-B-C, ID 2 got order B-C-A, and ID 3 got order C-A-B.
Participants
In total there were 18 participants in the study, out of which 66.7% were female and 33.3% male. ‘Other’ and ‘I’ prefer not to answer were also possible answers but none of the participants identified with those. The participants’ age ranged from 17 to 62. All participants reported normal or corrected-to-normal hearing. The participants were HTI.350 course students, or colleagues, friends, and relatives of them. 50% of the participants took part in the study at Tampere University in a lab environment, while the rest of the experiments were conducted in similar environments as the laboratory; a dark, quiet environment, utilizing similar equipment while keeping their eyes closed. The participants were asked to sign the “Informed consent form” followed by the ‘Background information form’ attached in Appendix 2 prior to participating in the study. The background information form included questions about the participant’s age, gender, and hearing (normal or corrected-to-normal), and whether they had ear diseases or disorders. The aim was to have people with the same hearing abilities participating in the study.
Apparatus
The study was conducted through a native Android mobile application called “The Sonic Fear App”, developed for Android mobile devices by one of the researchers. To conduct the experiment, an Android mobile device was presented to the participants, and earphones were provided to play the stimuli from the three audio environments through the app. The app provided instructions for playing the sound samples and after that displayed the rating scales, which the participants answered using the mobile device’s touch screen. The application collected all the participants’ responses and stored them in an excel sheet for simplified processing of data.
Stimuli
The experiment utilized six audio stimuli, two in each audio environment. The audio stimuli were designed with two main layers: 1) the background, which was constant across all environments and consisted of non-linear sounds and atmospheric synths commonly used in horror genres. The background sound design relied on sound level variability, high pitch levels and pitch variability, asynchrony, and tempo irregularity. 2) the foreground, which differed between the environments according to the environment’s theme, utilized binaural recordings to immerse the participant. The aforementioned layers were designed, produced, and mixed in Steinberg Cubase (a digital audio workstation), utilizing plug-ins such as DearVR and VST AmbiDecoder along with an extensive audio library for Native Instruments Kontakt 6 used for producing horror scores shown in Figure [1]. Increasing tempo and pitch, irregularities in frequency, intensity, and duration (instability), roughness, loudness, dissonance, asynchrony, and unpredictability were the main properties taken into consideration for producing the stimuli.
Figure 1: Cubase interface showing the utilized audio library, DearVR plugin and Audio Mixing panel for producing the stimuli.
The first environment utilized rough sounds of varying frequencies, sounds of explosions, and impact (e.g., slamming doors, windows, breaking furniture, heavy rain, thunder), creating the atmosphere of an environmental catastrophe in the first stimuli and a nuclear war in the second stimuli. In this scenario, the unsettling environment was the main audio source.
The second environment relied on high pitch sounds, sudden loudness, the presence of unfamiliar creatures surrounding the listener, fast tempo, and rising pitch contour. In this scenario, the increased level of unfamiliarity and the paranormal were the main audio source.
The third environment utilized binaural audio samples of humans in distress. Both stimuli in this environment included sounds of humans screaming and crying in pain, fear, or sadness. The stimuli were fast-paced, and overwhelming, and relied on pitch variability, asynchrony, and sound proximity. In this scenario, distressed humans were the main audio source.
Experimental task
The participants were instructed to keep their eyes closed for the duration of the actively incoming sound. They were tasked to listen to the audio stimuli, after which they were given the on-screen instructions to fill in the four in-app self-report surveys following each stimulus. The self-report surveys reported the levels of fear, discomfort, stress, and arousal experienced by the participants.
This resulted in 24 rating scales following the stimuli, in addition to a final in-app evaluation with 12 rating scales (four for each audio environment; to provide an overall picture of the perceived levels of fear, discomfort, stress, and arousal post task) (Figure [2]).
Figure 2: In-app screenshots of the self-report surveys (rating scales) presented after each stimulus.
Procedure
When the participant first arrived in the laboratory, they were instructed to read and sign the relevant documents attached in Appendix 2. They were also instructed to turn their mobile phones off and were provided with masks and single-use earphone covers. After this, they were instructed to sit in front of a table, which had a mobile device running the app placed on top of it. The moderator then explained how the app was going to work - using the ‘Instructions’ document as a base - attached in Appendix 2. After this, the moderator dimmed down the lights, the participant put on the earphones, and was instructed to close their eyes whilst listening to the audio samples. First, they went through a practice trial in the app which explained the process of listening to the sound samples and answering the self-report surveys included in the app. Following this, the moderator came to the room to inquire about the audio level, and whether it was too high or appropriate. Once the audio level was determined to be appropriate, the participant proceeded with the experimental trial by listening to the first sound sample from the first environment. During the experiment, the researchers and moderator sat in the adjacent room and were silently observing and documenting any visible movements, facial expressions, or sounds arising from the participant through the separating window between the two rooms. After each sound sample, the participant opened their eyes to answer questions about the stimulus. This was repeated for all three audio environments. At the end of the experiment, the participant answered a final survey regarding their perception of the three audio environments overall. After the experimental tasks were finished, the researchers inquired about any final thoughts or comments from the participants. After the more casual toned commentary round, the researchers thanked the participants for their time and wished them farewell. Each sound sample lasted 67 seconds and in total, conducting the whole experiment took approximately 15-20 minutes.
Data analysis
The experiment was a within-subjects design with sound as a factor. The Friedman test was used for statistical analysis. Bonferroni-corrected Wilcoxon signed-rank tests were used for post hoc tests when needed.
RESULTS
In testing the effect of sound on fear levels in participants, Friedman's test showed that there was a statistically significant effect of sound, Χ²(2) = 9.1, p < .05.
Post hoc pairwise comparisons with Wilcoxon signed-rank tests conducted using Bonferroni adjusted alpha levels of .017 per test (.05/3), showed that audio environment 3 caused significantly more fear than audio environment 1, Z = -2.90, p < .05. Other pairwise comparisons were not statistically significant. Figure [3] showcases the mean fear level and S.E.M.s for the three audio environments.
Figure 3: Mean fear level and S.E.M.s for the three audio environments.
In testing the effect of sound on discomfort levels in participants, Friedman's test showed that there was a statistically significant effect of sound, Χ²(2) = 17.7, p < .05. Post hoc pairwise comparisons with Wilcoxon signed-rank tests conducted using Bonferroni adjusted alpha levels of .017 per test (.05/3) showed that audio environment 3 caused significantly more discomfort than audio environment 1, Z = -3.30, p < .05. and significantly more discomfort than audio environment 2, Z = -2.80, p < .05. Other pairwise comparisons were not statistically significant. Figure [4] showcases the mean discomfort level and S.E.M.s for the three audio environments.
Figure 4: Mean discomfort level and S.E.M.s for the three audio environments.
In testing the effect of sound on stress levels in participants, Friedman's test showed that there was a statistically significant effect of sound, Χ²(2) = 8.9, p < .05.
Post hoc pairwise comparisons with Wilcoxon signed-rank tests conducted using Bonferroni adjusted alpha levels of .017 per test (.05/3) showed that audio environment 3 caused significantly more stress than audio environment 1, Z = -2.60, p < .05. and significantly more stress than audio environment 2, Z = -2.70, p < .05. Other pairwise comparisons were not statistically significant. Figure [5] showcases the mean stress level and S.E.M.s for the three audio environments.
Figure 5: Mean stress level and S.E.M.s for the three audio environments.
In testing the effect of sound on arousal levels in participants, Friedman's test showed that there was a statistically significant effect of sound, Χ²(2) = 9.0, p < .05.
Post hoc pairwise comparisons with Wilcoxon signed-rank tests conducted using Bonferroni adjusted alpha levels of .017 per test (.05/3) were not statistically significant. Figure [6] showcases the mean arousal level and S.E.M.s for the three audio environments.
Figure 6: Mean arousal level and S.E.M.s for the three audio environments.
DISCUSSION
With a glance at the charts and visualized data, it is noticeable that the third audio environment has evoked higher levels of fear, discomfort, stress, and arousal than environments 2 and 1. Moreover, environment 1 evoked the least amount of the negative emotions discussed above. It can be concluded that humans, because of relatedness, sense more danger when they hear another human in distress. Moreover, the sounds presented in the third environment are quickly recognizable and not unlikely to face in real-life situations. Environment 3 confirms previous studies on non-linear audio, mentioning how sounds, which are associated with threat, aggression, and danger, often mimic vocal expressions caused by physiological changes resulting from fear (Reymore, 2018; Yan et al., 2019).
The utilization of binaural audio aimed to compensate for the lack of accompanying visuals by creating immersive audio-only environments where the spatial origin of the sound can be tracked. Post experiment interviews with the participants confirmed that binaural audio allowed them to imagine the space they were in and visualize the audio sources. However, older participants noted that the provision of accompanying visuals could reduce the required cognitive load of analyzing what the sound is, and the time required to understand the environment, making the sounds scarier.
While conducting the experiments in a lab environment, the researchers observed the participants and took notes of their reactions and overall demeanor. Some participants showed signs of discomfort, constantly adjusting posture, lifting shoulders, and moving their heads backward and away from the sound source. Such body movements increased when high-pitched sounds were present and especially in stimuli that used sudden loudness, unpredictability, and increased audio proximity (e.g., human whispers) in the second and third environments.
The experiment was conducted using both participants in a lab environment and remote participants. Participants who took part in the experiment in a lab environment showed similar results to those who took the experiment outside the controlled environment, which indicates a good level of coherence between the experiment’s internal and external validity. Moreover, the study sample varied in age, ranging between 17-62 years, and cultural background. The results were consistent throughout the experiment, which is a good indicator of the generalizability of the findings. The experiment can be conducted in other quiet and dark environments mimicking a lab setting yet not as controlled. Remote participants used the developed Android application to take the test. However, this could be affected by the accessibility to suitable hardware such as headphones that support spatial sound, which can greatly affect the test results.
For future studies, we aim to expand the scope of our experiment and publish our android application to make it accessible to a wider audience to expand our data pool. Several factors including cultural background, digital literacy, familiarity with the horror genre, accessibility to appropriate hardware, etc. that may affect the test result should be taken into consideration.