· 8 years ago · May 09, 2018, 07:56 PM
1%%%%%
2%%
3%% Sample document ``thesis.tex''
4%%
5%% Version: v0.2
6%% Authors: Jean Martina, Rok Strnisa, Matej Urbas
7%% Date: 30/07/2008
8%%
9%% Copyright (c) 2008-2011, Rok Strniša, Jean Martina, Matej Urbas
10%% License: Simplified BSD License
11%% License file: ./License
12%% Original License URL: http://www.freebsd.org/copyright/freebsd-license.html
13%%%%%
14
15% Available documentclass options:
16%
17% <all `report` document class options, e.g.: `a5paper`>
18% withindex - enables the index. New index entries can be added through `\index{my entry}`
19% glossary - enables the glossary.
20% techreport - typesets the thesis in the technical report format.
21% firstyr - formats the document as a first-year report.
22% times - uses the `Times` font.
23% backrefs - add back references in the Bibliography section
24%
25% For more info see `README.md`
26\documentclass[withindex,glossary]{cam-thesis}
27
28% Citations using numbers
29\usepackage[numbers]{natbib}
30\usepackage{algorithm}
31\usepackage[noend]{algpseudocode}
32
33%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
34%% Thesis meta-information
35%%
36
37%% The title of the thesis:
38\title{Earables\\
39Wearable Computing for your Ear}
40
41%% The full name of the author (e.g.: James Smith):
42\author{Alexander Gubbay}
43
44%% College affiliation:
45\college{Darwin College}
46
47\collegeshield{CollegeShields/Darwin}
48
49
50%% Submission date [optional]:
51% \submissiondate{November, 2042}
52
53%% You can redefine the submission notice [optional]:
54\submissionnotice{This dissertation is submitted for the degree of MPhil in Advanced Computer Science}
55
56%% Declaration date:
57\date{June, 2018}
58
59%% PDF meta-info:
60\subjectline{Computer Science}
61\keywords{one two three}
62
63
64
65%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
66%% Abstract:
67%%
68\abstract{%
69This dissertation focuses on a new type of wearable device becoming increasingly popular - headphones equipped with sensor technology. Microphones are increasingly found embedded in the bud of the headphone itself, moving away from the traditional placement on the connection cable, due to advancements in cable-free Bluetooth headphones such as Apple AirPods and Google Pixel Buds, both of which contain microphones - currently primarily for voice control and phone calls. This paper will show how devices similar to these can be used to generate richer functionality. Two main examples of this are given - the first demonstrating how it may be possible to use sonic signals to detect when the user is wearing the device in their ear. The second shows how by sending an acoustic signal into the ear canal of the wearer, a response can be captured on the microphone that is sufficiently unique so as to allow biometric identification or verification of the wearer. The functionality of this system from a hardware and algorithmic perspective is presented in detail. The performance of the system is inspected through a user study carried out with prototype hardware. The results of this experiment are evaluated, showing greater than 99\% accuracy at identification, and a false accept rate of less than 2\% for verification. Good performance in challenging acoustic environments is also demonstrated. The changeability of a wearer's ear canal acoustic fingerprint is also discussed, motivated by a user study over a substantial time period, which reveals that the fingerprint generated is stable over long periods of time, potentially multiple months. The system's applicability and value added to a user explained, and a comparison to existing biometric techniques given. Finally, the work is concluded with a discussion on avenues for further work.
70}
71
72
73
74%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
75%% Acknowledgements:
76%%
77\acknowledgements{%
78 My acknowledgements.
79}
80
81
82
83%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
84%% Glossary [optional]:
85%%
86\newglossaryentry{HOL}{
87 name=HOL,
88 description={Higher-order logic}
89}
90
91
92
93%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
94%% Contents:
95%%
96\begin{document}
97
98%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
99%% Title page, abstract, declaration etc.:
100%% - the title page (is automatically omitted in the technical report mode).
101\frontmatter{}
102
103
104
105%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
106%% Thesis body:
107%%
108\chapter{Introduction}
109
110Fundamentally, headphones are devices that work with sound and advances in hardware technology is enabling manufacturers to embed these devices with sensor technology. Examples of this include Apple\footnote{https://www.apple.com/uk/shop/product/MMEF2/airpods}, Google\footnote{https://store.google.com/product/google\_pixel\_buds}, and Samsung\footnote{https://www.samsung.com/us/audio/headphones/in-ear/gear-iconx-black-sm-r150nzkaxar/}. The unique positioning of these sensors on the wearer's head gives an opportunity to collect data from a point on the body close to the point of perception for the user. Given these factors, it is a natural choice to consider equipping these new generations of sensor enabled headphones with microphones. They are both near the point of sound sensing for the wearer, and the combination of a speaker-microphone pair presents an opportunity employ active sensing techniques.
111
112As no prior art seeking to specifically characterise the data retrievable from such a device could be found, this will be presented here driven by the exploration into specific, potentially useful applications of this new data source.
113
114This report will introduce a prototype hardware device with microphones embedded in the inside and outside of a pair of headphones. This puts these sensors in a unique position relative to the wearer, for capturing data from within the ear canal itself, and also from the environment around the wearer as detected at the ear. Figure \ref{config} shows a diagrammatic example of the microphone placement on the prototype device. Although the actual form factor differs from the diagram, the positioning of the microphones and loudspeaker is identical. An initial investigation was carried out in order to gain an understanding of the broad kinds of data that can be captured from this device. Details of this are given. The results of this exploration have driven work into specific applications of the technology. As headphones are ubiquitous pieces of consumer technology, an effort is taken to ensure the real world usability of the tools presented here. This means that it is crucial to evaluate the systems against their resilience to challenging everyday environments, and also the applicability to as wide a range of people of people as possible. The value that the work presented here could add to a consumer/user of the hardware is discussed and used as a motivator for the research direction. With this in mind, several key areas of work are presented, and the performance of systems built from them has been quantitatively evaluated with a user study.
115
116\begin{figure}[h]
117 \centering
118 \includegraphics[width=0.6\linewidth]{Figures/headphone-configuration.png}
119 \caption{Diagrammatic example of headphone microphone positioning. Public domain image.}
120 \label{config}
121\end{figure}
122
123One potential application of this hardware is a biometric security system for mobile devices, by taking an acoustic fingerprint of the ear for either identification or verification. The lack of a need for advanced hardware to enable this (only a headphone with a simple MEMs microphone is required) could make it cost effective in comparison to specialised biometric hardware such as fingerprint scanners and facial recognition systems. For example, the iPhone X's facial recognition hardware is estimated to cost \$16.70 to produce\cite{AndrewRassweilerWayneLamJeremieBouchaud2017}. By comparison, a high quality MEMs microphone can be purchased for a few pennies\footnote{https://www.mouser.co.uk/Sensors/Audio-Sensors/MEMS-Microphones/\_/N-98yda?Ns=Pricing\%7C0}. This would represent a significant saving on the bill of materials for any manufacturer including biometric systems in their hardware.
124
125The ability to identify the wearer of headphones extends beyond security applications. For example, a positive identification match upon the user putting the headphones on could trigger personalisation on the specific system the headphones are connected to. It may load custom music playback settings such as queuing the last opened playlist, setting the last volume used, and loading the user's preferred EQ configuration. If the headphones are being used for phone calls, for example in an office, it could trigger an update of the user's out of office status and set the correct caller ID on the phone.
126
127A further application could be to use the wearer's acoustic fingerprint as a username/security token when performing identification over the phone. This system could form one of the three factors of authentication - something you have, something you know and something you are. Proving something you are is not a simple problem over the phone, given the lack of proximity between the two parties and the generally low bandwidth connection between them. Voice identification systems are in use for authentication in commercial systems\footnote{https://www.hsbc.co.uk/1/2/voice-id}. However recent advances in voice forgery techniques present a potential vulnerability for this system \cite{Galou2011}. We expose our voice to almost everyone we interact with, and it has been shown that strong attacks on voice identification systems can be carried out from voice forgeries generated with just a few minutes of recorded audio \cite{Mukhopadhyay}. An acoustic fingerprint of the ear canal is far less vulnerable to accidental leakage in this way. A system using acoustic fingerprints as one of the three authentication factors would not require additional hardware for many users. For example every Apple iPhone since 2015 has been equipped with a microphone inside the earpiece, resulting in a loudspeaker-microphone pair in the device's earpiece being pointed into the user's ear while making a call. This means, worldwide, an estimated 322 million Apple iPhone devices contain this microphone speaker combination\footnote{http://fortune.com/2017/03/06/apple-iphone-use-worldwide/}. Traditional phone calls sample at 8kHz, creating and extremely small frequency range in which fingerprinting could occur. However, voice over LTE (VoLTE) enables wideband voice transmission for phone calls between supported handsets, that supports transmission of sounds at frequencies up to 20kHz \cite{Barcelona2012}, which would provide the sample rate and bit depth required to run the fingerprinting algorithm presented here over the phone.
128
129Therefore, the first goal was to develop a system that can reliably identify the wearer of the hardware, based off the hypothesis that a given individual's ear canal is sufficiently unique in anatomical morphology, that it has specific acoustic properties that can be measured and used for identification purposes. Specifically, it is proposed that the ear canal of an individual has an echo response to a high frequency chirp that contains features that can be extracted to create a hash or fingerprint. These fingerprints can be compared to one-another in various combinations to generate a system for identification with a high degree of accuracy and reliability. The algorithms built to test this, and an experimental analysis are presented in this paper. The performance, reliability and applicability of the system is discussed.
130
131The second focus of the work has been to leverage the microphones to detect when the headphones have been seated in the ear. This could have applications such as automatically pausing music when the headphones are removed or answering a phone call when they are inserted. Work has been undertaken to characterise the signals received from the device when the headphones are either in or out of the ear. A system for performing this task is presented.
132
133There is the potential for these technologies to be combined, for example detecting when the headphone has been inserted and automatically triggering identification at that point to unlock a paired device or only answer a phone call if authorised.
134
135In addition to these two key areas of work, research into the general types of data that can be retrieved from the device is given. For example, it has been found that for some individuals that the pulse can be detected.
136
137\chapter{Previous Work}
138\section{The Rise of Wearables}
139The field of wearable technology has seen meteoric growth, and is forecast to be worth \$25 billion by 2019, growing 64\% from 2015 \cite{CCSInsight}. One of the key features of these wearable devices is the sensors they are are equipped with. Accelerometers, gyroscopes and heart rate monitors are common, along with WiFi and Bluetooth radios. These sensor rich devices are in a unique position of being physically attached to the user for large parts of the day, and in some cases at night too \cite{Smith2015}. This gives the opportunity for large amounts of data to be collected on a continuing basis. Naturally, these devices and the data they produce have been the focus of large amounts of attention from the academic community. The huge volumes of data produced by these devices as part of the internet of things creates new frontiers in data management, and in particular how to deliver information to the wearer without overloading them with useless metrics \cite{Swan2012}. Their role in improving quality of life through the use of wearable data in not just monitoring for health related metrics \cite{Conroy2014} \cite{Pantelopoulos2010}, but also by driving health care decision making has also been well explored \cite{Patel2015} \cite{Swan2012a}, along with the use of wearable data for other rich applications such as indoor positioning \cite{Gu2009}. The number of internet connected devices is expected reach 50 billion by 2020\footnote{https://blogs.cisco.com/diversity/the-internet-of-things-infographic}, and wearables will form a significant part of this - 310 million were sold in 2017\footnote{https://techcrunch.com/2017/08/24/global-wearables-market-to-grow-17-in-2017-310m-devices-sold-30-5bn-revenue-gartner/}.
140
141However, it has been shown that privacy and security issues presented by these devices remain a key concern for users \cite{Motti2015} \cite{Piwek2016}, and work has been carried out to explore the use of biometric authentication techniques with these new devices. For example using accelerometer based gait analysis to identify the wearer \cite{Gafurov}, which functions by detecting characteristics in the three axes of motion created when the wearer walks. This requires the device containing the accelerometer to be attached to the hip of the user. Much like fingerprinting for biometrics, it requires the user to be enrolled by capturing samples of their normal gait, against which a live sample can be compared. Other biometrics popular in mobile devices include fingerprinting in the classic sense, in which a feature vector is generated from a scan or photograph of the user's fingerprint. The live fingerprint taken at identification time is then compared to the previously captured enrolment samples for similarity. These systems are presented in terms of classification performance for identification, and equal error rates for verification enabling direct performance comparison to the system built here. Similar and in some cases better performance is shown in testing, along with greater usability than a device attached to the hip or a system requiring a user hold their finger on a scanner to take a fingerprint. This would be particularly true where acoustic fingerprinting is triggered automatically upon detecting the user inserting the headphone.
142
143Systems for identifying a user by a combination of biometrics was explored. The Mobile Biometrics Project \cite{Phil2012} aimed to fuse voice and face recognition into a system that could robustly identify users on mobile hardware by leveraging two ubiquitous sensors on smartphones - cameras and microphones. The facial recognition system would attempt to locate a face in an image, normalise it to account for lighting conditions and the relative angle of the face, and then finally extract a feature vector from it. The voice recognition functioned by representing the sound sample with cspectral analysis to decompose the frequency components and represent them on a mel scale. It will be shown that the system presented here outperforms either of these biometrics when they function in isolation, and has improved performance compared to their combined output.
144
145
146\section{In Ear Microphones}
147Several patents have been filed for in ear microphone hardware, dating as early as 1992 \cite{Davis1992} \cite{kruger1997} \cite{Jarle2000}. However, these devices do not contain a headphone driver to produce sound, and crucially do not contain an outward facing second microphone. This limits their relevance to the proposed work here. These devices are focused on the detection of speech through the vibration of the bony structure of the ear canal. Devices with multiple microphones designed to detect this resonance have been created, specifically designed to detect the vibration caused by speech conducted though the bones of the skull. A patent filed by Jawbone Acquisition LLC \cite{GregoryBurnett2005} shows how multiple microphones, including one mounted to make physical contact with the bony structures behind the wearers ear, can be used in combination with external microphones to create noise cancelled speech detection. The key difference between all of these devices and the prototype here is that they have all been designed to specifically to detect the conducted vibration from speech. This is not the case for our device, which has been placed to allow for the general acoustic environment in the outer ear canal to be measured, along with external sound.
148
149The bluetooth earpiece RippleBuds\footnote{http://ripplebuds.com} have a similar design to the prototype used here, having a microphone on the inside of the bud that is placed inside the ear to better pick up speech for phone calls in noisy environments. This device also has the ability to create sound from the earbud like a conventional headphone. This device however does not have a microphone facing externally like the prototype used here.
150%FIX THIS
151A significant body of work discussing the characteristics and behaviour of such a system could not be found. Research also did not reveal any evidence of current or previous investigations into using the changes in the sound profile measured by the device to detect when it has been inserted into a user's ear. Microphones in the ear canal have been previously used to measure the effect of the ear structure on hearing aid performance \cite{EarlR.Harford1980}, but this used a separate microphone inserted deep into the ear canal.
152
153Mainly, in ear microphones have been used to indirectly measure some other metric form the user. An example of this is a study that utilised in ear microphones to determine the user's eating habits though the measurement of chewing sounds transmitted through the jaw \cite{Nishimura2008}. In ear microphones have also been used to develop better speech recognition for control of military robots in noisy operational situations \cite{Kurcan2006}.
154
155\section{Detecting Biometrics Through the Ear}
156
157Previous work has shown that pressure changes in the ear can be used to measure the heart rate of the wearer \cite{MitsubishiChemicalHoldingsCorp2013}. This technology is proprietary so the signal processing and exact hardware employed is unknown. The system appears to consider the change in pressure an impulse event, rather than detecting the properties of the sound emitted from the blood moving through the blood vessels around the ear.
158Several consumer products exist for detecting a user's pulse through the ear.\footnote{https://www.bose.co.uk/en\_gb/products/headphones/earphones/soundsport-wireless-pulse.html} \footnote{https://www.jabra.co.uk/sports-headphones/jabra-sport-pulse-wireless} \footnote{https://www.jbl.com/wireless-headphones/UA+SPORT+WIRELESS+HEART+RATE.html} \footnote{http://www.samsung.com/uk/wearables/gear-iconx-r140/SM-R140NZKABTU/}
159Specific information on how these devices function is limited, but they all appear to employ light based techniques to perform photoplethysmography - detection of blood flow through the tissue of the ear through changes in its colour. This results in a bulky additional sensor that all of the devices have attached to the ear bud to house the light source and sensor. This also has an associated energy cost. It is feasible in this case that passively detecting the heart rate from a microphone is a solution that may use less power and be less bulky in the limited space available.
160
161The technology firm NEC has released a short summary demonstrating their research efforts into biometrics through sound echoes in the ear\footnote{https://www.nec.com/en/press/201603/global\_20160307\_01.html}. Specific information on the testing done or the definitions of accuracy used are not provided, and nor are any specific details of the technology that would make it replicable. There are several different characteristics of this work that differentiate it from that presented here. First, the signals are delivered much more slowly - over 100s of milliseconds. The chirp used for identification here is less than 10ms. Next, it appears that the system operates a lower frequency, from 200Hz to 9.9kHz. This is well within the audible hearing range of humans. The system here uses much higher frequency signals that approach ultrasound thus making them less audible. Also it appears from the supplied diagrams that the system only extracts peaks in the signal response, whereas the method used here attempts to characterise the shape of the entire signal. Given the lack of information provided, it is not possible to do a direct performance comparison. A patent for acoustic feature detection of the human ear has also been filed, but this uses a much higher sample rate of 200kHz, which is of limited use for the application of consumer grade hardware sought here. For example, the Apple iPhone's microphone has a maximum sample rate of 44.1kHz\footnote{https://developer.apple.com/library/content/documentation/Audio/Conceptual/AudioSessionProgrammingGuide/OptimizingForDeviceHardware/OptimizingForDeviceHardware.html}, several times less than this previous work leveraged. This limitation is common for most consumer grade devices.
162
163Measurements of the acoustic impedance features of the human tympanic membrane (ear drum) have been taken previously as this serves a key function in the estimation of sound pressure in the ear canal for various applications \cite{Hudde1983}. The small sample size of 6 individuals found this value to be detectably unique between these participants and repeatable over a time scale of several weeks.
164
165\chapter{Design and Implementation}
166\section{Introduction}
167This section discusses the work carried out in the project to leverage the prototype device to collect data from within and without the user's ear. The device contained two microphones per ear - one in the inside, in the cavity formed by inserting the device in the wearer's ear, and the other on the outside facing outwards from the wearer. The device took took the form factor of a standard pair of headphones, and was connected by a set of three 3.5mm TRS jack ended cables. The first enabled stereo audio playback, and the other two allowed for collection of audio recorded from the pair of microphones in each ear. The maximum rate at which this could be sampled using available hardware was 48kHz, meaning the theoretical maximum detectable frequency was around 24kHz - well into the ultrasonic range. Two USB sound cards were used to capture the audio data, with all 4 channels captured simultaneously. Figure \ref{experiment} shows a diagram of this set up. The audio from each microphone was saved as 16 bit PCM wave files. This simplified the processing of the data by saving each sample as a pair of integers. This was then parsed into MATLAB for processing, with the audio saved as an $n-by-m$ matrix where $n$ was the number of tracks recorded, and $m$ the number of samples.
168
169
170\begin{figure}[h]
171 \centering
172 \includegraphics[width=0.8\linewidth]{Figures/experiment-4.png}
173 \caption{Diagram showing experimental set up for data capture.}
174 \label{experiment}
175\end{figure}
176\section{Hardware}
177\subsection{Microphone Hardware Performance}
178
179The microphones in the headphone unit are microelectromechanical systems, similar to those found in many headsets used for telephone conversations. Each has a 2.2k$\Omega$ impedance at 500Hz. Each ear bud has a rubber flange which forms a seal with the wearer's ear canal when inserted.
180
181A test track was played with both headphones resting on a flat surface, with the entrance to the speaker cavity facing upwards. This test was used to establish a baseline for the sensitivity of the microphone speaker combination at different frequencies. Figure \ref{outside} shows representative example of this, plotting the frequency and volume of the detected signal from the chirp. The top graph shows the detected signal in the time domain from a chirp increasingly linearly in frequency. The bottom graph shows the same for the frequency domain. It is immediately clear that the microphone's response is not constant in the face of changing frequency. The level of the detected signal begins to drop steeply as the frequency approaches ultrasound, becoming virtually undetectable above 21kHz. This is likely intentional in the design, as for most consumer applications, ultrasound would simply unwanted noise. Perhaps unsurprisingly, the microphone's peak sensitivity is between 1kHz and 7kHz, frequencies common in music and speech. There is a drop in sensitivity, between 10kHz and 13kHz, with an increase noted between 13kHz and 18kHz. This was found to be the same for both the left and right internal microphones.
182
183\begin{figure}[h]
184 \centering
185 \includegraphics[width=0.6\linewidth]{Figures/outside.png}
186 \caption{Magnitude of response from in ear microphone over changing frequencies.}
187 \label{outside}
188\end{figure}
189
190The microphones are sensitive, and it is has been found that several biological processes can be detected. Any action involving the movement of the jaw such as eating or chewing are transmitted clearly, due to structure of the human jaw transmitting vibration into the tissue close to the position of the microphone. It has been found that bone is a particularly good conductor of sound\cite{Stenfelt2005}. This was expected and has been discussed in prior work. Speech is also transmitted clearly, due to the conductance of the bones in the skull that surround the ear canal. Again this was expected and has been leveraged in commercial products, as mentioned in the literature review. More interestingly, it was sometimes possible to observe the pulse of the user. By applying a bandpass filter, allowing only frequencies between 20Hz and 40Hz, it was found that regular peaking patterns could be found in the audio output trace, with a frequency and duration similar to that of a pulse. The repeatability of this phenomenon is investigated in this work. It is hypothesised that these peaks are caused by the changing pressure in the ear as a result of blood vessel expansion from the contraction of the heart. It was hoped that counting the resultant peaks over a given time that it would be possible to estimate heart rate, however the intensity of this signal was found to vary greatly between individuals.
191
192The sound of the wearer breathing is audible on the internally recorded track, even when the wearer is at rest and breathing lightly. If the signal could be isolated from the noise, the pattern of the wearers breathing could be established, and from this the breathing rate of the wearer could be determined and perhaps the intensity and duration of inhalation and exhalation. However, the signal is extremely broadband low in volume and occurs without well defined start and end points, making it extremely challenging to isolate effectively, particularly in the face of background noise.
193
194\section{Human Ear Acoustics}
195
196The structure of the ear is important to examine as it is a sensory organ that is specifically equipped to work with sound, and is adapted to gather and focus sound waves. This means that it has exploitable acoustic properties. The ear is a sensitive organ, and hearing damage from headphone use is a prevalent and serious problem \cite{Kim2009} so it is also important to take into account the structure and anatomy of the ear the when considering the comfort and safety of the user while conducting tests. Figure \ref{ear} shows the structure of the outer ear.
197
198\begin{figure}[h]
199 \centering
200 \includegraphics[width=0.6\linewidth]{Figures/earDiagram.png}
201 \caption{Structure of the outer ear. Public Domain Image \cite{Gray1918}}
202 \label{ear}
203\end{figure}
204
205The device sits in the entrance to the outer ear canal, with the external microphone housing resting on the ear lobe. The internal microphone and speaker driver sit approximately 7mm into the ear canal. A meta-analysis of ear canal morphology measurement studies found the average length of the ear canal in adults to be between 22.4 and 31.8mm, with men showing a slightly longer length than women, on average \cite{Ahmad2000}. The average diameter is 7mm, and the canal follows an "S" shaped path. Acoustic techniques have been developed for measuring the length of the ear canal \cite{Hiipakka2008}. These techniques employ high frequency sound to measure the echo time of a signal - much like radar.
206
207The range of frequencies an adult can hear is usually considered to range between 20Hz and 20kHz, although this decreases as the person ages \cite{Brant1990}. Sound is detected by the sound waves passing down the ear canal and impacting the tympanic membrane, which in turn causes small bones to vibrate in the middle ear. These bones pass the vibration into the cochlea, which is a curled structure that contains small hairs of varying length that move at different frequencies, creating a signal in the nervous system. The cochlea is effectively a transducer converting sound signals into electrical impulses. Importantly though, the cochlea is fluid filled, creating an interface of different viscosities though which the sound waves must pass - air and liquid. This difference in mechanical impedance results in a large proportion of the waves being reflected at this interface \cite{StevenSmith2011}.
208
209The walls of the ear canal are formed from external acoustic meatus, which has been shown to amplify sound between 3kHz and 12kHz \cite{Kemp1979}. Information on its behaviour with sounds with frequency above the limit of human hearing was not found.
210
211\chapter{Algorithm Design}
212
213\section{Headphone State Detection}
214It was hypothesised that it may be possible to detect when the user has the headphone inserted into their ear by detecting characteristics of the signals from the internal and external microphones. An initial investigation into the behaviour of the microphones when inserted versus removed revealed that there was very little difference between the signals received from the devices in each state. This made simply detecting the state by measuring some property of the signals that is different when inserted versus not inserted on a continuous basis unattractive - particularly in the presence of any kind of noise. Because of this the investigation focused on the transition of placing or removing the device from the user's ear. Two main methods were tried to enable exploitation of this transition, however the first method - using external ultrasonic signals to detect insertion - was quickly abandoned due to a lack of any supporting test results, and was therefore not tested in the user study. The performance and limitations of the second method are detailed in the results section.
215\subsection{Conductance of Ultrasound To Detect Insertion}
216It was hypothesised that it may be possible to use ultrasound from an external source, such as a smart watch or paired smart-phone that often accompany headphones to detect when the user had inserted the device.
217Exploratory tests were conducted using speaker emitting ultrasound placed against the skin of the wearer of the headphones. The purpose was to evaluate if it is possible to detect when the user had inserted the headphone into their ear through an increase in detected sound by the internal microphone, due to the ultrasonic signal being conducted through the wearer's body. However, it was found that there was no significant difference in the volume detected by the internal microphone between when the device was inserted into the user's ear and not. No evidence of any meaningful component of the signal being conducted through the body to the user's ear as opposed to escaping from the contact point with the user's skin could be found. The volume of the output signal being generated was verified through the use of a spectral analysis tool, which confirmed the source to be emitting at approximately 59dB during the tests. The test was conducted with the speaker placed at several points on the user's body, including their chest, both wrists and shoulders. Given the lack of promise exhibited in initial tests for this technique, it was not pursued further and alternative approaches were explored.
218
219\subsection{Characterisation of the Insertion Signal}
220The first step of the investigation into plausible methods for detecting insertion and removal was to understand what repeatable characteristics could be detected. The algorithm described here leveraged this to detect the transition of the user inserting or removing the device from their ear. The signal was informally characterised as follows:
221\begin{enumerate}
222\item The signal has a high volume, peaking the volume scale at maximum magnitude for sometimes several hundred samples at a time, giving it a high RMS.
223\item The signal is highly variable, and is non constant in volume, showing sharp peaks and many sudden changes in magnitude, giving it a high standard deviation over time.
224\item The signal is very broadband, not favouring any particular frequency, but is mainly contained of lower frequency sounds. This gives it a specific frequency power distribution that forms a logarithmic curve.
225\end{enumerate}
226Figure \ref{raw} below shows a volume-time representation of a recorded test sample. The large "noisy" sections at either end of the recording is the signal generated from the insertion and removal of the headphones from the ear of the participant, the signal sought to be characterised.
227\begin{figure}[h]
228 \centering
229 \includegraphics[width=0.7\linewidth]{Figures/signal-raw.png}
230 \caption{Raw time volume graph of a test sample recording.}
231 \label{raw}
232\end{figure}
233
234Given these, three parameters arise to characterise the signal against: high RMS, high standard deviation, and a specific frequency power curve. Each of these findings was built into an algorithm. It is important that this algorithm can be operated against an unbounded data stream and can make decisions about the state of the headphone with an acceptable degree of latency. Thus a windowing approach has been used, taking the last 5000 samples as the current window and the previous 10 windows as the lookback space. This gives a decision making delay of approximately 10ms and requires approximately 1 second of previous audio. Testing to choose optimal values for the window size and lookback space was undertaken using the experimental data set. Initial values have been chosen based on testing with the full experimental data set. With this in mind, the algorithm functions as follows:
235
236\begin{enumerate}
237\item The signal from both microphones is passed through a low pass filter to remove additional higher frequency noise. The absolute vales of the output is taken, as it is not important if the variance is positive or negative.
238\item The difference between the standard deviations of the inside and outside microphones is taken. If the external microphone std dev is larger than the interior, the source of the noise is external and uninteresting, so the difference value is set to 0. This allows signals from the external microphone to be ignored, but also for signals picked up on both microphones to be discounted as the insertion and removal signal is specific to the internal microphone. Figure \ref{stddev} visualises the data for the same recording in figure \ref{raw}. Again, two bursts are visible at either end of the recording, this time with all values appearing between 0 and 1, aiding in comparison. The insertion and removal points detected by the algorithm have been marked in red. From this is is clear to see that the signal created by the insertion and removal movement is significantly different from any other feature on the signal.
239
240\begin{figure}[h]
241 \centering
242 \includegraphics[width=0.7\linewidth]{Figures/signal-std-dev.png}
243 \caption{Raw time volume graph of a test sample recording.}
244 \label{stddev}
245\end{figure}
246
247
248
249\item This normalised standard deviation difference is compared to that of the last 10 windows. The standard cumulative sum (CUSUM) algorithm for detecting trends in mean in time series data is used. This returns a non zero value when the upper cumulative sum of the input drifts above five standard deviations above target mean determined from the historic vales.
250\item If this value is found to be greater than zero, the maximum upper cumulative sum is taken from the output and compared to a threshold. If larger than 0.3, the threshold, the algorithm marks the window as containing a potential change point. Figure \ref{trig} shows the trigger points generated for the signal input discussed above.
251
252\begin{figure}[h]
253 \centering
254 \includegraphics[width=0.7\linewidth]{Figures/signal-trigger.png}
255 \caption{Raw time volume graph of a test sample recording.}
256 \label{trig}
257\end{figure}
258
259\item Once triggered, the algorithm will not re-trigger for a fixed number of windows, so as to avoid detecting the same insertion multiple times, as the insertion signal tends to last for up to a second.
260\end{enumerate}
261
262The algorithm has been tested against ground truth data, for tuning of the thresholds employed, and to validate the performance of the system on both insertion and removal signals. Details of the performance is given in the results section.
263
264\section{Heartbeat Detection}
265It was found that in some users, it was possible to passively detect their heartbeat on the internal microphone, through the sound of their pulse in their ear canal. This was confirmed initial testing by comparing the period of the signal on the audio to that of the wearer's heartbeat using a heart rate monitoring device. Figure \ref{heart} shows these peaking patterns for two different wearers. This detected change in pressure is likely due to the expansion and contraction of the deep auricular artery, which passes through the wall of the acoustic meatus that forms the wall of the ear canal to supply the skin of the ear canal and the tympanic membrane with blood \cite{Standring2008}. This blood vessel has been associated with sensation of pulsing in the ear, and pulsatory tinnitus \cite{Levine2008}, suggesting that some individuals can sense the same pressure changes picked up by the microphone. This effect was particularly pronounced on the left-hand recording, it is suspected that this is due to the fact it is closer to the heart.
266
267\begin{figure}[h]
268 \centering
269 \includegraphics[width=0.7\linewidth]{Figures/HeartRate.png}
270 \caption{Two magnitude-time graphs showing regular peaks confirmed to be a result of cardiac contraction in the wearer.}
271 \label{heart}
272\end{figure}
273
274When tested on a small set of users, it was found that this sound was particularly intense immediately after the user had been doing heart rate raising activity. Figure shows an example of the signal recorded immediately after the wearer had stopped exercising, and had a ground truth heart rate of 121BPM.
275
276\begin{figure}[h]
277 \centering
278 \includegraphics[width=0.7\linewidth]{Figures/elevated_HR.png}
279 \caption{Magnitude-time graph showing increased regular peaks as a result of post-exercise cardiac contraction in the wearer.}
280 \label{heart}
281\end{figure}
282
283The signal was found to exist between 10Hz and 40Hz, which made it hard to isolate, as a large amount of background noise is in these ranges. For example the friction of the signal cable on the side of the user's head and the impact noise generated when taking steps. This made it impossible to detect the signal during any kind of movement from the wearer. It was also tested if the signal could be detected while the user was using the headphones for another function, such as listening to music. Unfortunately it was found that this effectively masked the signal, and could even lead to misleading results from the percussive beats of the music. Efforts were made to build a system to reliably extract this data from the 15 seconds of silence captured in the user experiment, but this was not achieved as the signal did not appear in the recordings of different users at any significant level above general background noise (from air conditioning etc.) and the noise floor of the microphones themselves. It appears that the morphology of the individual plays an important role in the detectability of the pulse in the ear as it was present across multiple repeats for some users but not others.
284\section{User Fingerprinting}
285This section describes the algorithms involved in each step of the fingerprinting process, from raw audio captured at the microphone to a fixed length feature vector which forms the fingerprint. The steps in this process are described in figure \ref{fingersteps}.
286
287\begin{figure}[h]
288 \centering
289 \includegraphics[width=0.7\linewidth]{Figures/steps_to_finger.png}
290 \caption{Flow chart showing process from raw audio to similarity score output.}
291 \label{fingersteps}
292\end{figure}
293Two methods for isolating the signal in the input audio are presented along with two different methods of fingerprinting the audio and an algorithm for each to enable comparison between them.
294\subsection{Fingerprint Chirp Isolation} \label{finger:isolate}
295This algorithm detects and extracts the chirps from the input audio recorded in the wearer's ear. It was a requirement that the algorithm be able to detect the start of the chirp as close to the actual event possible. This is because the fingerprinting algorithm compares the value at each point in the two fingerprints, so if the alignment is significantly off it could reduce the correctness of the system. Two independent methods have been generated to achieve this, each of which offers its own set of compromises -
296\begin{itemize}
297\item \textbf{Algorithm 1: }More precise in detecting the correct start point of the signal, less able to work with quieter echoes, or signals with significant noise. This means that it fails for test chirps with extremely low volume.
298\item \textbf{Algorithm 2: }The second algorithm is much better at detecting a quiet or noisy signal, however it does not have the same degree of precision as the first, thus reducing the quality of the output.
299\end{itemize}
300The outputs of both algorithms is the indexes of the start and end of each chirp in the data stream, defining the start and end point of each fingerprint.
301
302Figure \ref{isolate} below visualises the output of this component of the system. The blue line shows the recorded signal amplitude over time. To the observer it is clear to see that there are three chirps of significant magnitude, separated by a period of silence. The vertical green lines are the detected start and end points of the signal.
303\begin{figure}[h]
304 \centering
305 \includegraphics[width=0.9\linewidth]{Figures/isolated-signal.png}
306 \caption{Example chirp response envelope (blue) with the automatically detected change points (green).}
307 \label{isolate}
308\end{figure}
309
310\subsubsection{Chirp Algorithm 1}
311
312The algorithm takes as input the signal envelope and the number of chirps to detect, which serves as a hint for when the system has found the correct number of signals.
313
314\begin{enumerate}
315\item A threshold is defined as the 35th percentile of the signal magnitude.
316\item The first time the signal exceeds this value is taken as the start of the signal.
317\item As the length of the chirp in samples is known, the algorithm then skips through the signal to just before the point at which it is expected that the signal will end.
318\item The first time the signal drops below the threshold value is taken as the end of the chirp.
319\end{enumerate}
320
321This works well when the signal is loud compared to the noise. It is precise as no smoothing is applied to the signal to distort its shape over time. However, due to the fixed threshold and the lack of smoothing, a signal with a low SNR is liable to become lost in noise using this method. If this algorithm fails to find an acceptable signal, it will fall back to the second algorithm defined below.
322
323The threshold value was determined through correctness evaluation on the test data. The precision and recall of the system was compared to manually annotated ground truth data at 0.5\% threshold intervals. For the test data collected, 35\% was found to offer the best compromise between early and late starts.
324
325\subsubsection{Chirp Algorithm 2}
326
327The input for this method is the same as for algorithm 1. Figure \ref{algo2} shows a simplified example of the algorithm's operation. It proceeds as follows:
328
329\begin{enumerate}
330\item The signal is passed through a band pass filter, to remove noise above and below the frequencies of the chirps themselves and is smoothed by means of a simple moving average. This is plotted as the blue series in the example figure.
331\item A threshold value is generated initially as the 50th percentile of the signal magnitude. This is then placed on the signal magnitude as a theoretical horizontal line. This is the orange series in figure \ref{algo2}.
332\item The number of times the magnitude envelope crosses this threshold is calculated.
333\item The threshold is adjusted up or down until the number of intersections is equal to the twice the number of chirps: one for the ascending leg of the envelope, and one for the descending. Part 2 of the example shows this point.
334\item Once this is found, the threshold will decrease as far as possible before the correct number of intersections is lost, as seen in part 3 of the example.
335\item The algorithm will then backtrack to the lowest point at which exactly 2 intersections were detected. This can be seen in part 4.
336\item The indexes of these intersections is taken as the start and end points of the chirp.
337\end{enumerate}
338
339\begin{figure}[h]
340 \centering
341 \includegraphics[width=0.9\linewidth]{Figures/simplified-isolation.png}
342 \caption{Chirp Isolation Algorithm 2}
343 \label{algo2}
344\end{figure}
345
346Should the algorithm fail to find the correct number of intersections at any threshold, it is likely do moving average having not smoothed the envelope enough to represent the start and end of each chirp as as a single ascending and descending line. Thus, the moving average window is increased and the algorithm retried. This reduces the likelihood of a jagged signal causing a false negative due to too many intersections, but reduces precision. This is tried up to three times, at which point the algorithm fails.
347
348The advantage of this algorithm is its ability to work with smoothed data, and to also iterate on the threshold to improve accuracy. In the face of a quiet signal with respect to the noise level, the algorithm is also able to adjust its sensitivity. However the inaccuracy introduced by smoothing the data reduces the precision of the result.
349\subsection{Fingerprint Signal Preprocessing}
350The next step is to take the isolated chirp and preprocess it so a fingerprint can be built from it.
351The first two steps are preprocessing of the signal, and are identical for both algorithms, the input to both is a waveform containing a single chirp as isolated by the algorithms discussed earlier. Figure \ref{single} shows an example of a chirp used to build a fingerprint, which will be used to illustrate the steps each algorithm takes.
352
353
354
355\begin{enumerate}
356\item The envelope of the signal is calculated. Figure \ref{envelope} shows this.
357\item A smoothing function is applied to the envelope to remove high frequency noise, while preserving the shape of the signal. Figure \ref{smooth} shows the smoothed envelope for the raw signal in figure \ref{single}.
358\end{enumerate}
359
360
361\begin{figure}[h!]
362 \centering
363 \includegraphics[width=0.7\linewidth]{Figures/finger-single.png}
364 \caption{A single isolated chirp response.}
365 \label{single}
366\end{figure}
367\begin{figure}[h!]
368 \centering
369 \includegraphics[width=0.7\linewidth]{Figures/finger-env.png}
370 \caption{Smoothed analytic signal envelope.}
371 \label{envelope}
372\end{figure}
373
374\begin{figure}[h!]
375 \centering
376 \includegraphics[width=0.7\linewidth]{Figures/finger-smooth.png}
377 \caption{A single isolated chirp response analytic signal envelope.}
378 \label{smooth}
379\end{figure}
380
381\newpage
382
383\subsection{Frequency Density Histogram Fingerprint}
384Once the echo responses have been isolated and preprocessed, the next step in the system is to produce a fingerprint of the signal to enable comparison.
385
386With this in mind, several requirements were generated for the fingerprints:
387
388\begin{itemize}
389\item It should have a very large number of possible fingerprints, so as to avoid collision for non-equal signals.
390\item For a given signal, the fingerprint should be deterministically calculated, such that fingerprint(x) = fingerprint(y) if x = y.
391\item Intuitively, no two captures are going to be sample-for-sample identical so the system should have tolerance for signals that are similar, but not identical. Factors such as the fit of the device in the ear, the humidity of the air in the ear canal and the homeostatic state of the wearer are likely to all have an effect on the output, and the system should be resilient to these changes.
392\item It should be eventually plausible that the fingerprinting and comparison operations could be performed on mobile hardware - i.e they should not require large amounts of compute or memory. This leads to a goal of making the algorithm as computationally simple as possible.
393\item It should be easy and fast to compare the produced fingerprints to determine similarity.
394\end{itemize}
395
396
397
398
399\subsubsection{Frequency Histogram Vector Generation}
400
401The range of frequencies known in the input signal is known, so their presence (or absence) in the response signal can be used to characterise it. This is performed by grouping frequency ranges into buckets and measuring the relative average power at each of those frequency bands. The resultant spectral power distribution thus forms a feature vector which can enable comparison.
402
403The steps to the algorithm are as follows:
404
405\begin{enumerate}
406\item An empty histogram structure is created with 100 buckets. Each represents a frequency range of 100Hz from 10kHz to 20kHz.
407\item The average power of the signal in that frequency band is taken at 100Hz intervals, using a bandpower function.
408\item This average power value is inserted into the histogram in the bucket corresponding to the frequency range.
409\item Once filled, the histogram is then normalised between 1 and 0.
410\item The completed histogram is outputted as the fingerprint.
411\item A mean fingerprint for multiple input chirps is generated by creating an individual histogram for each input chirp and taking the mean of each bucket.
412\end{enumerate}
413
414By normalising the values in the histogram, differences in absolute power (volume) are ignored, creating tolerance in the system to different fits in the user's ear.
415
416\begin{algorithm}
417\caption{Histogram Fingerprint Generation}\label{histogen}
418\begin{algorithmic}[1]
419\Procedure{HistoGenerate}{$A$}\Comment{Preprocessed signal}
420
421\State $b \gets []$
422\State $f \gets 10,000$ \Comment{Starting frequency}
423\\
424\For{$i \gets 0$, $100$}
425 \State $p \gets bandpower(A,f,f+100)$ \Comment{Takes signal and frequency range}
426 \State $q \gets [q : p]$\Comment{Append power to array tail}
427 \State $f \gets f + 100$ \Comment{Move up frequency range}
428\EndFor
429\\
430
431\Return{f} \Comment{Final fingerprint}
432\EndProcedure
433\end{algorithmic}
434\end{algorithm}
435
436\subsubsection{Histogram Comparison}
437
438Each histogram feature vector contains 100 buckets, each with a non negative number between 1 and 0 representing the power detected in that frequency range. Given two histogram fingerprints, the comparison algorithm described below will output a percentage similarity between 0 and 100. It proceeds as follows:
439
440\begin{enumerate}
441\item An initial similarity score of 100 is assigned.
442\item The histograms are iterated though, and the absolute difference between the corresponding buckets taken.
443\item The difference is subtracted from the similarity score.
444\item The final similarity score is outputted as the result.
445\end{enumerate}
446
447The closer the bucket power densities are to each other, the smaller the difference value that is subtracted, thus the better the final similarity score.
448\subsection{Waveform Feature Vector Fingerprint}
449
450The fingerprint described previously is reliant on the absolute frequencies that the fingerprint response features occur at remaining constant. The second algorithm ignores as much of the absolute values of the chirp response as possible, instead focusing on the \textit{shape} of the response. The hypothesis behind this is that such an algorithm will have better resistance to noise in the fingerprint as it is not dependent on the exact frequencies in the signal. The algorithm below ignores exact frequencies and volumes, instead focusing on the relationship between different segments of the signal.
451
452\subsubsection{Waveform Feature Vector Generation}
453
454The second fingerprinting algorithm is specified as follows. It takes as input the preprocessed chirp. Again, the example in figure \ref{single} will be used here again to illustrate. The system outputs an array of values, proportional to the length of the input chirp. Each value in the array is either 1, 0 or -1. The combination of these values forms the fingerprint. The algorithm applies the following steps:
455
456
457\begin{enumerate}
458\item The preprocessed signal is divided into 100 segments, and the mean signal magnitude in each is taken. This mean value is described as the \textit{segment-point}. This is stored in an array structure, $q$.
459\item An empty fingerprint structure, $m$, is created and a tolerance factor, $t$, of 2\% either side of the mean of the entire signal magnitude is calculated. This gives the system a reference point of the scale of the signal allowing signals of different magnitudes to be fingerprinted with consistent sensitivity.
460\item Each segment-point $q[i]$ is iterated over, and compared to its following neighbour $q + 1$ to determine the relative difference in magnitude - $\Delta$. The $\Delta$ value is plotted as an intermediate state in figure \ref{inter}.
461\begin{figure}[h]
462 \centering
463 \includegraphics[width=0.7\linewidth]{Figures/finger-inter.png}
464 \caption{Intermediate fingerprint $q$ plot, showing magnitude of difference between segments.}
465 \label{inter}
466\end{figure}
467\begin{itemize}
468\item If the segment-point is smaller than its neighbour minus the tolerance factor, the fingerprint is marked with a 1 at that point, indicating that the segment is an ascending one.
469\item If the reverse is true: the segment-point is larger than its neighbour plus the tolerance factor, the fingerprint is marked with -1 at that point, indicating a descending segment.
470\item If neither of the above apply, the difference falls within the tolerance - no significant change is detected and the fingerprint is marked with a 0 at that point.
471\end{itemize}
472\item The filled fingerprint structure $m$ is returned as the result. Figure \ref{result} shows a graphical representation of the data within the fingerprint data structure for the initial fingerprint signal seen in figure \ref{single}.
473\begin{figure}[h]
474 \centering
475 \includegraphics[width=0.7\linewidth]{Figures/finger-finger.png}
476 \caption{Final fingerprint $m$ plot.}
477 \label{result}
478\end{figure}
479\end{enumerate}
480Algorithm \ref{wavegen} specifies the algorithm formally.
481\begin{algorithm}
482\caption{Waveform Fingerprint Generation}\label{wavegen}
483\begin{algorithmic}[1]
484\Procedure{WaveGenerate}{$A$}\Comment{Preprocessed signal}
485
486\State $q \gets [ ]$
487\State $m \gets [ ]$
488\State $t \gets 0.02 * mean(A)$ \Comment{Tolerance: 2\% entire signal mean}
489\\
490\For{$i \gets 0$, $100$, $length(A)$}\Comment{Iterate in jumps of 100 samples.}
491 \State $\Delta = mean(A[i]...A[i+99]$
492 \State $q \gets [q : \Delta]$\Comment{Append mean to array tail.}
493\EndFor
494\\
495\For{$i \gets 0$, $length(q)-1$}
496 \If{$q[i] < (q[i+1]) - t$}
497 \State $m[i] \gets 1$
498 \ElsIf{$q[i] > (q[i+1]) + t$}
499 \State $m[i] \gets -1$
500 \Else{}
501 \State $m[i] \gets 0$
502 \EndIf
503\EndFor
504\Return{m} \Comment{Final fingerprint}
505\EndProcedure
506\end{algorithmic}
507\end{algorithm}
508
509The fingerprint of each chirp is calculated individually, and can then be combined into a single final fingerprint, calculated as the median of the values in segment point. This is shown in algorithm \ref{wavemed}.
510\begin{algorithm}
511\caption{Median Fingerprint Generation}
512\label{wavemed}
513\begin{algorithmic}[1]
514\Procedure{WaveMedian}{$m1,m2,...,mn$}\Comment{Fingerprints to median}
515\State$ q \gets []$
516\State $l \gets min(length(m1),length(m2),...,length(mn))$
517\For{$i\gets 1$,$l$}
518 \State $x \gets median(m1[i],m2[i],...,mn[i])$ \Comment{median of segment $i$}
519 \State $q \gets [q : x]$ \Comment{Append x to tail of q}
520\EndFor
521\Return{q} \Comment{Final median fingerprint}
522\EndProcedure
523\end{algorithmic}
524\end{algorithm}
525
526Effectively, the fingerprint marks at each point in the signal if the magnitude is increasing or decreasing with respect to its neighbour, reducing each segment to a simple increasing, decreasing or steady value. This approach of abstraction has several advantages. First, it completely ignores any detail of the absolute volume and frequency, making it resilient to changes in these domains between samples. Furthermore, the exact gradient of the change is ignored, again making the fingerprint resistant to small changes in the exact response. The small size of the segment leaves room for a huge number of possible fingerprints. For example, if a chirp response contains 100 segments there are 5.15e+47 possible combinations.
527
528Furthermore, the fingerprint itself is small in terms of space, consisting of only 100 integers. This makes the system space efficient as once the fingerprint has been generated, the audio signal can be discarded.
529
530\subsubsection{Waveform Comparison}
531
532An algorithm has been developed to enable the comparison of two waveform fingerprints, $A and B$, outputting a percentage similarity score. These fingerprints do not need to be the same length. 100\% indicates that at every segment, the two fingerprints hold the same value. 0\% indicates that the fingerprints are opposite at every point. The method is simple, making $ O(n) $ comparisons where $ n $ is the number of segments in the shortest input fingerprint. For the implementation described here, this results in around 100 comparisons per two input fingerprints.
533
534Each segment in the shortest input fingerprint (input lengths may differ) is given percentage weighting, $w$.
535For example, with 100 segments, each segment is weighted at 1\%. Each fingerprint is then compared segment by segment, and the similarity weights summed to create the final score $x$:
536
537\begin{itemize}
538\item 1 x percentage weight for an exact match
539\item 0.5 x percentage weight for a difference of 1 - i.e (1 and 0) or (0 and -1)
540\item 0 x percentage weight for a difference of 2 - i.e (1 and -1)
541\end{itemize}
542
543This is shown formally in algorithm \ref{wavecomp}.
544
545\begin{algorithm}
546\caption{Waveform Fingerprint Comparison}\label{wavecomp}
547\begin{algorithmic}[1]
548\Procedure{WaveCompare}{$A,B$}\Comment{Fingerprints A and B}
549
550\State $l \gets min(len(A),len(B))$ \Comment{Shortest fingerprint length}
551\State $m \gets 0$ \Comment{Initial Similarity Score}
552\State $w \gets 100 \div l$ \Comment{Segment weighting}
553
554\For{$i\gets 1$, l}
555\If{abs($A[i]-B[i]$) == 0}
556 \State $m \gets m + w $
557\ElsIf{abs($A[i]-B[i]$) == 1}
558 \State $m \gets m + (0.5 * w) $
559\Else
560 \State $m \gets m + (0 * w) $ \Comment{Shown for clarity}
561\EndIf
562\EndFor
563\State \textbf{return} $m$\Comment{Final Score}
564\EndProcedure
565\end{algorithmic}
566\end{algorithm}
567
568
569These weighted percentages are summed and the final value returned as the similarity result.
570Figure \ref{compare} demonstrates an example of the comparison of two fingerprints, from the left and right ear of a wearer. The areas highlighted in green show a match in the fingerprint segments, yellow a partial match and red a mismatch. The similarity score for these two fingerprints is 69.5\%.
571
572\begin{figure}[h]
573 \centering
574 \includegraphics[width=0.9\linewidth]{Figures/comp-result.png}
575 \caption{Comparison of two fingerprints.}
576 \label{compare}
577\end{figure}
578
579\subsection{Identification Through Direct Comparison}
580This method allows for the identification of wearer by taking a supplied live fingerprint and comparing against the data corpus of all known fingerprints. It uses either of the fingerprinting algorithms described before, simply swapping the comparison algorithm as required. The system attempts to answer the question "who is this user?", by comparing the live fingerprint to the database of enrolled fingerprints in a one to many comparison. It requires that for each user a set of fingerprints is stored that can be compared against at identification time. This therefore requires that a wearer be enrolled in the system a before identification can take place.
581Either one or two new fingerprints are taken to identify - one per ear. Both operation modes are shown in algorithms \ref{idnentAlgoSingle} and \ref{idnentAlgoDual}, respectively.
582
583The system iterates through the database of enrolled prints and uses one of the previously described comparison algorithms to determine their similarity. The mean similarity per ear per enrolled user is recorded. Then, for each ear, the enrolled user with the highest mean similarity is taken as the match. If the same user is identified for both ears independently, that user is returned as the match. If different users are returned per ear, the match with the highest degree of certainty is taken - provided the confidence level is at least 15\% greater. If not, the identification fails due to lack of confidence in the matches.
584
585The time complexity of the matching algorithm scales linearly with the number of enrolled users. Alternative methods were explored for reducing the number of required comparisons, such as comparing the unidentified fingerprint to the median of each user's enrolled fingerprint, thus reducing the number of comparisons per user to 1, but this was found to compromise performance.
586
587\begin{algorithm}
588\caption{Single Fingerprint Identification}\label{idnentAlgoSingle}
589\begin{algorithmic}[1]
590\Procedure{IDENT1}{$C,D$}\Comment{Challenging fingerprint $C$, and database $D$}
591
592\State $m \gets 0$ \Comment{Max confidence}
593\State $u \gets null$ \Comment{Best guess user}
594\For{$i \gets 1$,$length(D)$}
595 \State$ s \gets similarity(C,D[i])$
596 \If{$s > m$}
597 \State $ m \gets s$
598 \State $ u \gets i$
599 \EndIf
600\EndFor
601\Return{$m,u$}\Comment{Return best guess and confidence.}
602\EndProcedure
603\end{algorithmic}
604\end{algorithm}
605
606\begin{algorithm}
607\caption{Dual Fingerprint Identification}\label{idnentAlgoDual}
608\begin{algorithmic}[1]
609\Procedure{IDENT2}{$L,R,D$}\Comment{Challenging fingerprints $L,R$, and database $D$}
610
611\State $mL, mR \gets 0$ \Comment{Max confidence L + R}
612\State $uL, uR \gets null$ \Comment{Best guess user L + R}
613
614\For{$i \gets 1$,$length(D)$}
615 \State$ sL \gets similarity(L,D[i])$
616 \State$ sR \gets similarity(L,D[i])$
617 \If{$sL > mL$}
618 \State $ mL \gets s$
619 \State $ uL \gets i$
620 \EndIf
621 \If{$sR > mR$}
622 \State $ mR \gets s$
623 \State $ uR \gets i$
624 \EndIf
625\EndFor
626\If{$uR == uL$}\Comment{Both agree}
627 \Return{$mean(mL,R),uL$}
628\Else \Comment{No agreement}
629 \If{$mL > mR + 15$}
630 \Return{$mL,uL$}
631 \ElsIf{$mR > mL + 15$}
632 \Return{$mR,uR$}
633 \Else
634 \space$ $\Return{$0,null$}
635 \EndIf
636\EndIf
637\EndProcedure
638\end{algorithmic}
639\end{algorithm}
640
641\subsection{Verification by Thresholding}
642
643This method is the alternative functional mode of a biometric system, which attempts to verify the authenticity of the supplied fingerprint in a one to one comparison. In other words, "Is the user who they say they are?". An example use of this technique is for example for unlocking a mobile phone. It is not important to know \textit{who} the user is, just if their fingerprint matches the person who is authorised to use the device. This therefore reduces the number of comparisons required, as only one user's set of prints must be considered. It does however make the problem of deciding when to accept a fingerprint as valid more difficult, as the frame of reference for what constitutes a good match provided by the comparisons with other users is lost.
644
645The approach taken to solve this problem, shown in algorithm \ref{verifAlgo} is to compare the new print to the set of known prints, and to then take some aggregate statistic of their similarities. If the aggregate similarity is above some threshold, then accept the new print as being a match for the saved prints. Otherwise a match cannot be attained with sufficient confidence so the fingerprint is rejected as a non match.
646
647This presents a performance trade off between a higher false positive rate with a lower confidence threshold, and a higher false negative rate a higher thresholds. This ultimately forms a trade off between convenience and security. The performance of the system at different thresholds on the user data set is evaluated in the results section.
648
649\begin{algorithm}
650\caption{Single Fingerprint Verification}\label{verifAlgo}
651\begin{algorithmic}[1]
652\Procedure{VERIF}{$C,K,T$}\Comment{New f/print $C$, known f/prints $D$, Threshold $T$}
653
654\State $m \gets []$ \Comment{Confidences}
655
656
657\For{$i \gets 1$,$length(D)$}
658 \State$ s \gets similarity(C,D[i])$
659 \State$ m \gets [m : s]$
660\EndFor
661q = mean(m) \Comment{Mean of all similarities}
662\If{$q \geq T$}
663 \Return{True} \Comment{Accepted}
664\Else
665 \space\Return{False} \Comment{Rejected}
666\EndIf
667\EndProcedure
668\end{algorithmic}
669\end{algorithm}
670
671\subsection{Data Fusion}
672By combining the waveform feature vector and the frequency density histogram, a more robust biometric was created. This gives a total of 4 feature vectors per user, when sampling from both ears. The identification or verification procedure was carried out as described previously. In order to combine the signals the following techniques were applied:
673\begin{itemize}
674\item For identification: The strict majority was taken. If no strict majority was present, the outcome with the highest cumulative confidence was taken. This was calculated by summing the confidence for each outcome possibility.
675\item For Verification: The mean of the verification scores was taken. This results in an acceptance area as described in figure \ref{accept}, where a signal scoring poorly in one dimension may still be accepted if it scores well in the other.
676\end{itemize}
677
678\begin{figure}[h]
679 \centering
680 \includegraphics[width=0.4\linewidth]{Figures/hybrid-accept.png}
681 \caption{Shape of accepted inputs when using thresholding system based on mean of two metrics, x and y.}
682 \label{accept}
683\end{figure}
684
685\chapter{Evaluation}
686
687The experiment carried out has aimed to target the three most promising areas of potential data. The first of these was detecting and characterising the signal produced when the headphones are inserted or removed from the user's ear. Next, passive audio data was collected, while the user was sat still and breathing normally. This was to determine if the pulse of the wearer could be reliably detected. Lastly, the echo response from high frequency chirps was collected, to analyse the performance of the fingerprinting system. With this in mind, the following procedure was devised and carried out on a set of volunteer participants.
688
689A total of 32 participants have been tested. Of these, 52\% were male and 48\% were female. The minimum age was 18, the maximum 62. The mean age was 35. Participants younger than 18 were not tested as the headphones are specified to fit the median adult, and thus were found to be too large for all of the participants attempted under 18. The tests were carried out in a variety of locations, with 22 being carried out in a quiet environment, 5 in an area with moderate background noise - an office location, and 5 in a noisy cafe environment. In all of these environments, no particular controls over background noise were made.
690
691The use of an oscilloscope to capture data from the microphones was considered but discarded for several reasons. The first being that the system would require custom hardware to power the microphone and amplify the output, which combined with the size and power requirements of the scope significantly reduced portability. Second, the microphones in the headphones are designed to capture audio mainly in the audible spectrum. Thus they are only rated to 20kHz, a frequency that can be captured at the laptop sample rate of 48kHz.
692
693\section{User Study Design}
694
695Participants were asked to place the headphones in their ears once recording was started. Once the wearer had the headphones seated comfortably, the test audio track was started. This would play a tone, marking the start of the test, followed by 15 seconds of silence. The data in this section was used for pulse detection. Next, a battery of 10 chirps was played, each lasting 20ms and ranging from 10kHz to 20kHz. This battery was repeated 8 times, each time reducing the volume of input signal by 10\%. This gave a total of 400 chirps recorded per user per ear. Each battery was separated by a mid-frequency tone, chosen to be at the peak sensitivity of the microphones, so that they were detected clearly on the output. This aided in automatic parsing of the test data. At the end of the test batteries, the user removed the headphones after which the recording was stopped. This ensured that both the insertion and removal signal was captured. Each test was repeated 5 times. This resulted in a total of 800 individual chirps recorded per participant, including left and right ears, 10 insertion signals, and 10 removal signals.
696
697The tones played at the start and end of the test track acted as ground truth points delimiting the sections of the recording where the user has the headphones in or out of their ear. By reliably detecting this, the correctness of any insertion/removal algorithm were verified easily - it simply must independently detect insertion before the start signal, and independently detect removal after the end signal, but before the end of the recording.
698
699
700\subsection{Sound and Ultrasonic Safety}
701
702The test track chirp used to generate the fingerprint has a range of 10kHz to 20kHz, which is approaching ultrasonic frequencies, as the limit of human hearing is usually around 20kHz. To ensure that the signal used to generate the fingerprint response was safe to expose to participants and was of an acceptable volume so as not to risk discomfort or hearing damage, an investigation to the safe limits of sound exposure was carried out, with a particular focus on high frequency to ultrasonic sounds. Overexposure to ultrasound can have severe adverse consequences. One study of ultrasonic noise reported to create sensations of dizziness, nausea and tinnitus in those exposed to it at high volumes. These symptoms were reported at sound levels over 85dB for frequencies above 20kHz \cite{Smagowska2013}.
703
704Concrete guidance on safety when working with ultrasound is given by the UK Health and Safety Executive, and is presented with a set of research into previous studies informing their guidance \cite{Lawton2001}. The research found few subjective complaints and no measurable temporary hearing loss in subjects exposed to ultrasound at volume levels that lead to significant temporary hearing loss and complaints at audible frequencies. Crucially they note "These results indicated that air-borne ultrasound is considerably less hazardous than audible sound." \cite{Lawton2001} The work cites previous investigations carried out on ultrasonic sound safety limits. These papers take the "work day exposure" approach, which assumes that a worker would be exposed to sound at this level for a full 8 hour work day, on a regular basis. All recommend exposure is kept below 120dB - 75dB for signals above 20kHz, decreasing to 80dB - 75dB for signals below 10kHz. Canada, the USA and Australia have work-day exposure limits of 110dB for sound over 20kHz.
705
706In an abundance of caution, the work here used high frequency and ultrasound in short bursts and well below these levels, while also using the highest frequency possible. Work investigating the maximum sound pressure output of commercially available headphones found it to be in the range of 91dB and 112dB \cite{FligorBrianJ.;Cox2004}, which is approaching but not exceeding the safe limit set out by the work discussed. A careful but common sense approach of ensuring sound settings on devices were kept well below their maximum settings was used, as it appears headphones are not capable of emitting hearing-damage inducing levels of ultrasound, particularly not in shorter bursts than the 8-hour assumption used in the calculation of the safety guidance. This approach was used for any work with sound in audible frequencies.
707
708The WHO and the American Speech-Language-Hearing Association recommend not listening to sound through headphones at more than 60\% of the maximum output for more than an hour a day \cite{WorldHealthOrganisation2011}. Clearly while devices vary, this was taken into account and this 60\% threshold acted as the absolute limit of output volume, and was limited by software on the device at all times to prevent accidental breach of this limit while using ultrasound. Software for implementing this in MacOS exists\footnote{http://www.anoshkin.net/products/mac/volimiter/}, the development platform used here.
709\section{User Study Results}
710\subsection{Captured Data}
711
712The total data recorded exceeded 8GB in size, and was over 5 hours worth of audio. Several functions have been defined in MATLAB to aid processing the data, each defined in such a way that it adds value to the data even in isolation of the broader experimental context. The first locks onto the tones that delimit the experiment start and end - every point in the test is relative to these. To do this, a copy of the signal is passed through a notch filter that isolated the specific frequencies of the tone. The first point found to have a relative magnitude greater than 0.5 is then taken as the start point of the first tone. From this point, the start and end points of all the other tests in the audio can be located. Figure \ref{setupsound} shows diagramtically the sections of the audio track that are delineated. It consists of the following sections:
713
714\begin{enumerate}
715\item Section containing the signal generated from headphone insertion.
716\item Test start tone.
717\item 15 seconds of silence for use in heartbeat detection.
718\item Fingerprint chirp test start signal.
719\item 8X fingerprint chirp batteries - each contains 10 chirps of decreasing volume. The chirps are not isolated this way - a greater degree of precision was required, as this system provides the sections of the experimental track separated with a resolution of 100-200 samples. Furthermore, as this system is specifically designed to work with the experimental test track, this method would not be useful in general application of the technology.
720\item Test end tone.
721\item Section containing the signal generated from headphone removal.
722\end{enumerate}
723
724\begin{figure}[h]
725 \centering
726 \includegraphics[width=0.7\linewidth]{Figures/TestSetup.png}
727 \caption{Sections of the audio captured during the user study.}
728 \label{setupsound}
729\end{figure}
730
731\subsubsection{Analysis of Fingerprint Data}
732
733This section analyses the source of the data within in the fingerprints themselves that leads to the identifiable differences between users, and consistency between readings of the same wearer. To reveal this, a test was set up to systematically remove segments, comparing the changes in precision and recall for these sub-fingerprints. This is equivalent to removing chunks of frequencies from the input signal, to produce a reduced frequency range over which the fingerprint was built. The result has been compiled into an F1 score for the thresholding technique operating on different sub-segments of the fingerprints. Figure \ref{subfingerhigh} shows the performance of the system as the fingerprint size is increased, by initially starting only with the lowest frequencies present. This shows that the wider the sine sweep in the chirp, the better the performance and does not suggest that the data of interest is embedded in any particular high frequency area of the signal.
734
735\begin{figure}[h]
736 \centering
737 \includegraphics[width=0.6\linewidth]{Figures/trimhigh.png}
738 \caption{Thresholding F1 Score When Removing High Frequency Fingerprint Segments}
739 \label{subfingerhigh}
740\end{figure}
741\begin{figure}[h]
742 \centering
743 \includegraphics[width=0.6\linewidth]{Figures/trimlow.png}
744 \caption{Thresholding F1 Score When Removing Low Frequency Fingerprint Segments}
745 \label{subfingerlow}
746\end{figure}
747
748
749However, this test does not reveal the importance of the lowest frequencies in the fingerprint, as the smallest size tested here is 5 segments wide. To analyse this, the same tests were run, this time excluding the first 5 segments from all of the sub-segments. The results of this are plotted in figure \ref{subfingerlow}. It can be seen that removing even just the first 5 segments reduces the system's F1 score to below 0.6. Removing further segments quickly reduces the system's performance further as it approaches a classification quality of below 0.5. This suggests that the first 5 segments, corresponding to the first 500Hz of the sweep are extremely important, and that the low frequency segments in general are vital. Considering the two results together suggests that the low frequencies are particularly important, but that data at all frequencies in the chirp plays an important role in achieving a low error rate.
750
751\subsection{Fingerprinting Performance}
752This section analyses the performance of the acoustic fingerprint as a biometric. It covers
753\begin{itemize}
754\item The ability to isolate the fingerprint signal form background noise.
755\item The power of the similarity algorithms to discriminate between matching and non matching fingerprints
756\item Whole system identification and verification performance and resistance to environmental noise.
757\item Ability to detect insertion and removal of the headphones.
758\item Stability of the fingerprints over time.
759\end{itemize}
760
761\subsubsection{Fingerprint Signal Extraction}
762Accuracy was deemed to be acceptable when the system isolated the chirp start and end point within less than 50 samples of the ground truth. This is because any offset greater than this would impact the fingerprint produced. The sections isolated here are passed into the algorithms defined in section \ref{finger:isolate}.
763The pair of algorithms used to extract the start and stop points of the chirps for fingerprinting proved robust, showing good performance as the volume of the input signal was decreased. At 80\% input volume, the system was able to detect all fingerprints with an acceptable degree of accuracy from all 32 test participants.
764
765It has been found that algorithm 1 was able to find a signal in the vast majority of cases, with the performance decreasing as the volume of the input chirp was decreased. For the standard signal at 0.8 relative magnitude, algorithm 1 was able to to find a signal 86.1\% of the time. This decreased as volume decreased, however in all but the very quietest case both algorithms combined were able to find a signal. This can be seen in table \ref{table-signal}.
766
767\begin{table}[h]
768\centering
769\caption{Signal capture functionality against decreasing volume.}
770\label{table-signal}
771\begin{tabular}{|l|lll}
772\hline
773\textbf{Volume (relative)} & \multicolumn{1}{l|}{\textbf{Algorithm 1}} & \multicolumn{1}{l|}{\textbf{Algorithm 2}} & \multicolumn{1}{l|}{\textbf{Failed}} \\ \hline
774\textbf{80\%} & 86.1\% & 13.9\% & 0\% \\ \cline{1-1}
775\textbf{70\%} & 72.6\% & 27.4\% & 0\% \\ \cline{1-1}
776\textbf{60\%} & 71.4\% & 28.6\% & 0\% \\ \cline{1-1}
777\textbf{50\%} & 71.5\% & 28.5\% & 0\% \\ \cline{1-1}
778\textbf{40\%} & 71.5\% & 28.5\% & 0\% \\ \cline{1-1}
779\textbf{20\%} & 68.5\% & 31.5\% & 0\% \\ \cline{1-1}
780\textbf{10\%} & 65\% & 33.7\% & 1.3\% \\ \cline{1-1}
781\end{tabular}
782\end{table}
783
784These percentage values are taken from performance of the algorithms on the experimental data, which contained 2400 test batteries in total.
785
786
787\subsubsection{Fingerprinting as a Hash Function}
788
789A key measure of the effectiveness of the acoustic fingerprints as biometrics is the degree of similarity between samples known to have the same ground truth. This information is useful when contextualised with the aggregate degree of similarity between fingerprints known to have different ground truth values. An ideal system would show 100\% similarity between two signals classified as having the same ground truth - i.e two fingerprints from the same wearer. The same system would also show very low similarity between fingerprints known to be from different users. The gap between these similarities is the space in which the discrimination between a match and non-match can occur so the larger this is the better.
790
791Table \ref{sims} shows this data for the waveform and histogram function. The hybrid system is not included as it does not make a separate similarity judgement, instead working with judgements provided by the waveform and histogram functions. It can be seen that the waveform function is better able to discriminate between matching and non matching samples, as the gap between the similarities of the two groups is larger. This suggests that the waveform fingerprint will have better performance when used for verification and identification, as the similarity gap between a match and non match is larger, making the difference easier to detect. This is confirmed in the following sections.
792
793\begin{table}[]
794\centering
795\caption{Similarity between fingerprints matching and non matching.}
796\label{sims}
797\begin{tabular}{|l|ll|ll}
798\hline
799 & \multicolumn{2}{c|}{\textbf{Histogram}} & \multicolumn{2}{c|}{\textbf{Waveform}} \\ \hline
800 & \multicolumn{1}{l|}{Mean Similarity} & std.d & \multicolumn{1}{l|}{Mean Similarity} & \multicolumn{1}{l|}{std.d} \\ \hline
801Matching Samples & 94.46\% & 1.82\% & 85.64\% & 3.99\% \\ \cline{1-1}
802Unmatched Samples & 84.84\% & 1.45\% & 68.57\% & 1.69\% \\ \cline{1-1}
803\end{tabular}
804\end{table}
805
806\subsubsection{Identification Performance}
807The performance of the system when comparing an unknown fingerprint to a corpus of known fingerprints for identification is good. The results are presented below in table \ref{identity-wave}. When just using the waveform fingerprint the F1 score for identifying the matching data entry given a single live sample is 0.950, giving an equal error rate (EER) of 5.02\%. However, given samples from the left and right ear known to the belong to the same user, the two data points can be combined. This improves the system's performance yielding an F1 performance of 0.984 for identifying both the wearer and the ear the sample was taken from. This is an equal error rate of 1.6\% over 32 identification attempts. In the study, this corresponded to correct identification of all participants except one. In 81.2\% of cases, the system provided a match on both ears of the wearer, allowing for increased confidence in the identification result. A published example of fingerprint based identification showed equal error rates of 4.2\% on 110 fingers \cite{Tuyls2005}. Notably however, this work includes template protection to ensure the security of the enrolled data, without this the equal error rate was 1.4-1.4\%. Work published in 2018 showed a minimum equal error rate of 0.32\% for partial fingerprint identification \cite{Qin2018}. While the work presented here performed slightly below this level, it is worth noting that this is state of the art performance in an established identification technique, so there is likely room for significant performance improvement with the ear based fingerprinting technique.
808
809
810\begin{table}[h]
811\centering
812\caption{Identification through fingerprinting performance for each operational mode.}
813\label{identity-wave}
814\begin{tabular}{|l|lll|lll|lll}
815\hline
816\textbf{} & \multicolumn{3}{c|}{\textbf{Histogram}} & \multicolumn{3}{c|}{\textbf{Waveform}} & \multicolumn{3}{c|}{\textbf{Hybrid}} \\ \hline
817\textbf{} & \multicolumn{1}{l|}{F1} & \multicolumn{1}{l|}{Precision} & Recall & \multicolumn{1}{l|}{F1} & \multicolumn{1}{l|}{Precision} & Recall & \multicolumn{1}{l|}{F1} & \multicolumn{1}{l|}{Precision} & \multicolumn{1}{l|}{Recall} \\ \hline
8181 Ear & \textbf{0.89} & 1.0 & 0.81 & \textbf{0.93} & 1.0 & 0.87 & \textbf{1.0} & 1.0 & 1.0 \\ \cline{1-1}
8192 Ears & \textbf{0.95} & 1.0 & 0.91 & \textbf{0.98} & 1.0 & 0.97 & \textbf{1.0} & 1.0 & 1.0 \\ \cline{1-1}
820\end{tabular}
821\end{table}
822
823The system's performance can be improved further by applying a the hybrid approach to identification - using both the waveform and histogram fingerprints together. The resultant system was able to correctly identify the user in every case in the experimental data, giving an F1 score of 1.0. To better understand the increased robustness offered by using 4 data points per user, The number of times all four data points agreed on a match was measured, along with the number of samples for which a majority of statistics agreed. A majority was found in 93.20\% of cases, and of these 38.53\% were unanimous. These figures show that while for the sample size of the experiment the outcome improvement was only small (+1.6\% correct identifications), the extra data gives the ability to make much stronger decisions and provides further scope over which to discriminate. It is therefore likely that besides the extra security provided by requiring two separate features of the signal to be forged rather than one, that the system would have a lower false reject rate and a significantly lower false accept rate in a larger test than either of the other fingerprints alone.
824
825\subsubsection{Resistance to Noise}
826
827The fingerprints taken in noisy environments were compared for similarity, in order to understand if the extra background noise led to high variation in the signals, and thus worse performance. This was not found to be the case. For the samples taken in noisy environment, the mean similarity between fingerprints from the same ear was found to be 4.3\% lower than the mean for the fingerprints taken in quiet environments. The lower similarity in noisy environments falls just outside the standard deviation of all samples, suggesting that the system is affected by moderate noise from external environments but not significantly so. It is unavoidable that loud external noises would have an impact.
828
829\subsubsection{Verification Performance}
830
831Good performance for verification was found, when compared to other published systems. The best performing system was the waveform fingerprint, showing better precision and recall than either the histogram or hybrid system The maximum F1 score for verification was 0.984. The hybrid system showed a lower recall than the waveform alone, which is due to the fact that 4 measurements must satisfy the threshold rather than just two, increasing the burden of proof. This makes the system more secure, but also more prone to rejecting a legitimate authentication attempt. The hybrid system did not show a decrease in precision when compared to the waveform only version, indicating that it did not make any more false positives than the waveform system. This The F1 score, recall and precision of each of the verification modes can be seen in table \ref{verification-table}. This suggests that for a compromise between usability and security, the waveform only version is best. However, the addition of the histogram fingerprint in the hybrid version would likely result in a more secure system, at the cost of more false negatives.
832
833\begin{table}[h]
834\centering
835\caption{Verification performance of each operational mode.}
836\label{verification-table}
837\begin{tabular}{|l|lll}
838\hline
839 & \multicolumn{1}{l|}{\textbf{F1 Score}} & \multicolumn{1}{l|}{\textbf{Precision}} & \multicolumn{1}{l|}{\textbf{Recall}} \\ \hline
840\textbf{Histogram} & 0.909 & 0.882 & 0.938 \\ \cline{1-1}
841\textbf{Waveform} & 0.984 & 1.000 & 0.969 \\ \cline{1-1}
842\textbf{Hybrid} & 0.968 & 1.000 & 0.938 \\ \cline{1-1}
843\end{tabular}
844\end{table}
845Figure \ref{f1score} below shows the F1 score of each of the operation modes at thresholds from 1 to 100, in 0.1 increments. It shows that for all of the operation modes, initially the performance increases as the threshold increases, as the number of false positives decreases. All versions experience a dramatic performance collapse as the threshold approaches 100, caused by the threshold increasing beyond the mean similarity for matching samples, resulting in a quickly increasing number of false negatives. The graph suggests that the histogram fingerprint is less effective than the waveform version, as performance of the histogram does not start to increase until a much higher threshold, due to a high number of false positives at thresholds that the waveform and hybrid system see good performance and a low rate of false acceptance. This is due to that fact that the difference in similarity between a match and non matching histogram fingerprint is smaller than that for a waveform fingerprint. Notably, the hybrid system finds its optimal performance at a lower threshold than either the waveform or histogram fingerprints. This is because in combining the two fingerprints, a majority agreement can be reached between the indicators where either of the indicators on their own would have rejected.
846
847\begin{figure}[h]
848 \centering
849 \includegraphics[width=0.9\linewidth]{Figures/Verification-Performance.png}
850 \caption{Thresholding F1 Score for each operational mode, per threshold in 0.1 increments.}
851 \label{f1score}
852\end{figure}
853
854In order to compare the verification performance here to other biometric systems it is necessary to convert the data above into statistics in terms of the false accept rate (FAR) and false reject rate (FRR), as many other biometric verification systems are measured this way. These figures along with those of other biometric systems examples can be seen in table \ref{verify-compare} below. For brevity, only the waveform fingerprint system is considered here, as it was the best performing out of the three modes presented.
855
856\begin{table}[]
857\centering
858\caption{False accept and reject rates of other forms of biometrics.}
859\label{verify-compare}
860\begin{tabular}{|l|llllll}
861\hline
862\textbf{} & \multicolumn{1}{l|}{\textbf{Wave}} & \multicolumn{1}{l|}{\textbf{Gait} \cite{Gafurov}} & \multicolumn{1}{l|}{\textbf{Fingerprint V1}\footnote{This system has an EER of 1.4-1.6\% without template protection.} \cite{Tuyls2005}} & \multicolumn{1}{l|}{\textbf{V2} \cite{Qin2018}} & \multicolumn{1}{l|}{\textbf{Face+Voice} \cite{Phil2012}} & \multicolumn{1}{l|}{\textbf{Voice}\cite{Burget2007}} \\ \hline
863FAR & \textbf{\textless{}0.01\%} & 5\% - 9\% & 6.1\% & 100\% & $\sim$0.5\% - $\sim$4.25\% & 4.7\% - 4.9\% \\ \cline{1-1}
864FRR & \textbf{3.1\%} & 5\% - 9\% & 3.4\% & 100\% & $\sim$0.5\% - $\sim$4.25\% & 4.7\% - 4.9\% \\ \cline{1-1}
865\end{tabular}
866\end{table}
867
868\subsection{Insertion Detection Performance}
869The algorithm described previously was tested against the data captured in the user study. This was 320 insertion and removal signals, from 32 different users in total. The performance of the algorithm was mixed, as can be seen in table \ref{in-out}. Detecting insertion proved to be an easier task than detecting removal, as the signal generated was much louder. This is reflected in the data, which shows the F1 score for insertion to be 0.935, with a high degree of precision. The system made very few false positive identifications, with the main source of error being false negatives due to a quieter than expected signal. The performance of the algorithm at detecting removal was poor. The signal generated from removal was much quieter than for insertion, so a very high number of false negative identifications occurred. This resulted in a poor F1 score of 0.235. Further work would be required to determine a method for characterising the signal that did not suffer from a high false positive rate when attempting to detect removal, due to the low volume of the signal.
870
871\begin{table}[]
872\centering
873\caption{Performance of insertion/removal detection algorithm}
874\label{in-out}
875\begin{tabular}{|l|lll}
876\hline
877 & \multicolumn{1}{l|}{\textbf{F1}} & \multicolumn{1}{l|}{\textbf{Precision}} & \multicolumn{1}{l|}{\textbf{Recall}} \\ \hline
878\textbf{In Detection} & 0.935 & 0.975 & 0.898 \\ \cline{1-1}
879\textbf{Out Detection} & 0.235 & 0.776 & 0.138 \\ \cline{1-1}
880\end{tabular}
881\end{table}
882
883\subsection{Fingerprint Stability Over Time}
884From the original participant group of 32 individuals, 12 were asked to repeat the test using the same hardware 5 weeks after the original sample. This was done in order to understand if the features of the ear change over time, or are highly dependent on the conditions the sample were taken in. The mean similarity between waveform of the same user was 85.2\% $\pm$ 2.6\%. This falls within the standard deviation of sample similarity calculated from the original test, showing that the fingerprint repeatability is stable. In other words, it was not some uncontrolled condition in the original experiment that lead to the high degree of similarity between samples. The similarity between fingerprints known not to match was 68.8\% $\pm$ 1.5\%. This again falls within the standard deviation of the same metric for the original test data. Combined, these metrics show that the system's repeatable ability to repeatably generate fingerprints containing meaningful features has not faded. Next, the stability of the fingerprints created was evaluated, by using the original test data as the enrolled sample. The system was then asked to identify the wearer using the new data as the live sample. This simulated the situation in which a user was enrolled, and then returned over a month later to be identified. The results can be seen in table \ref{laterid}. A small drop in accuracy was observed, but the accuracy of the system was still high, with hybrid F1 = 0.960.
885
886\begin{table}[h]
887\centering
888\caption{Identification Performance after 5 weeks}
889\label{laterid}
890\begin{tabular}{|l|lll}
891\hline
892 & \multicolumn{1}{l|}{\textbf{F1 Score}} & \multicolumn{1}{l|}{\textbf{Precision}} & \multicolumn{1}{l|}{\textbf{Recall}} \\ \hline
893\textbf{Histogram} & 0.870 & 0.769 & 1.000 \\ \cline{1-1}
894\textbf{Waveform} & 0.917 & 0.846 & 1.000 \\ \cline{1-1}
895\textbf{Hybrid} & 0.960 & 0.923 & 1.000 \\ \cline{1-1}
896\end{tabular}
897\end{table}
898
899
900
901
902This corresponds to an equal error rate of 4.0\%, and a correct classification rate of 96.0\% which is good when compared to the wearable gait analysis biometric system, which showed an equal error rate of between 5\% and 9\% (correct classification rate between 95\% and 91\%) \cite{Gafurov} when using a thresholding technique for samples recorded in the same session. Other vision based gait recognition techniques have shown correct classification rates of between 87\% \cite{Lam2005} and 100\% \cite{Hayfron-Acquah2003}, however both of these systems report on classification rates for data captured in the same session, and require external hardware that can capture video of a pedestrian, making them unsuitable for wearable specific applications.
903
904Verification performance was also strong despite the time gap. The results are shown in table \ref{laterver}. Interestingly, the histogram version of the system performed the best, with F1 = 0.937 despite the fact it performed the worst in the original test. This may suggest it is the more stable of the two fingerprinting methods. The performance of each method each threshold increment is shown in figure \ref{f1retrial}.
905
906\begin{table}[h]
907\centering
908\caption{Verification Performance after 5 weeks}
909\label{laterver}
910\begin{tabular}{|l|lll}
911\hline
912 & \multicolumn{1}{l|}{\textbf{F1 Score}} & \multicolumn{1}{l|}{\textbf{Precision}} & \multicolumn{1}{l|}{\textbf{Recall}} \\ \hline
913\textbf{Histogram} & 0.937 & 1.000 & 0.881 \\ \cline{1-1}
914\textbf{Waveform} & 0.867 & 1.000 & 0.769 \\ \cline{1-1}
915\textbf{Hybrid} & 0.870 & 1.000 & 0.769 \\ \cline{1-1}
916\end{tabular}
917\end{table}
918
919\begin{figure}[h]
920 \centering
921 \includegraphics[width=0.9\linewidth]{Figures/retrial-performance.png}
922 \caption{Thresholding F1 Score for each operational mode, per threshold in 0.1 increments.}
923 \label{f1retrial}
924\end{figure}
925
926\chapter{Conclusions and Future Work}
927\section{Conclusion}
928This work has presented an introduction into the potential utility of a novel kind of hardware - headphones suitable for consumer use with microphones on the inside and outside of the ear. This was approached in three main parts. The first section introduced the specifics of the hardware prototype being used and the tools used to capture data from the device. Then, an understanding of the nature of the data that could be collected passively from the microphones was built. This was motivated with research into the state of the art for wearable technology, and an overview of the features of the human ear - and in particular how they may interact with an acoustic sensing device. Armed with this understanding, several hypotheses were constructed, and for each an algorithm built in an attempt to test it. The first was that it is possible to detect if the user has the headphone in their ear or not, based on the signals from the external or external microphones. An algorithm is presented that attempts to make classifications based on three observed characteristics of the insertion and removal transition signal. This system showed promise but is ultimately a problem probably best suited to machine learning based classification. In order to perform this however, a much larger data set would be required for training and verification.
929
930The ability to detect the heartbeat in some users was demonstrated and verified with a small study comparing the signals to ground truth readings. An explanation of the source of the signal based on the anatomy of the human auditory system was given, and linked to existing work on the phenomenon. Two issues prevent this from being exploited further. The first is that the signal was not reliably detectable in many of the users tested. This appeared to be due at least in part to the user's anatomy, so it is not clear if there would be a way to rectify this by for example using a more sensitive microphone. A problem that a different microphone configuration may solve however is the system's susceptibility to noise. Any external noise or movement from the user would render the signal undetectable.
931
932The main work in this section presented two methods for building a fingerprint of a wearer's ear canal, that may have applications in biometric security and for other general identification purposes. These systems test the hypothesis that the morphology of an individual's ear canal is sufficiently different such that it can be detected by measuring its acoustic properties. Each functions by building a fixed length feature vector from the signal returned when a high frequency chirp is emitted into the wearer's ear canal. Key to their functioning is the ability to build a simple representation of the fingerprint, and to also be resistant to noise and other variations in samples, as no two recordings is ever sample for sample identical. The first method, the frequency density histogram, functions by separating the range of frequencies in the response into bucket and measuring their relative intensity. The second measures the rate of change in magnitude between segments of the signal, and assigns an ascending, descending or constant flag to that section, building a barcode like fingerprint. Both systems enable comparison between two fingerprints to generate a percentage similarity score.
933
934A key section of work that is accessory to the fingerprinting but vital for the overall function was the design of an algorithm that could automatically detect and isolate the response from the emitted chirp. This was particularly important as any system based on this biometric must be able to effectively isolate the area of interest so that the fingerprint is built from the correct data. Two algorithms are presented that serve to detect and isolate the fingerprinting chirp from the audio, with one acting as a lower precision but less noise sensitive backup to the first. As they form such an important part of the overall system, their performance was also evaluated in the results section.
935
936Two modes of the fingerprinting system were presented: identification and verification in line with standard biometric system operational modes. Each has the ability to work with the histogram or waveform vectors. A third method for merging both fingerprints to provide extra data points for identification or verification was also presented.
937
938A user study was designed and carried out on 32 participants in order to evaluate the hypotheses and the performance of the systems built to test them. It showed that the algorithms built to extract the fingerprints were effective, with the system acceptably isolating the chirps in 99.8\% of the test cases. The tests allowed an understanding of how well the algorithms were able to discriminate between two separate samples from the same user, and two samples from different users. The waveform fingerprint showed mean similarity of 85.64\% for matching fingerprints, and 68.08\% for non matching fingerprints. The histogram figures were 94.46\% and 84.84\% respectively.
939
940The fingerprinting system's identification ability was strong. When using the waveform vector to detect the user based on a single ear, the F1 score was 0.984. The error rates were found to be similar or better than existing biometric identification techniques, such as voice recognition and some fingerprinting systems. Combining the two fingerprint techniques further improved performance, and the system was able to make a correct identification of every user in the test when using the hybrid approach, yielding F1 = 1.0.
941
942Verification was also tested, and the results presented were somewhat counter intuitive. The waveform feature vector performed significantly better than either the histogram or hybrid system. This is opposite to the performance in verification, in which the hybrid was the best performing. Overall however, the waveform showed a lower false accept rate than many of the other biometric systems compared to, and similar performance for false rejections. The false reject rate was lower than the gait recognition system and was within the range of error rates given for the face and voice bimodal verification system discussed in the related work. Compared to a template protected fingerprinting system the performance was slightly better. This remained true for the same fingerprinting system without template protection. The system outperformed the published results of a voice recognition system based on GMMs \cite{Burget2007}.
943
944One area of weaker performance was resistance to noise, which was found to reduce the similarity between two matching fingerprints by 4.3\%. Improving the system's resistance to external sources of noise would be a worthwhile area of future work as it would add real value in terms of usability of the system.
945
946Overall, the work here presented two potential methods of biometric authentication and identification using a person's ear canal acoustic properties. The user study provided initial validation of its performance and showed that the waveform feature vector is a good method for verification, with a low false negative rate and an even lower false positive rate. The histogram feature vector did not score as well in verification, but when combined with the waveform vector it improved the identification performance of system. These results provide promising evidence for a biometric system that relies on a minor enhancement to an already ubiquitous consumer device with an inexpensive sensor. Further work however would be needed however, to fully asses the suitability of the system for protecting a wearable device with this biometric. This may include a larger scale user study, and the inclusion of cryptographic concepts to protect the data recorded as part of enrolment in the system. This is discussed in the Future Work section below.
947\section{Future Work}
948%Live recognition
949%Mobile device code
950%Template protection
951%Multimodal authentication
952There is significant room for additional work in this area, as this report served as initial exploration into the potential for in-ear wearables to create useful data, and to discuss how this could be used to add value for the user. Detecting the insertion/removal of the headphone from the ear would be an ideal application for machine learning techniques to implement classification. However, in order to do this a much larger data set would be required, needing data from many more participants and possibly more data points per user. A larger user study would also have the benefit of providing a larger data set over which to test the biometric systems, creating a stronger case for their performance and security in general. This is left to future work.
953
954The use of in ear acoustics for biometrics is promising, and has significant scope for further endeavour. One key area would be to create an implementation of the system that can accept a sample and produce an identification/verification result in near real time. Currently the speed of the algorithm was good - taking under 8ms per match on a commodity dual core laptop, but the system has not been implemented on a mobile system. This would also require translation of the codebase into a language supported by the mobile device from MATLAB, possibly Java for Android or Swift for iOS. None of the code exploits MATLAB exclusive features so the system can already be exported in C as a starting point for this process.
955
956Some of the biometric systems compared against in this work leverage template protection to ensure the security of the enrolment data \cite{Tuyls2005} \cite{Qin2018}. This is particularly important for biometrics, as in contrast to passwords which can be changed, once a biometric is compromised it is compromised permanently. Therefore ensuring the representation used to store the enrolment data can allow effective and efficient comparison with new fingerprints is a key component of an fully functional biometric system. This problem is more complex than simply adding a salt and hashing the enrolment data, as one would do with a stored password. This is because the enrolled and live data rarely match exactly, and hash functions will often produce wildly different outputs for a small change in input. Prior art on template protection for biometrics exists \cite{Jain2008} \cite{Yang2010} \cite{Breebaart2009}, so it may be possible to apply these techniques to the feature vectors used here.
957
958We have seen previously that it is possible to fuse biometrics together to create a system that has better performance than either of the biometrics operating in isolation \cite{Phil2012}. It would be possible to fuse the acoustic fingerprinting system here with some other form of biometric to yield better performance in the case of a genuine authentication attempt as more data is being captured to reduce the possibility of a false negative, but to also improve security as an imposter must forge two separate aspects of a victim's identity rather than just one. A suitable candidate for the secondary biometric may be speaker identification as the headphones contain two external microphones ideally placed for picking up the speech of the wearer.
959
960One of the major limitations of the system as it stands is the fact that even small movements from the wearer result in large amounts of noise recorded on the internal microphone track. This is due to the cable connection rubbing against the side of the wearer's head, and the noise caused by the friction travelling up the cable into the headphone. This problem could be mitigated with the use of a Bluetooth version of the headset. However, this would then introduce a new requirement that the amount of data transmitted is as small as possible, so as to reduce the impact on the battery in the device. Testing the algorithms built here with a smaller bit depth and lower bit rate would be a worthwhile study into the feasibility of running the systems over a Bluetooth wireless connection.
961
962% Wireless version
963% Wider user study.
964
965
966%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
967% Bibliography:
968
969% \cleardoublepage
970\phantomsection
971\addcontentsline{toc}{chapter}{Bibliography}
972\bibliographystyle{plainnat}
973\bibliography{thesis}
974
975
976
977%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
978%% Appendix:
979%%
980
981\appendix
982
983
984
985
986%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%
987%% Index:
988%%
989\printthesisindex
990
991\end{document}