2.1 Air vs. Bone Conduction
The human auditory system perceives sound through two distinct pathways [1]:
1. Air conduction: Sound waves travel through the external ear canal, vibrate the tympanic membrane, and are
transmitted via the ossicular chain (malleus, incus, stapes) to the oval window of the cochlea. This pathway
preserves the full audible frequency spectrum (20 Hz – 20 kHz).
2. Bone conduction: Vibrations of the skull bones are transmitted directly to the cochlea, bypassing the outer and
middle ear. This pathway acts as a mechanical lowpass filter with a cutoff at approximately 4 kHz [1].
The internal voice perception is the sum of both pathways, while a recording captures only the air-conducted
component. This difference explains why recorded voice sounds unfamiliar — it lacks the bone-conducted low-
frequency energy and the occlusion-effect mid-frequency boost.
2.2 The Occlusion Effect
When the ear canal is blocked (as occurs naturally when the vocal tract radiates sound toward the skull), the bone-
conducted signal is enhanced in the low-to-mid frequency range. This is known as the occlusion effect [2]. The effect
peaks at approximately 2.8 kHz with a gain of 10–15 dB relative to open-ear conditions [3].
The occlusion effect has three contributing mechanisms: – Osseotympanic: Skull vibrations radiate into the ear canal
and pressurize the occluded volume – Inertial: The ossicular chain’s mass resists skull vibration, creating relative
motion at the oval window – Compressional: The cochlea is compressed and expanded by skull vibration, creating a
traveling wave
2.3 Skull Resonance
Bone conduction transfer functions exhibit a resonance in the 600–1000 Hz range, corresponding to the mechanical
resonance of the middle ear structures [4]. Stenfelt and colleagues measured bone conduction sensitivity and
demonstrated that:
- Below 500 Hz, bone conduction is approximately 10 dB less sensitive than air conduction
- Between 500 Hz and 4 kHz, the sensitivity gap narrows
- Above 4 kHz, bone conduction sensitivity drops sharply
2.4 From Psychoacoustics to Computation
Our goal is to model this perceptual transformation computationally. By designing a digital filter that approximates
the bone conduction transfer function, we can simulate the internal voice from an external recording, then compare
the two using speaker embedding similarity