2. Background: Bone Conduction Psychoacoustics

2.1 Air vs. Bone Conduction

The human auditory system perceives sound through two distinct pathways [1]:

1. Air conduction: Sound waves travel through the external ear canal, vibrate the tympanic membrane, and are

transmitted via the ossicular chain (malleus, incus, stapes) to the oval window of the cochlea. This pathway

preserves the full audible frequency spectrum (20 Hz – 20 kHz).

2. Bone conduction: Vibrations of the skull bones are transmitted directly to the cochlea, bypassing the outer and

middle ear. This pathway acts as a mechanical lowpass filter with a cutoff at approximately 4 kHz [1].

The internal voice perception is the sum of both pathways, while a recording captures only the air-conducted

component. This difference explains why recorded voice sounds unfamiliar — it lacks the bone-conducted low-

frequency energy and the occlusion-effect mid-frequency boost.

2.2 The Occlusion Effect

When the ear canal is blocked (as occurs naturally when the vocal tract radiates sound toward the skull), the bone-

conducted signal is enhanced in the low-to-mid frequency range. This is known as the occlusion effect [2]. The effect

peaks at approximately 2.8 kHz with a gain of 10–15 dB relative to open-ear conditions [3].

The occlusion effect has three contributing mechanisms: – Osseotympanic: Skull vibrations radiate into the ear canal

and pressurize the occluded volume – Inertial: The ossicular chain’s mass resists skull vibration, creating relative

motion at the oval window – Compressional: The cochlea is compressed and expanded by skull vibration, creating a

traveling wave

2.3 Skull Resonance

Bone conduction transfer functions exhibit a resonance in the 600–1000 Hz range, corresponding to the mechanical

resonance of the middle ear structures [4]. Stenfelt and colleagues measured bone conduction sensitivity and

demonstrated that:

  • Below 500 Hz, bone conduction is approximately 10 dB less sensitive than air conduction
  • Between 500 Hz and 4 kHz, the sensitivity gap narrows
  • Above 4 kHz, bone conduction sensitivity drops sharply

2.4 From Psychoacoustics to Computation

Our goal is to model this perceptual transformation computationally. By designing a digital filter that approximates

the bone conduction transfer function, we can simulate the internal voice from an external recording, then compare

the two using speaker embedding similarity

Discover Your Inner Voice: Psychoacoustic Style-Delta Extraction via Diffusion Based Voice Conversion

Abstract

This Blog presents “Discover Your Inner Voice,” a method for modeling the perceptual difference between how a

speaker hears their own voice (via bone conduction) and how it is heard by others (via air conduction). Using CAM++

speaker embeddings from the Seed-VC diffusion-based voice conversion pipeline, we extract a 192-dimensional style

delta vector ΔV = V_perceived − V_raw, where V_raw encodes the externally recorded voice and V_perceived

encodes the voice after psychoacoustic skull-resonance filtering. Our key finding — ‖ΔV‖ = 1.191 with cosine

similarity 0.9967 — reveals that CAM++ is inherently invariant to spectral filtering, a desirable property for speaker

verification but a limitation for perceptual style extraction. We discuss the implications, propose a filtered reference

override solution, and outline a roadmap toward clinical applications in laryngeal cancer rehabilitation.

1. Introduction

The question is as old as recorded audio: “Why does my voice sound different on recording?” Every speaker

experiences the disconnect between their internal self-perception and the external recording. This phenomenon arises

from a fundamental psychoacoustic difference: we hear our own voice through two simultaneous pathways — air

conduction through the ear canal and bone conduction through the skull.

While this discrepancy is a familiar curiosity, its computational modeling has remained unexplored. Modern voice

conversion systems can clone a speaker’s voice from a few seconds of reference audio, but they operate on the

external voice — the voice as heard by others. The internal voice — the voice as perceived by the speaker — is lost.

This project introduces the Inner Voice framework: a method to capture the transformation between a speaker’s

external and internal voice using speaker embeddings from a state-of-the-art voice conversion pipeline. We pose the

following research questions:

  1. Can a psychoacoustic skull-resonance filter produce a measurable shift in speaker embedding space?

2.Can this shift be encoded as a style delta vector (ΔV) for use in voice conversion?

3.Does the ΔV carry perceptually meaningful information about the internal voice?

The scope of this work is a proof-of-concept on a single subject, with a full implementation pipeline including a

Gradio web interface, real-time metrics, and a clinical application framework for laryngeal cancer rehabilitation.

Contributions: – A psychoacoustic skull-resonance filter modeling bone conduction effects – The ΔV = V_perceived

− V_raw style delta extraction pipeline – An override_style parameter for the Seed-VC voice conversion system – A

three-tab Gradio interface with real-time visualization – The experimental finding that CAM++ is filter-invariant to

spectral coloration, quantifying this with ‖ΔV‖ = 1.191 – Open-source release of all code, vectors, and documentation

The remainder of this paper is organized as follows. Section 2 reviews bone conduction psychoacoustics. Section 3

describes the Seed-VC architecture. Section 4 covers the CAM++ speaker encoder. Section 5 presents the proposed

ΔV extraction method. Section 6 details the system implementation. Section 7 reports results and analysis. Section 8

discusses a clinical application. Section 9 concludes with future work.

1.1 Related Work

Voice conversion has evolved through several generations. Early systems used Gaussian mixture models (GMMs) to

map source-to-target spectral envelopes [11]. Parallel data was required — the same utterance from both speakers.

The introduction of cycle-consistent adversarial networks (CycleGAN-VC) [12] enabled non-parallel training,

followed by variational autoencoders (VAEs) and autoencoder-based approaches that disentangled content from

speaker identity.The current state-of-the-art uses diffusion models and self-supervised learning. DiffWave [13], SpecDiff [14], and

VoiceBox [15] demonstrated that diffusion models produce high-fidelity speech. Seed-VC [5] achieved the first truly

zero-shot voice conversion competitive with speaker verification systems, reaching a speaker embedding cosine

similarity (SECS) of 0.867.

Concurrent work in speaker embedding analysis includes studies on embedding invariance [16], which showed that

verification embeddings are robust to channel effects. Our work extends this investigation to the specific case of bone-

conduction modeling.

No prior work has attempted to model the bone-conduction vs. air-conduction perceptual difference using speaker

embeddings. The Inner Voice concept is, to our knowledge, the first attempt to computationally capture the internal

voice perception.

Harmonix Series: Accessible Digital Musical Instruments for Mindfulness and Creativity


The “Harmonix Series” by Wing Hei Cheryl Hui and Patrick Hartono represents a visionary bridge between human-computer interaction and therapeutic art, moving beyond the technical novelty of Digital Musical Instruments to address a profound need for the democratization of creativity. What I find most compelling about this work is its commitment to radical accessibility; by shifting the focus from mastering a complex tool to simply experiencing a soundscape, the authors empower users of all physical and cognitive abilities to become creators. This is further elevated by the intentional integration of mindfulness into the interface design, which transforms the act of music making into a meditative process for emotional regulation rather than just a performance. The synergy between robust technical implementation and a sophisticated aesthetic sensibility is palpable, resulting in an instrument that doesn’t just function, but truly resonates on a human level. Ultimately, this project serves as a vital reminder that the future of music technology should prioritize deep human impact and digital wellbeing, treating the user as a whole person seeking connection and calm.

How an Immersive (3D) Mixing Bottleneck Inspired Me

After a decade of producing electronic music, I understood that a great mix is about creating a sense of space. When I transitioned into object-based audio (like Dolby Atmos), I encountered a paradox: While I could place a vocal perfectly in 3D space using precise coordinates, the next step getting the reverb right felt like stepping back into the Stone Age.

The Bottleneck of Spatial Incoherence

In spatial mixing, we use digital metadata (x,y,z) to define an object’s location. However, to make that object sound physically plausible far away, close, or high up the Reverb Send Level must be manually adjusted to match that coordinate.

I discovered that this manual adjustment was the core problem:

  1. Inconsistency: What sounds right in Scene A might be wrong in Scene B. Maintaining spatial realism across a two hour film or album tracklist became a massive, repetitive task.
  2. Subjectivity: The decision relied entirely on my ears, not physics. This meant spending creative hours tweaking parameters that should be calculated, not guessed.

The Revelation: If the “correct” reverb level is a function of the sound’s position and the room’s acoustics, it is an objective, solvable problem.

The Solution Concept: We could use AI (Deep Learning) to perform this complex, repetitive DRR calculation instantly. My thesis idea was born: Train an AI to map a vocal’s characteristics and its 3D position directly to the DRR-derived Reverb Send Level.

This project is the culmination of my journey turning creative frustration into a rigorous technical solution that aims to inject speed, consistency, and scientific backing into the art of spatial vocal mixing.

Crossing the Bridges

“Crossing the bridges” can refer to a specific film, a musical album, a metaphor for spiritual or professional journeys, or a proverb about not worrying about future problems. The meaning depends entirely on the context, such as a 1992 drama film, a 2013 film about returning to a village, a 2005 documentary about Istanbul, a musical album, or the idiom “don’t cross that bridge until you come to it”

 My name is Meriç, and I’m a sound designer, music producer, and electronic dance music performer  known as numeric.

 I started making electronic music about eight-ten years ago, and over time it turned into something much deeper a lifelong search for sound, emotion, and space.  

Music and sound have always been central to my life. This passion led me to explore questions like how to create sound and arrange electronic music compositions. This curiosity drove me to research more, investigate deeply, and develop a passion for music and sound. I started my journey as a DJ in 2014, performing at various venues. Through this experience, I realized the need to create compositions that could convey my stories through sound. Over the years, I have immersed myself in music composition and production, constantly striving to enhance my sound and musical quality. I have learned sound design, production, effect techniques, mixing techniques, arranging, and experimented with various hardware/software synths.

My passion for music, especially electronic music composition, led me to pursue a Master’s degree in Music at Istanbul Technical University . During my studies, I honed my skills in sound design, mixing, and mastering techniques, solidifying my desire to build a career in the sound design area. My time at ITU provided a rich exploration of the audio realm. I deepened my understanding of musical structure and harmony through music theory courses and explored sound manipulation through audio programming . I gained insights into contemporary music trends and techniques through my coursework. I also had to meet with great professors, artists and well minded fellow students.And I think most important I discovered how interesting and at the same time beautiful to research what you are passionate about .

 Sound design, for me, is both a craft and an inquiry. I spend countless hours refining my techniques, creating unique pieces and constructing custom systems for performance and media art. This personal research runs parallel to my artistic output, which lies at the intersection of electronic music, experimental audio, and immersive experience.I’m fascinated by how sound behaves in space, how it moves around the listener, and how technology can make those experiences more alive.Studying sound design gives me the tools to connect artistic ideas with technical reality.It helps me understand how sound works, but also how to use it to create emotion, meaning, and atmosphere.In my creative work, sound design is not separate from music Its the biggest part of the music.It’s the texture, the emotion, a tool which shapes a raw material.I see it as a way to build worlds, whether in a live performance, movie, theatre  installation, or interactive setting or advertisement. In the future, I want to continue my journey in sound design on both an academic and an artistic level, developing my practice professionally through research, creative projects, and collaborations.In the next two semesters, I would like to focus on the intersection of creative sound design, spatial audio, and intelligent systems .During my previous master’s studies, I worked on projects related to immersive and 3D sound, and that’s where my real passion for this field began to grow.

That experience deeply shaped my artistic direction and was one of the main reasons I decided to join the Sound Design program at IEM, a place that, in my opinion, holds a tremendous position in the field of immersive sound.

Now, I want to take this interest to a much higher level, both creatively and conceptually.

 In the coming semesters, I plan to explore how immersive sound can be used to create emotionally powerful and spatially engaging experiences, bridging artistic expression and technological research.