2. Background: Bone Conduction Psychoacoustics

2.1 Air vs. Bone Conduction

The human auditory system perceives sound through two distinct pathways [1]:

1. Air conduction: Sound waves travel through the external ear canal, vibrate the tympanic membrane, and are

transmitted via the ossicular chain (malleus, incus, stapes) to the oval window of the cochlea. This pathway

preserves the full audible frequency spectrum (20 Hz – 20 kHz).

2. Bone conduction: Vibrations of the skull bones are transmitted directly to the cochlea, bypassing the outer and

middle ear. This pathway acts as a mechanical lowpass filter with a cutoff at approximately 4 kHz [1].

The internal voice perception is the sum of both pathways, while a recording captures only the air-conducted

component. This difference explains why recorded voice sounds unfamiliar — it lacks the bone-conducted low-

frequency energy and the occlusion-effect mid-frequency boost.

2.2 The Occlusion Effect

When the ear canal is blocked (as occurs naturally when the vocal tract radiates sound toward the skull), the bone-

conducted signal is enhanced in the low-to-mid frequency range. This is known as the occlusion effect [2]. The effect

peaks at approximately 2.8 kHz with a gain of 10–15 dB relative to open-ear conditions [3].

The occlusion effect has three contributing mechanisms: – Osseotympanic: Skull vibrations radiate into the ear canal

and pressurize the occluded volume – Inertial: The ossicular chain’s mass resists skull vibration, creating relative

motion at the oval window – Compressional: The cochlea is compressed and expanded by skull vibration, creating a

traveling wave

2.3 Skull Resonance

Bone conduction transfer functions exhibit a resonance in the 600–1000 Hz range, corresponding to the mechanical

resonance of the middle ear structures [4]. Stenfelt and colleagues measured bone conduction sensitivity and

demonstrated that:

  • Below 500 Hz, bone conduction is approximately 10 dB less sensitive than air conduction
  • Between 500 Hz and 4 kHz, the sensitivity gap narrows
  • Above 4 kHz, bone conduction sensitivity drops sharply

2.4 From Psychoacoustics to Computation

Our goal is to model this perceptual transformation computationally. By designing a digital filter that approximates

the bone conduction transfer function, we can simulate the internal voice from an external recording, then compare

the two using speaker embedding similarity

Discover Your Inner Voice: Psychoacoustic Style-Delta Extraction via Diffusion Based Voice Conversion

Abstract

This Blog presents “Discover Your Inner Voice,” a method for modeling the perceptual difference between how a

speaker hears their own voice (via bone conduction) and how it is heard by others (via air conduction). Using CAM++

speaker embeddings from the Seed-VC diffusion-based voice conversion pipeline, we extract a 192-dimensional style

delta vector ΔV = V_perceived − V_raw, where V_raw encodes the externally recorded voice and V_perceived

encodes the voice after psychoacoustic skull-resonance filtering. Our key finding — ‖ΔV‖ = 1.191 with cosine

similarity 0.9967 — reveals that CAM++ is inherently invariant to spectral filtering, a desirable property for speaker

verification but a limitation for perceptual style extraction. We discuss the implications, propose a filtered reference

override solution, and outline a roadmap toward clinical applications in laryngeal cancer rehabilitation.

1. Introduction

The question is as old as recorded audio: “Why does my voice sound different on recording?” Every speaker

experiences the disconnect between their internal self-perception and the external recording. This phenomenon arises

from a fundamental psychoacoustic difference: we hear our own voice through two simultaneous pathways — air

conduction through the ear canal and bone conduction through the skull.

While this discrepancy is a familiar curiosity, its computational modeling has remained unexplored. Modern voice

conversion systems can clone a speaker’s voice from a few seconds of reference audio, but they operate on the

external voice — the voice as heard by others. The internal voice — the voice as perceived by the speaker — is lost.

This project introduces the Inner Voice framework: a method to capture the transformation between a speaker’s

external and internal voice using speaker embeddings from a state-of-the-art voice conversion pipeline. We pose the

following research questions:

  1. Can a psychoacoustic skull-resonance filter produce a measurable shift in speaker embedding space?

2.Can this shift be encoded as a style delta vector (ΔV) for use in voice conversion?

3.Does the ΔV carry perceptually meaningful information about the internal voice?

The scope of this work is a proof-of-concept on a single subject, with a full implementation pipeline including a

Gradio web interface, real-time metrics, and a clinical application framework for laryngeal cancer rehabilitation.

Contributions: – A psychoacoustic skull-resonance filter modeling bone conduction effects – The ΔV = V_perceived

− V_raw style delta extraction pipeline – An override_style parameter for the Seed-VC voice conversion system – A

three-tab Gradio interface with real-time visualization – The experimental finding that CAM++ is filter-invariant to

spectral coloration, quantifying this with ‖ΔV‖ = 1.191 – Open-source release of all code, vectors, and documentation

The remainder of this paper is organized as follows. Section 2 reviews bone conduction psychoacoustics. Section 3

describes the Seed-VC architecture. Section 4 covers the CAM++ speaker encoder. Section 5 presents the proposed

ΔV extraction method. Section 6 details the system implementation. Section 7 reports results and analysis. Section 8

discusses a clinical application. Section 9 concludes with future work.

1.1 Related Work

Voice conversion has evolved through several generations. Early systems used Gaussian mixture models (GMMs) to

map source-to-target spectral envelopes [11]. Parallel data was required — the same utterance from both speakers.

The introduction of cycle-consistent adversarial networks (CycleGAN-VC) [12] enabled non-parallel training,

followed by variational autoencoders (VAEs) and autoencoder-based approaches that disentangled content from

speaker identity.The current state-of-the-art uses diffusion models and self-supervised learning. DiffWave [13], SpecDiff [14], and

VoiceBox [15] demonstrated that diffusion models produce high-fidelity speech. Seed-VC [5] achieved the first truly

zero-shot voice conversion competitive with speaker verification systems, reaching a speaker embedding cosine

similarity (SECS) of 0.867.

Concurrent work in speaker embedding analysis includes studies on embedding invariance [16], which showed that

verification embeddings are robust to channel effects. Our work extends this investigation to the specific case of bone-

conduction modeling.

No prior work has attempted to model the bone-conduction vs. air-conduction perceptual difference using speaker

embeddings. The Inner Voice concept is, to our knowledge, the first attempt to computationally capture the internal

voice perception.

Linz and Ars Electronica

Ars Electronica

Overall, I liked the Ars Electronica visit. The building itself is nice, the facilities are well designed, and Linz is already a beautiful city. I can imagine that during the Ars Electronica Festival it becomes a much more exciting place with more people, exhibitions, and events.

The project that caught my attention the most was Solar Synthesizer 0.4. I liked the idea of making electronic music with solar energy and controlling it just by changing the light with your hand. It is a simple concept, but I think it works well because it combines music, interaction, and sustainability. Even though I found the technology behind it a bit outdated compared to what is possible today, I still think the idea is interesting and creative.

My only criticism of the exhibition in general is that some of the technology and especially the AI-related installations felt a bit outdated. AI is developing so fast that I expected to see newer and more advanced examples. I think there are now other places in the world where more experimental and up-to-date technologies can be experienced.

Still, as a student studying sound and interaction design, I think it was definitely worth visiting. It is a good place to get inspiration and see different approaches to interactive media, even if not every installation feels cutting-edge anymore.

Paris Blog

Paris and IRCAM Excursion

Paris was beautiful. Since it was my first time seeing the city, I do not really have any complaints about the city itself. The museums we visited and the areas we explored were useful, carefully maintained and special places. It was also good to experience the atmosphere of Paris outside of only reading or hearing about it.

Naturally, the trip was very expensive. My criticism is more about the support around the excursion. I think the school did not provide enough support, and some of the support that was promised did not arrive in a timely way. For example, the IRCAM conference ticket was not a surprising or unexpected cost. It was a expected fee from the beginning, so I think the school could have at least covered that part. They did not, and that was frustrating. Still, in the end, I had the opportunity to see Paris, and I am happy about that.

Apart from the financial side, I liked visiting IRCAM. It is clearly a very established institution, and it has pioneered many things in the field of sound, music technology and computer music. It is also valuable that they hold this conference every year and bring people together around these topics.

At the same time, I also felt a little disappointed. As many of my friends noticed as well, IRCAM felt a bit outdated and overly commercialised in some parts. Maybe this is normal for a large institution with a long history, but I expected some of the work to feel more experimental and more forward-looking.

Some of the people I met were helpful, and some were not, but that is not the most important point for me. What bothered me more was that some presentations felt too outdated for an institution that sees itself as a frontrunner in audio technologies. For example, there was a presentation about audio in gambling. In short, the presenter sonified gambling data in Pure Data. Of course, the patch itself was fine, and I do not want to disrespect the work. But as a topic, it felt very simple and outdated for a place like IRCAM.

This is only my personal opinion, but I expected a stronger level of experimentation and technological relevance. Still, the excursion was useful overall. Seeing Paris, visiting IRCAM and observing both the strong and weak sides of such an important institution helped me understand the field more realistically.

Blog_10 Reflection

10. Current State and Personal Reflection

At this point, the project has reached a working prototype stage. It is not finished as a full thesis yet, but it is no longer only an idea or a sketch. There is now a real system that can be opened, configured and listened to. The listener can move through a virtual source layout, the system calculates direct-sound behaviour, and REAPER applies the values through the audio chain.

Looking back, the most important progress was not only technical. Of course, I learned a lot about Python, OSC, REAPER routing, mcfx, IEM plug-ins and binaural rendering. But maybe the bigger development was learning how to reduce the project to something I can actually defend. At the beginning, I wanted to work with physical sound superposition in a room. That idea is still important to me, but during the semester I understood that I first needed a controlled reference system.

The current prototype is therefore not trying to be everything. It does not model the full room. It does not include reflections, reverberation, diffraction, occlusion or source directivity. Instead, it focuses on the direct path between each source and the listener. This limitation is not only a weakness; it is also what makes the system understandable. I can explain what is calculated, what is sent to REAPER, and what is heard over headphones.

One useful lesson was that wording matters. If I call the system a room simulation, I create expectations that the prototype cannot fulfil. If I call it a direct-sound auralisation or a controlled reference model, the project becomes more honest. This was something I had to learn through feedback. Some parts of the system were already working, but my explanation was sometimes too broad or too confident. The feedback helped me clean that up.

The project also changed how I think about tools. I started with Pure Data and Max/MSP, but the final direction needed a custom interface and a more flexible calculation layer. PyGame and Python gave me that. REAPER then became the audio engine where the calculated values could be tested in a real plug-in chain. This combination is not always elegant, and sometimes it creates annoying technical problems, but it also gave me a system I understand from the inside.

The next step is to compare this controlled prototype with the physical room. That is where the original idea returns. The CUBE can show what the real room adds: reflections, room response, loudspeaker behaviour, interference and perception. I think this comparison is now the most interesting direction for the thesis. The prototype gives me the clean version; the physical room gives me the complicated version.

For me, the value of this semester is that the project became concrete. I now have a technical baseline, a clearer research direction and a better understanding of the limits. The system is not perfect, but it works, and more importantly, I can explain why it works the way it does. That feels like a good point to continue from next semester.

Blog_9 Future

9. Next Semester: Bringing the Prototype to the Physical Room

The next step is to bring the project back to the physical room. The current prototype gives an idealised direct-sound condition. The CUBE will add the things that are deliberately missing: reflections, reverberation, room response, loudspeaker behaviour, source directivity, modal effects and real listener perception.

The goal is not to prove that the auralisation is “better” or that the real room is “worse”. That would be the wrong question. The more interesting question is: what changes when the same source layout moves from a controlled headphone-based auralisation into a physical multi-loudspeaker environment?

To make this comparison meaningful, the next semester needs a clearer method. One possible approach is to create selected source layouts in the GUI, then recreate equivalent layouts in the CUBE. The same or similar sine-tone frequency sets can be used, such as 440 Hz, 444 Hz, 448 Hz and related intervals. This keeps the connection to the original superposition idea.

For measurement, I do not want to claim a complete protocol yet. But the basic plan would be to use fixed source and listener positions. Level could be measured with a measurement microphone at selected listener positions and compared with the calculated direct-sound attenuation. Delay or time of flight could be checked using short impulse-like test signals and looking at the first arrival. The later reflections would then show what the room adds beyond the direct path.

Perception will also matter. The system can predict level and delay, but the listener may experience the physical room differently because of reflections, beating, instability or localisation changes. This is where a listening-test method could become important. It should focus on movement, localisation, beating and perceived spatial stability.

The game/interface side project may or may not connect to the thesis later. During the semester, I became interested in the PyGame controller and even made a Doom-like game, later moving it to Godot. That is not the main thesis right now, but it influenced how I think about movement and interaction. Maybe there will be a connection later; maybe not. For now, the safest thesis direction is the comparison between controlled direct-sound auralisation and real-room behaviour.

The project has reached a working baseline. The next step is to test what this baseline can reveal when it meets the physical space again.

Blog_8 Feedback

8. Feedback, Corrections and a More Honest Framing

The second presentation and professor feedback were very useful because they exposed weak points in how I was explaining the project. Some parts were technically working, but the wording was not always careful enough. This was especially true around terms like simulation, room model, coordinate convention and perception.

One important correction was the word simulation. It is too strong for the current system. The prototype does not simulate a complete room. It does not include reflections, reverberation, diffraction, occlusion, source directivity or measured room response. A better term is direct-sound auralisation or controlled reference model. That makes the project more honest.

Another correction was the phrase room model. It can sound like the system models the room acoustics. It does not. It calculates source-listener geometry and direct-sound behaviour. So it is better to say Python-based direct-sound calculation script or calculation model, not full room model.

The coordinate drawing also needed correction. The angle mapping used in the system should not be presented as a universal coordinate convention. It is a plug-in angle mapping used to send direction values to the IEM plug-ins. In this mapping, front is 0 degrees, left is positive 90 degrees, right is negative 90 degrees and back is plus or minus 180 degrees. That is fine as an implementation mapping, but it should not be overgeneralised.

Some theoretical claims also had to be removed or softened. For example, distance-dependent HRTF changes and air absorption were not strong enough to present as central elements of the current prototype. Temperature-based sound speed also should not be mentioned casually unless the exact formula and assumptions are shown. These details can easily create questions that distract from the actual working system.

This feedback improved the project. It forced me to separate what is technically implemented from what could become future work. That is uncomfortable sometimes, but it is good for the thesis. The project becomes stronger when its limits are clear.

The current framing is now much better: the system is a working direct-sound auralisation baseline. It is controlled, testable and ready for comparison. It does not pretend to be the real room. Instead, it prepares a clean reference condition so that the real room can later be studied more clearly.

Blog_7 Endgame

7. REAPER, mcfx and IEM Ambisonics

The REAPER signal chain is where the calculated values become audible. Each source track has a practical plug-in chain. The first mcfx gain-delay instance generates the sine tone and applies the first gain stage. Then a custom JSFX called TEZ_Smooth_Propagation_Delay applies the time-of-flight delay in a smoothed way. After that, a second mcfx gain-delay instance applies extra gain reduction if it is needed.

The second mcfx instance exists because one gain stage has a limited controllable range in my OSC/REAPER setup. For example, the model might calculate that a source at 20 metres should be about minus 26 dB relative to 1 metre. If one mcfx stage can safely cover about minus 18 dB, the remaining minus 8 dB is applied by the second stage. The acoustic model still calculates one value. The two-stage split is only an implementation detail.

This was confusing at first, but the logic is simple: the model says how much quieter the sound should be. REAPER then needs to apply that level reduction. If one gain control cannot apply the full value, the reduction is split across two serial gain controls. In dB, serial gain values add together, so minus 18 dB plus minus 8 dB becomes minus 26 dB.

The delay side had another technical issue. If delay values are changed too abruptly while the listener moves, clicks or crackling can happen. That is why the custom smooth delay FX exists. It receives the calculated time-of-flight delay and changes it more smoothly, so movement sounds more stable.

After the source tracks, each signal is routed into the Ambisonic bus. The IEM MultiEncoder receives the source direction values. The IEM SceneRotator receives the listener head rotation. The EnergyVisualizer helps check the spatial energy visually, and the BinauralDecoder converts the Ambisonic scene into headphone playback.

This means the current prototype is headphone-based. The binaural decoder assumes that the left signal reaches only the left ear and the right signal reaches only the right ear. If the output is played over two loudspeakers, some things like level changes and beating can still be heard, but the binaural head-tracking result is not reliable because of acoustic cross-talk.

So the REAPER chain is not only playback. It is the part where the mathematical model becomes a listening experience.

Blog_6 Calculation Script

6. The Direct-Sound Calculation Model

The calculation part of the project is intentionally simple. It does not try to model a complete room. It calculates the direct path between each source and the listener. This is important because it keeps the system testable and avoids making claims that are too large.

The model calculates three main things: distance, level and time of flight. First, it calculates the difference between source and listener position on the x, y and z axes. These differences are dx, dy and dz. From these values, it calculates the straight-line distance between source and listener. This distance is called r.

Second, the model calculates level attenuation. The system uses 1 metre as the reference distance. At 1 metre, the level change is 0 dB. As the source gets farther away, the level decreases according to free-field pressure attenuation. This follows the relationship p proportional to 1/r. In dB, every doubling of distance gives about minus 6.02 dB. So 1 metre is 0 dB, 2 metres is about minus 6 dB, 4 metres is about minus 12 dB, and so on.

This is sometimes easy to misunderstand. It is not minus 6 dB per metre. It is minus 6 dB per distance doubling. That is why a 20 metre test position gives about minus 26 dB, not minus 120 dB.

Third, the model calculates time of flight. Sound does not arrive instantly. If a source is farther away, the sound arrives later. The delay is calculated as distance divided by the speed of sound. For example, at 20 metres, the time of flight is about 58 milliseconds.

There is also a practical limitation: the 1 metre reference clamp. Below 1 metre, the system does not increase the level further. This prevents unrealistic gain close to an idealised point source. However, this also means that if the goal were to simulate a very small studio or very close source-listener distances, this part would need to be reconsidered.

The direction calculation is also part of the model. The source position is converted into azimuth and elevation. Listener yaw is not added to the source azimuth. Instead, source direction goes to the MultiEncoder, and head rotation goes separately to the SceneRotator. This separation prevents double rotation.

For me, this is the core of the project: a simple model that is honest about what it does.