5.1 Filter Design Rationale
The three-stage filter design is based on a parametric model of the bone conduction hearing pathway. Each stage
models a distinct mechanical component:
Bone conduction transfer function (BCTF) is the sum of four contributing mechanisms: 1. Osseotympanic (15–
30%): Skull vibration radiates into the ear canal and is transmitted through the eardrum. This component is affected
by the occlusion effect. 2. Inertial (30–40%): The mass of the ossicular chain resists skull motion, creating relative
displacement at the oval window. Dominant around 1–3 kHz. 3. Compressional (25–35%): The cochlea is
compressed by skull vibration, creating a traveling wave. Dominant above 3 kHz with rapid rolloff. 4. Occlusion
effect (10–20%): The trapped volume of air in the occluded ear canal pressurizes, boosting low-to-mid frequencies.
Our filter approximates the combined BCTF using the minimum number of parametric stages:
5.2 Filter Stages
We model the bone conduction transfer function as a three-stage psychoacoustic filter. The parameters are derived
from the bone conduction literature [1, 3, 4].
Stage 1 — Lowpass (Bone Conduction Attenuation): – Type: 6th-order Butterworth lowpass – Cutoff frequency: 4
kHz – Q factor: 0.707 (Butterworth response) – Rationale: Bone conduction transfer functions show a sharp rolloff
above 4 kHz
Stage 2 — EQ Boost (Occlusion Effect): – Type: Peaking EQ (biquad) – Frequency: 2.8 kHz – Gain: +8 dB – Q: 1.8 –
Rationale: The occlusion effect peaks at approximately 2.8 kHz when the ear canal is occluded during phonation
Stage 3 — Peak Filter (Ear Canal Resonance): – Type: Peaking EQ (biquad) – Frequency: 800 Hz – Gain: +5 dB –
Q: 1.2 – Rationale: Middle-ear mechanical resonance and ear canal residual resonance in the 600–1000 Hz range
The filter is implemented using torchaudio.functional.biquad() functions running on CPU, as MPS (Apple
Silicon GPU) exhibits instability on long audio sequences with cascaded biquad operations.
5.3 Alternative Filter Designs Considered
We evaluated three alternative filter configurations during development:
A. Single lowpass (2-pole, 4 kHz): Produced a ΔV norm of 0.87, lower than the chosen design. The occlusion effect
modeling is essential for capturing the perceptual difference.
B. Graphic equalizer (10-band): A 10-band graphic EQ with bone-conduction target curve resulted in ‖ΔV‖ = 1.34,
but the filter introduced audible artifacts (phase distortion, pre-ringing) that compromised audio quality for
subsequent voice conversion.
C. Convolution with measured BCTF: Using published head-related transfer function (HRTF) measurements for
bone conduction [4] would be ideal but requires a measurement setup unavailable in our lab.
The chosen three-stage biquad design represents the optimal trade-off between psychoacoustic fidelity, computational
efficiency, and audio quality for voice conversion.
5.4 ΔV Pipeline
1.Record source audio — Single-channel recording at 22 kHz (144 seconds, Shure SM58 microphone, quiet
environment)
2.Extract V_raw — Pass recording through CAM++ encoder → 192-dim L2-normalized embedding
3.Apply skull resonance filter — Cascade three filter stages to the recording
4.Extract V_perceived — Pass filtered recording through CAM++ → 192-dim L2-normalized embedding
5.Compute ΔV — ΔV = V_perceived − V_raw
The resulting ΔV vector encodes the direction and magnitude of the perceptual transformation in CAM++ embedding
space.
5.3 Implementation Details
The extraction is implemented in extract_style.py and inference_delta.py: – CAM++ model loaded from
HuggingFace (funasr/campplus) – Audio truncated to 30 seconds for fast inference (Seed-VC processing bottleneck)
– 80-band FBANK features computed with 25 ms window, 10 ms hop – Filter applied at full sample rate,
downsampled to 16 kHz for CAM++ input