5. Proposed Method: ΔV Extraction

5.1 Filter Design Rationale

The three-stage filter design is based on a parametric model of the bone conduction hearing pathway. Each stage

models a distinct mechanical component:

Bone conduction transfer function (BCTF) is the sum of four contributing mechanisms: 1. Osseotympanic (15–

30%): Skull vibration radiates into the ear canal and is transmitted through the eardrum. This component is affected

by the occlusion effect. 2. Inertial (30–40%): The mass of the ossicular chain resists skull motion, creating relative

displacement at the oval window. Dominant around 1–3 kHz. 3. Compressional (25–35%): The cochlea is

compressed by skull vibration, creating a traveling wave. Dominant above 3 kHz with rapid rolloff. 4. Occlusion

effect (10–20%): The trapped volume of air in the occluded ear canal pressurizes, boosting low-to-mid frequencies.

Our filter approximates the combined BCTF using the minimum number of parametric stages:

5.2 Filter Stages

We model the bone conduction transfer function as a three-stage psychoacoustic filter. The parameters are derived

from the bone conduction literature [1, 3, 4].

Stage 1 — Lowpass (Bone Conduction Attenuation): – Type: 6th-order Butterworth lowpass – Cutoff frequency: 4

kHz – Q factor: 0.707 (Butterworth response) – Rationale: Bone conduction transfer functions show a sharp rolloff

above 4 kHz

Stage 2 — EQ Boost (Occlusion Effect): – Type: Peaking EQ (biquad) – Frequency: 2.8 kHz – Gain: +8 dB – Q: 1.8 –

Rationale: The occlusion effect peaks at approximately 2.8 kHz when the ear canal is occluded during phonation

Stage 3 — Peak Filter (Ear Canal Resonance): – Type: Peaking EQ (biquad) – Frequency: 800 Hz – Gain: +5 dB –

Q: 1.2 – Rationale: Middle-ear mechanical resonance and ear canal residual resonance in the 600–1000 Hz range

The filter is implemented using torchaudio.functional.biquad() functions running on CPU, as MPS (Apple

Silicon GPU) exhibits instability on long audio sequences with cascaded biquad operations.

5.3 Alternative Filter Designs Considered

We evaluated three alternative filter configurations during development:

A. Single lowpass (2-pole, 4 kHz): Produced a ΔV norm of 0.87, lower than the chosen design. The occlusion effect

modeling is essential for capturing the perceptual difference.

B. Graphic equalizer (10-band): A 10-band graphic EQ with bone-conduction target curve resulted in ‖ΔV‖ = 1.34,

but the filter introduced audible artifacts (phase distortion, pre-ringing) that compromised audio quality for

subsequent voice conversion.

C. Convolution with measured BCTF: Using published head-related transfer function (HRTF) measurements for

bone conduction [4] would be ideal but requires a measurement setup unavailable in our lab.

The chosen three-stage biquad design represents the optimal trade-off between psychoacoustic fidelity, computational

efficiency, and audio quality for voice conversion.


5.4 ΔV Pipeline

1.Record source audio — Single-channel recording at 22 kHz (144 seconds, Shure SM58 microphone, quiet

environment)


2.Extract V_raw — Pass recording through CAM++ encoder → 192-dim L2-normalized embedding


3.Apply skull resonance filter — Cascade three filter stages to the recording


4.Extract V_perceived — Pass filtered recording through CAM++ → 192-dim L2-normalized embedding


5.Compute ΔV — ΔV = V_perceived − V_raw

The resulting ΔV vector encodes the direction and magnitude of the perceptual transformation in CAM++ embedding

space.

5.3 Implementation Details

The extraction is implemented in extract_style.py and inference_delta.py: – CAM++ model loaded from

HuggingFace (funasr/campplus) – Audio truncated to 30 seconds for fast inference (Seed-VC processing bottleneck)

– 80-band FBANK features computed with 25 ms window, 10 ms hop – Filter applied at full sample rate,

downsampled to 16 kHz for CAM++ input

Leave a Reply

Your email address will not be published. Required fields are marked *