arXiv is now an independent nonprofit! Learn more
License: CC BY 4.0
arXiv:2609.29203v1 [eess.AS] 24 Sep 2026
MAE
mean absolute error
CNN
convolutional neural network
CRNN
convolutional recurrent neural network
DoA
direction of arrival
DNN
deep neural network
DRR
direct-to-reverberant ratio
ELU
exponential linear unit
GMM
Gaussian mixture model
GRU
gated recurrent unit
IID
interaural intensity difference
ITD
interaural time difference
LSTM
long short-term memory
LPC
linear predictive coding
MSE
mean squared error
MLP
multilayer perceptron
RIR
room impulse response
RNN
recurrent neural network
SELD
sound event localization and detection
SNR
signal-to-noise ratio
SVM
support vector machines
STFT
short-time fourier transform

Few-Shot Calibration for Sim-to-Real Single-Channel Speaker Distance Estimation

Michael Neri     Archontis Politis     Tuomas Virtanen 
Abstract

Speaker distance estimators are trained almost exclusively on simulated room acoustics, because real recordings annotated with the true talker-to-microphone distance are scarce. We show that models trained this way transfer poorly. On three real corpora we evaluate, simply predicting the average distance of the corpus is more accurate than any learned model. Then, we ask how few labelled real utterances are needed to make a frozen, synthetic-trained estimator useful, and study post-hoc calibration maps that rescale its output without gradients or retraining. An analysis of the achievable error shows that what the calibration is not limited by the absolute accuracy of the estimator, but how well it orders utterances by distance, since a constant bias or a wrong output scale is removed exactly by the calibration itself. Balancing this against the cost of estimating each coefficient from few samples yields a criterion that accounts for which map wins on which corpus and at which annotation budget, together with a shrinkage variant that requires no hard decision. Our findings suggest selecting synthetic checkpoints by linear correlation with true distances rather than by absolute error. Code, datasets, and analysis are available at https://github.com/michaelneri/audio-distance-estimation.

Index Terms: 
Acoustics, Speaker Distance Estimation, Room Impulse Response, Multi-task learning, Sim-to-Real Transfer.
††address: Faculty of Information Technology and Communication Sciences, Tampere University, Finland
{michael.neri, archontis.politis, tuomas.virtanen}@tuni.fi

1 Introduction

Estimating the distance between a talker and a microphone from a single microphone is of practical interests for several applications, ranging from hearing aid devices [1], hands-free communication [2], and speech recognition [3]. The cue is intrinsically acoustic: the direct-to-reverberant ratio, the early-reflection pattern and the spectral envelope all vary systematically with distance [4, 5, 6]. When these pieces of information are degraded, both human and machines distance perception deteriorate considerably [7, 8].

Early approaches to the more generic source distance estimation focused on binaural recordings and hand-crafted features derived from the DRR (DRR) [9] or from inter-channel correlations. These features were used to fit GMM and SVM (SVM) classifiers to distinguish distance between discrete bins [10, 11, 12]. More recent studies revisited the problem with DNN, yet most remain limited to binary far/near classification [13, 14] or to coarse distance bins over a narrow range [8, 15, 16]. Yet, these methods typically required careful hyperparameter tuning and generalised poorly across different acoustic environments.

In fact, the practical obstacle is the availability of audio data with labeled distances. The research community has therefore converged on training with synthetic data. The work in [17] was the first to frame speaker distance as a continuous regression problem, demonstrating that a DNN operating on STFT (STFT) phase features can achieve centimeter-level accuracy in simulated environments. It is worth noting that the centimeter-level accuracy reported in [18, 17] was obtained on simulated data with time- and level-calibration. Similarly, in [19] the authors proposed a data augmentation approach for generating realistic room-acoustic conditions to support speaker distance estimation using the CRNN (CRNN) model in [18]. Most directly relevant to this work, in [4] it was shown that learning-based single-channel distance estimators rely predominantly on early reflections rather than other RIR (RIR) components, but it did not study the sim-to-real transfer for speaker distance estimation without the knowledge of the time calibration. The only result available in this scope is in [18], which provided near-chance results on zero-shot experiments on three real corpora with a simulation-based estimator.

To mitigate this sim-to-real gap, the contributions of this work are: (i) we carried out zero-shot speaker distance estimation analysis, analyzing how much mismatch is present when training on synthetic audios and testing on real data. (ii) We devised a few-shot calibration that adapts a frozen simulation-based estimator to a new corpus with NN labelled utterances by fitting a two-parameter affine map, requiring no gradients and no retraining. (iii) Based on [4], we modified the head of the CRNN to predict acoustic-based features of the enclosure. While this leaves simulation-based MAE (MAE) largely unchanged, it substantially improves correlation with true distance labels under domain shift, which is what governs how well the estimator can be calibrated.

2 Few-shot calibration

2.1 Problem statement

Let fθf_{\theta} be a distance estimator trained exclusively on synthetic data 𝒟s\mathcal{D}_{s} with learned parameters θ\theta and kept frozen. Deployed on a target corpus 𝒟t\mathcal{D}_{t} it predicts an uncalibrated distance u=fθ​(x)u=f_{\theta}(x), nominally in meters from the recording x∈ℝTsx\in\mathbb{R}^{T_{s}} with TsT_{s} samples. We assume access to a small calibration set 𝒞N={(ui,di)}i=1N\mathcal{C}_{N}=\{(u_{i},d_{i})\}_{i=1}^{N} drawn from the training split of that corpus, with NN of the order of a few tens of samples, and find a map g:ℝ+→ℝ+g:\mathbb{R}^{+}\!\to\mathbb{R}^{+} producing a calibrated estimate d^=g⁡(u)\hat{d}=g(u), requiring neither gradients, nor access to θ\theta, nor retraining.

2.2 Foundations of distance calibration

Two systematic effects displace uu from dd. First, a regressor trained under a squared-error objective shrinks towards the prior mean of 𝒟s\mathcal{D}_{s}, whereas 𝒟t\mathcal{D}_{t} concentrates its distances elsewhere, causing an offset. Second, since the amplitude cue contributes negligibly [4], the estimator maps early-reflection structure to distance under the room-geometry and absorption statistics seen in training; a target enclosure whose reverberant statistics differ therefore induces a multiplicative scale error. Both can be corrected by the affine map

d^=a​u+b,\hat{d}=a\,u+b, (1)

with (a,b)(a,b) estimated by least squares on 𝒞N\mathcal{C}_{N}, without any domain adaptation or fine-tuning.

The constant predictor d^=d¯\hat{d}=\bar{d} is the reference any calibration must beat: it consumes the same NN labels, estimates the average d¯=𝔼⁡[𝒞N]\bar{d}=\mathbb{E}[\mathcal{C}_{N}], and ignores fθf_{\theta} entirely. Constraining either coefficient of (1) gives two one-parameter members of the same family, offset-only (a≡1a\!\equiv\!1), which corrects the prior shift alone, and scale-only (b≡0b\!\equiv\!0), which corrects the scale alone. Both trade modelling bias for estimation variance.

2.3 Estimating the affine mapping

Let μd,σd\mu_{d},\sigma_{d} and μu,σu\mu_{u},\sigma_{u} denote the target-corpus label and raw predicted distance statistics, and ρ\rho the Pearson correlation between uu and dd. Theoretically with unlimited number of samples (N→+∞N\rightarrow+\infty), the coefficients of (1) that minimise the population risk R⁡(a,b)=𝔼⁡[(d−d^)2]=𝔼⁡[(d−a​u−b)2]R(a,b)=\mathbb{E}[(d-\hat{d})^{2}]=\mathbb{E}[(d-au-b)^{2}] are obtained by solving ∂R⁡(a,b)/∂a=0\partial R(a,b)/\partial a=0 and ∂R⁡(a,b)/∂b=0\partial R(a,b)/\partial b=0, yielding the optimal coefficients

b⋆=μd−a⋆​μu,a⋆=ρ​σdσu,b^{\star}=\mu_{d}-a^{\star}\mu_{u},\qquad a^{\star}=\rho\,\frac{\sigma_{d}}{\sigma_{u}}, (2)

Substituting back to R⁡(a,b)R(a,b) yields after some calculations

𝔼⁡[(d−d^)2]=σd2​(1−ρ2).\mathbb{E}\big[(d-\hat{d})^{2}\big]=\sigma_{d}^{2}\,(1-\rho^{2}). (3)

It is worth noting that ρ\rho is invariant to affine transformations of uu, so a constant bias or a wrong output scale leaves (3) unchanged. What remains is how well the model correlates with distances within the target corpus. A model with large zero-shot MAE and high ρ\rho calibrates better than an accurate but unordered one. With finite NN, (a,b)(a,b) are themselves estimated from 𝒞N\mathcal{C}_{N}, and their sampling error adds to the test risk. An affine map (p=2p=2 free parameters) has expected test risk

RNaff≈R⁡(a⋆,b⋆)​(1+2N)=σd2​(1−ρ2)​(1+2N),R_{N}^{\mathrm{aff}}\;\approx\;R(a^{\star},b^{\star})\Big(1+\frac{2}{N}\Big)=\sigma_{d}^{2}\,(1-\rho^{2})\Big(1+\frac{2}{N}\Big), (4)

The constant predictor is the special case a=0a=0, so the same reasoning applies with p=1p=1,

RNconst≈R⁡(0,μd)​(1+1N)=σd2​(1+1N).R_{N}^{\mathrm{const}}\;\approx\;R(0,\mu_{d})\Big(1+\frac{1}{N}\Big)=\sigma_{d}^{2}\Big(1+\frac{1}{N}\Big). (5)

Having both the risk of the affine and the constant calibrations, we can estimate when it is worth fitting a slope, i.e., when RNaff<RNconstR_{N}^{\mathrm{aff}}<R_{N}^{\mathrm{const}}, which yields the inequality

ρ2>1N+2.\rho^{2}\;>\;\frac{1}{N+2}. (6)

Using this inequality which includes the number of calibrating samples and the correlation between ground-truth and predicted distances of the target corpus, we can set the threshold in function of the cardinality of the calibration set, i.e. ρ>0.38\rho>0.38 at N=5N\!=\!5, 0.290.29 at N=10N\!=\!10 and 0.210.21 at N=20N\!=\!20. When the inequality does not hold, the least-squares slope is dominated by its own sampling noise.

Offset-only admits an analogous test. Shifting the estimator without rescaling it leaves the score spread intact, so its risk is R⁡(1,μd−μu)=σd2+σu2−2​ρ​σd​σuR(1,\mu_{d}-\mu_{u})=\sigma_{d}^{2}+\sigma_{u}^{2}-2\rho\sigma_{d}\sigma_{u}. As it estimates a single parameter, like the constant predictor, both carry the same variance penalty and the (1+1/N)(1+1/N) factors cancel; the comparison therefore involves no NN at all and reduces to

σu/σd< 2​ρ,\sigma_{u}/\sigma_{d}\;<\;2\rho, (7)

which holds only when the estimator under-disperses relative to the target labels.

It is worth noting that the correlation between the estimator and the calibration set is itself estimated from few samples, and thus noisy at small NN. We denote this estimate ρ^\hat{\rho}. Then, we propose an additional mapping that avoids a hard threshold on (6) and instead shrinks the slope in proportion to the evidence for it [20],

a^λ=ρ^2ρ^2+1/N​a^,b^λ=d¯−a^λ​u¯,\hat{a}_{\lambda}=\frac{\hat{\rho}^{2}}{\hat{\rho}^{2}+1/N}\;\hat{a},\qquad\hat{b}_{\lambda}=\bar{d}-\hat{a}_{\lambda}\,\bar{u}, (8)

which reduces to the constant predictor as ρ^→0\hat{\rho}\!\to\!0, restores (1) as N→∞N\!\to\!\infty, and halves the slope at ρ^2=1/N\hat{\rho}^{2}=1/N, that is a smoothed form of (6).

3 Materials

3.1 Distance estimator

We use as distance estimator the convolutional-recurrent regressor of [17] as the baseline. It processes STFT log-magnitude and sine/cosine phase features, a three-block CNN (CNN) with max/average pooling, a two-layer bidirectional GRU (GRU), and a per-frame distance head pooled over time. Input is 1010 s of mono audio at 1616 kHz. We additionally evaluate a full-stack variant that adds, over the baseline, (i) auxiliary regression heads that predict log⁡T60\log T_{60}, log⁡Tmix\log T_{\mathrm{mix}}, and log⁡V\log V from the mean-pooled recurrent features, adding roughly 100100k learnable parameters. Here TmixT_{\mathrm{mix}} is the mixing time, the boundary between early and late reflections in the RIR, obtained from the echo-density measure of [21]. These parameters are chosen because they summarise the reverberant statistics that carry the distance cue; (ii) SpecAugment [22] (22 time masks, 22 frequency masks) applied after feature extraction; and (iii) waveform augmentation (polarity inversion, ±6\pm 6 dB gain, ±5\pm 5 ms shift, each with probability 0.50.5). Both variants are trained with Adam at 10−310^{-3} under two regimes: clean-trained, and noise-trained with WHAM! [23] noise mixed in at randomised SNR (SNR) spanning uniformly 0−500-50 dB. Following [18], both the clip-level prediction y^\hat{y} and the per-frame prediction 𝐲^t\hat{\mathbf{y}}_{t} are supervised, which encourages the recurrent layers to produce sharper per-frame estimates, giving

ℒ=12​[ℓ⁡(y^,y)+ℓ⁡(𝐲^t,𝐲t)]+λ​∑k∈𝒦ℓ⁡(u^k,log⁡k),\mathcal{L}=\tfrac{1}{2}\bigl[\ell(\hat{y},y)+\ell(\hat{\mathbf{y}}_{t},\mathbf{y}_{t})\bigr]+\lambda\!\!\sum_{k\in\mathcal{K}}\!\ell(\hat{u}_{k},\log k), (9)

where ℓ\ell is the MSE (MSE), 𝒦={T60,Tmix,V}\mathcal{K}=\{T_{60},T_{\mathrm{mix}},V\} collects the auxiliary targets with predictions u^k\hat{u}_{k} and weight λ=0.3\lambda=0.3, and y∈{log⁡d,d}y\in\{\log d,d\} depending on whether the log-distance target is enabled. Auxiliary terms are dropped when the corresponding head is disabled in an ablation. All variants use five-fold cross-validation, giving five checkpoints each.

3.2 Synthetic dataset

We employ the same uncalibrated dataset (neither time- nor amplitude-calibration) as in [4]. Specifically, anechoic speech recordings obtained from the EARS dataset [24] are convolved with the simulated RIR from pyroomacoustics [25]. The experiments include 25002500 audio files of 1010 s duration at 1616 kHz. The samples are randomly assigned to 55 folds to assess the performance in a 55-fold cross-validation fashion. By doing so, each cross-fold iteration assigns 15001500, 500500, and 500500 audios to training, validation, and testing sets, respectively. We follow the five-fold cross-validation protocol of [4]: for each fold i∈{0,…,4}i\!\in\!\{0,\dots,4\}, fold ii is held out for testing, fold (i+1)mod5(i{+}1)\bmod 5 for validation, and the remaining three folds are used for training. The simulated rooms span a wide acoustic range: source-to-microphone distances from 1.001.00 to 11.0011.00 m (mean 5.85.8 m), reverberation times T60T_{60} from 0.280.28 to 2.162.16 s (mean 0.770.77 s), room volumes from 8787 to 883​m3883\,\text{m}^{3} (mean 441​m3441\,\text{m}^{3}), and mixing times TmixT_{\text{mix}} from 8484 to 166166 ms (mean 129129 ms).

3.3 Real datasets

We evaluate sim-to-real transfer on three corpora spanning two levels of realism. VoiceHome2 [26] which encompasses smart-home commands recorded by twelve speakers in twelve rooms across four houses, under quiet and noisy conditions with uncontrolled domestic interferers (competing talkers, TV, appliances) and no SNR annotation. Five source positions per room, standing and sitting, with a fixed 8-microphone MEMS array on a cubic baffle where we use the first channel only. In total, the dataset encompasses 752752 recordings of 1010 s, with distances spanning 1.01.0-4.54.5 m (mean 2.272.27 m). STARSS23 [27]. Multi-speaker interaction scenes recorded at Tampere University and Sony in eleven rooms with an Eigenmike array, from which a single omnidirectional channel is extracted. We use 29342934 single-speech excerpts that do not overlap with other annotated directional sources. The corpus is the most challenging of the three: speakers move and change orientation, and diffuse and directional ambient noise is present at significant levels. Distances span 1.51.5–2.92.9 m (mean 2.19 m). QMUL-TIMIT [18]. Measured omnidirectional RIR captured in three rooms at Queen Mary University of London: a 7.5×9×3.5 m classroom (236236 m3, 130130 RIR), the Octagon, an eight-walled Victorian hall with a 2121 m domed ceiling (≈9500\approx 9500 m3, 169169 RIR), and the Great Hall (169169 RIRs over a 12×12 m region). Each RIR is convolved with 55 anechoic TIMIT utterances, yielding 23402340 recordings, with RIR split 7070/1010/2020 into training/validation/testing. We report clean and 00 dB conditions with WHAM! noise, matching the synthetic protocol.

4 Results

4.1 Results on synthetic data

Table 1 reports the MAE at the clean and 00 dB extremes. Training the baseline with noise already recovers most of the 00 dB degradation (1.83→1.471.83\to 1.47 m); on top of the noise-aware regime the multi-task and augmentation stack gives a smaller further gain (1.47→1.411.47\to 1.41 m) while also improving the clean condition (1.30→1.231.30\to 1.23 m). The per-component contributions are not monotonic: the log-distance target alone significantly increases 00 dB MAE under clean training (p<0.05p<0.05), and neither the multi-task heads nor SpecAugment alone produce a significant change. A plausible explanation is that the log-distance loss reweights gradients toward small distances, which sharpens the target but offers no benefit without the implicit regularisation that noise diversity provides; under noise-aware training the same target is harmless. Only the full stack yields a significant improvement in both regimes.

Table 1: Test average MAE in meters with 95%95\% confidence half-widths (5 cross-validation folds, tt-distribution) at the two extreme SNR conditions. Best per column in bold. Markers give the paired tt-test against the baseline at 00 dB: ∗p<0.05{}^{*}p<0.05, p∗⁣∗<0.01{}^{**}p<0.01. The Random row predicts the training-set mean, serving as a sanity check.

Clean-trained Noise-trained Model clean 00 dB clean 00 dB Baseline [4] 1.30 ±\pm0.08 1.83 ±\pm0.17 1.29 ±\pm0.06 1.47 ±\pm0.11 + log-distance target 1.38 ±\pm0.13 2.01 ±\pm0.14∗ 1.31 ±\pm0.09 1.46 ±\pm0.13 + multi-task heads 1.35 ±\pm0.06 1.92 ±\pm0.19 1.34 ±\pm0.12 1.51 ±\pm0.10 + SpecAugment 1.32 ±\pm0.13 1.78 ±\pm0.20 1.29 ±\pm0.12 1.49 ±\pm0.11 + speech aug. (full stack) 1.23 ±\pm0.13 1.69 ±\pm0.13∗∗ 1.23 ±\pm0.09 1.41 ±\pm0.11∗ Random 2.47 ±\pm0.06

4.2 Zero-shot real data analysis

We revisit zero-shot analysis proposed in [18] using the same two real corpora (VoiceHome2 [26] and STARSS23 [27]) and the hybrid QMUL-TIMIT [18] to analyze the sim-to-real gap without any fine-tuning. For the comparison, we report two baselines that draw the label-prior floor. The first, Random (synth.), predicts the synthetic training-set mean for every sample, without requiring any real-data labels. The second, Random (real), predicts the training-split label mean of each target corpus. However, it is worth noting that the Random (real) baseline is not strictly zero-shot as it requires label access to the target domain.

Table 2: Zero-shot real-data test MAE in meters (mean [l​o,h​i][lo,\,hi] 95 % CI) and Pearson correlation ρ\rho per corpus, across five fold checkpoints (∗p<0.05{}^{*}p<0.05, ∗∗∗p<0.001{}^{***}p<0.001). QMUL-TIMIT (real measured RIR) is reported clean and at 00 dB SNR. Random baselines are constant predictors (ρ=0\rho{=}0 by definition). Best learned model per corpus in bold.

Model VoiceHome2 STARSS23 QMUL-TIMIT clean 0 dB Baseline, clean-trained 2.38​[1.66, 3.11]2.38\;{\scriptstyle[1.66,\,3.11]} 3.07​[2.28, 3.86]3.07\;{\scriptstyle[2.28,\,3.86]} 5.97​[5.36, 6.58]\mathbf{5.97}\;{\scriptstyle[5.36,\,6.58]} 5.54​[5.02, 6.05]5.54\;{\scriptstyle[5.02,\,6.05]} (ρ=+0.08∗)(\rho{=}{+}0.08^{*}) (ρ=+0.09∗∗∗)(\rho{=}{+}0.09^{***}) (ρ=+0.18∗∗∗)(\rho{=}{+}0.18^{***}) (ρ=+0.01)(\rho{=}{+}0.01) Baseline, noise-trained 2.04​[1.86, 2.22]2.04\;{\scriptstyle[1.86,\,2.22]} 2.79​[2.32, 3.25]2.79\;{\scriptstyle[2.32,\,3.25]} 7.30​[6.57, 8.03]7.30\;{\scriptstyle[6.57,\,8.03]} 4.84​[4.60, 5.07]\mathbf{4.84}\;{\scriptstyle[4.60,\,5.07]} (ρ=+0.01)(\rho{=}{+}0.01) (ρ=+0.15∗∗∗)(\rho{=}{+}0.15^{***}) (ρ=+0.17∗∗∗)(\rho{=}{+}0.17^{***}) (ρ=+0.05∗)(\rho{=}{+}0.05^{*}) Full stack, clean-trained 1.91​[1.35, 2.47]1.91\;{\scriptstyle[1.35,\,2.47]} 1.57​[1.19, 1.95]1.57\;{\scriptstyle[1.19,\,1.95]} 7.66​[7.04, 8.28]7.66\;{\scriptstyle[7.04,\,8.28]} 6.18​[5.63, 6.73]6.18\;{\scriptstyle[5.63,\,6.73]} (ρ=+0.08)(\rho{=}{+}0.08) (ρ=+0.14∗∗∗)(\rho{=}{+}0.14^{***}) (ρ=+0.55∗∗∗)(\rho{=}{+}0.55^{***}) (ρ=+0.28∗∗∗)(\rho{=}{+}0.28^{***}) Full stack, noise-trained 1.79​[1.63, 1.96]\mathbf{1.79}\;{\scriptstyle[1.63,\,1.96]} 1.51​[1.13, 1.88]\mathbf{1.51}\;{\scriptstyle[1.13,\,1.88]} 8.32​[7.84, 8.81]8.32\;{\scriptstyle[7.84,\,8.81]} 5.62​[5.14, 6.09]5.62\;{\scriptstyle[5.14,\,6.09]} (ρ=+0.05)(\rho{=}{+}0.05) (ρ=+0.18∗∗∗)(\rho{=}{+}0.18^{***}) (ρ=+0.63∗∗∗)\mathbf{(\rho{=}{+}0.63^{***})} (ρ=+0.29∗∗∗)\mathbf{(\rho{=}{+}0.29^{***})} Random(synth)⋆ 3.683.68 3.813.81 4.574.57 4.574.57 Random(real)⋆⋆ 0.800.80 0.450.45 3.013.01 3.013.01 ⋆Predicts the synthetic training-set mean (6.0 m).
⋆⋆Predicts the training-split label mean of each corpus (2.272.27 m, 1.761.76 m, 8.958.95 m for VoiceHome2, STARSS23, QMUL-TIMIT). Requires real-data label access; included as a label-prior lower bound.

Results of the zero-shot analysis are shown in Table 2. On VoiceHome2 and STARSS23 every learned model outperforms the synthetic constant predictor (3.683.68 m and 3.813.81 m), confirming that some transferable distance cue survives the domain shift. On QMUL-TIMIT the ordering reverses: Random(synth) reaches 4.574.57 m while the learned models score 5.975.97–8.328.32 m, despite correlating distances far better. On VoiceHome2 the correlation never exceeds +0.08+0.08 and is significant in only one of four configurations, so the MAE gain over the constant predictor reflects a distribution shift rather than genuine distance correlation. This is not a range effect: STARSS23 spans a strictly narrower interval (σd=0.26\sigma_{d}{=}0.26 m against 0.950.95 m) yet reaches ρ=+0.18\rho{=}{+}0.18. On STARSS23 all models achieve significant correlations (ρ=0.09\rho=0.09–0.180.18, p<0.001p<0.001), the full-stack noise-trained model attaining ρ=+0.18\rho{=}{+}0.18 alongside the lowest MAE of 1.511.51 m, so even in this regime the full-stack architecture extracts a weak but reliable distance signal. The top row of Fig. 1 shows what this looks like: predictions form near-vertical smears with little dependence on the true distance.

Refer to caption
Figure 1: Zero-shot predictions of the full-stack noise-trained model on the four real test conditions. The dashed line is the identity; colour encodes point density. Correlations are Pearson, as in Table 2 (∗∗∗p<0.001{}^{***}p<0.001).
Table 3: Calibration MAE in meters for every map of Section 2, full-stack noise-trained model. Uncalibrated is the zero-shot MAE of Table 2. NN labelled samples are drawn from the target training split (mean over 500 draws); evaluation on the full test split. Best per column in bold. Markers denote a significant improvement over the constant predictor (paired tt-test across folds): ∗p<0.05{}^{*}p<0.05, p∗⁣∗<0.01{}^{**}p<0.01, ∗∗∗p<0.001{}^{***}p<0.001.

Calibration map N=5N{=}5 N=10N{=}10 N=20N{=}20 N=50N{=}50 VoiceHome2 Uncalibrated (d^=u\hat{d}{=}u) 1.79 Constant d¯\bar{d} 0.85 0.82 0.81 0.80 Offset-only (a=1a{=}1) 1.69 1.62 1.58 1.56 Scale-only (b=0b{=}0) 1.17 1.12 1.09 1.07 Affine (1) 1.00 0.87 0.83 0.81 Shrunk affine (8) 0.92 0.85 0.82 0.80 STARSS23 Uncalibrated (d^=u\hat{d}{=}u) 1.51 Constant d¯\bar{d} 0.47 0.45 0.45 0.45 Offset-only (a=1a{=}1) 2.01 1.93 1.90 1.88 Scale-only (b=0b{=}0) 1.10 1.11 1.12 1.12 Affine (1) 0.51 0.46 0.44∗ 0.44∗∗ Shrunk affine (8) 0.49 0.45 0.44∗ 0.44∗∗ QMUL-TIMIT, clean Uncalibrated (d^=u\hat{d}{=}u) 8.32 Constant d¯\bar{d} 3.20 3.08 3.05 3.02 Offset-only (a=1a{=}1) 2.88∗∗ 2.78∗∗ 2.71∗∗ 2.69∗∗ Scale-only (b=0b{=}0) 2.79∗∗ 2.60∗∗ 2.52∗∗ 2.47∗∗ Affine (1) 3.01∗ 2.54∗∗∗ 2.39∗∗∗ 2.30∗∗∗ Shrunk affine (8) 2.91∗∗ 2.59∗∗∗ 2.46∗∗∗ 2.33∗∗∗ QMUL-TIMIT, 00 dB Uncalibrated (d^=u\hat{d}{=}u) 5.62 Constant d¯\bar{d} 3.19 3.09 3.05 3.02 Offset-only (a=1a{=}1) 2.88∗∗∗ 2.76∗∗∗ 2.70∗∗∗ 2.67∗∗∗ Scale-only (b=0b{=}0) 3.17 2.99∗∗ 2.88∗∗ 2.82∗∗∗ Affine (1) 3.78 3.07 2.82∗∗ 2.70∗∗ Shrunk affine (8) 3.37 2.99 2.82∗∗ 2.72∗∗

On QMUL-TIMIT the evaluation separates correlation from absolute calibration, as the bottom row of Fig. 1 illustrates. Under clean conditions the full-stack model achieves ρ=0.63\rho{=}0.63 (p<0.001p{<}0.001), far above the baseline (ρ=0.17\rho{=}0.17), despite a much higher absolute MAE (8.328.32 m vs. 7.307.30 m): the predictions form a tight, clearly ordered band lying well below the identity line, so the model is badly miscalibrated in scale relative to QMUL-TIMIT’s distance distribution while preserving the correct ordering. At 00 dB the contrast persists, as the baselines retain essentially no correlation (ρ≤+0.05\rho\leq{+}0.05) while the full-stack model reaches ρ=+0.29\rho{=}{+}0.29 (p<0.001p{<}0.001), showing that the full-stack recipe is markedly more robust under domain shift.

4.3 Few-shot calibration

We draw NN samples from the training split of each corpus, fit every map of Section 2 by least squares, and evaluate on the full test split; Table 3 reports the mean over 500500 draws for the full-stack noise-trained model. No single map dominates, and the winner is set by the variance budget of (6) together with the spread ratio of (7). On VoiceHome2 and STARSS23 the estimator over-disperses (σu/σd=1.8\sigma_{u}/\sigma_{d}=1.8 and 4.44.4) while ρ≈0\rho\!\approx\!0, so both conditions fail: on VoiceHome2 every map is significantly worse than the constant predictor at every NN (p<0.01p<0.01), and on STARSS23 the affine map only overtakes it from N=20N{=}20 and then by 0.010.01 m. Offset-only, which retains the inflated spread, is the worst variant on both (1.691.69 and 2.012.01 m at N=5N{=}5). On QMUL-TIMIT the estimator under-disperses (0.170.17 clean, 0.280.28 at 00 dB) and the picture reverses: offset-only significantly beats the constant predictor at every NN in both conditions, winning every column at 00 dB and beating the affine map even at N=50N{=}50 (2.672.67 vs. 2.702.70 m). The affine map leads only where ρ\rho is large enough to pay for its second parameter, on QMUL-TIMIT clean from N=10N{=}10 onwards (2.542.54 m against 3.083.08 m, p<0.001p<0.001). The shrunk estimator (8) follows the same rationale without committing to a hard decision. It tracks the better of the affine and constant maps within 1%1\% for N≥10N\!\geq\!10, inherits the significant gains on QMUL-TIMIT in both conditions, and on the corpora where the slope carries no information it degrades gracefully towards the constant predictor rather than towards the affine map, costing 0.070.07 m on VoiceHome2 at N=5N{=}5 and matching it thereafter.

5 Conclusion

We studied few-shot calibration of a frozen, synthetic-trained single-channel distance estimator. Zero-shot transfer to real corpora is poor enough that a constant predictor at the corpus mean beats every learned model on all three datasets. Analysing the population risk of the affine map and its one-parameter restrictions shows that the calibrated error depends on how well the estimator correlates with real distances, not its absolute error. Based on the finite-sample cost of each fitted coefficient turns this into a condition, ρ2>1/(N+2)\rho^{2}>1/(N+2), that accounts for which map wins on which corpus and at which calibration budget, and into a shrinkage estimator that interpolates between the affine and constant maps without a hard threshold. The practical consequence is that the optimal distance estimator should be selected on linear correlation rather than MAE, since calibration repairs scale but cannot repair uncorrelated predictions.

References

  • [1] V. Hamacher, J. Chalupper, J. Eggers, E. Fischer, U. Kornagel, H. Puder, and U. Rass (2005) Signal processing in high-end hearing aids: State of the art, challenges, and future trends. EURASIP Journal on Advances in Signal Processing 2005 (18), pp. 1–15. Cited by: §1.
  • [2] S. Oh, V. Viswanathan, and P. Papamichalis (1992) Hands-free voice communication in an automobile with a microphone array. In IEEE ICASSP, Vol. , pp. . Cited by: §1.
  • [3] M. Omologo, P. Svaizer, and M. Matassoni (1998) Environmental conditions and acoustic transduction in hands-free speech recognition. Speech Communication 25 (1-3), pp. 75–95. Cited by: §1.
  • [4] M. Neri, A. Politis, and T. Virtanen (2026) Dependence on Early and Late Reverberation of Single-Channel Speaker Distance Estimation . In IWAENC, Vol. , pp. . External Links: Document Cited by: §1, §1, §1, §2.2, §3.2, Table 1.
  • [5] P. Zahorik, D. S. Brungart, and A. W. Bronkhorst (2005) Auditory distance perception in humans: a summary of past and present research. ACTA Acustica united with Acustica 91 (3), pp. 409–420. Cited by: §1.
  • [6] C. Mendonça, P. Mandelli, and V. Pulkki (2016) Modeling the perception of audiovisual distance: bayesian causal inference and other models. PloS one 11 (12), pp. e0165391. Cited by: §1.
  • [7] E. C. Cherry (1953) Some experiments on the recognition of speech, with one and with two ears. The Journal of the Acoustical Society of America 25 (5), pp. 975–979. Cited by: §1.
  • [8] E. Georganti, T. May, S. van de Par, A. Harma, and J. Mourjopoulos (2011) Speaker Distance Detection Using a Single Microphone. IEEE Transactions on Audio, Speech, and Language Processing 19 (7), pp. 1949–1961. External Links: Document Cited by: §1, §1.
  • [9] Y.C. Lu and M. Cooke (2010) Binaural Estimation of Sound Source Distance via the Direct-to-Reverberant Energy Ratio for Static and Moving Sources. IEEE Transactions on Audio, Speech, and Language Processing 18 (7), pp. 1793–1805. External Links: Document Cited by: §1.
  • [10] S. Vesa (2007) Sound Source Distance Learning Based on Binaural Signals. In IEEE WASPAA, Vol. , pp. . Cited by: §1.
  • [11] S. Vesa (2009) Binaural Sound Source Distance Learning in Rooms. IEEE Transactions on Audio, Speech, and Language Processing 17 (8), pp. 1498–1507. External Links: Document Cited by: §1.
  • [12] E. Georganti, T. May, S. van de Par, and J. Mourjopoulos (2013) Sound Source Distance Estimation in Rooms based on Statistical Properties of Binaural Signals. IEEE Transactions on Audio, Speech, and Language Processing 21 (8), pp. 1727–1741. External Links: Document Cited by: §1.
  • [13] K. Patterson, K. Wilson, S. Wisdom, and J. R. Hershey (2022) Distance-Based Sound Separation. In Interspeech, Cited by: §1.
  • [14] D. A. Krause, A. Politis, and A. Mesaros (2021) Joint direction and proximity classification of overlapping sound events from binaural audio. In IEEE WASPAA, pp. . Cited by: §1.
  • [15] M. Yiwere and E. J. Rhee (2019) Sound source distance estimation using deep learning: An image classification approach. Sensors 20 (1), pp. 172. Cited by: §1.
  • [16] A. Sobhdel, R. Razavi-Far, and V. Palade (2024) A Few-Shot Learning Approach for Sound Source Distance Estimation Using Relation Networks. In ICMLA, Vol. , pp. . External Links: Document Cited by: §1.
  • [17] M. Neri, A. Politis, D. A. Krause, M. Carli, and T. Virtanen (2023) Single-Channel Speaker Distance Estimation in Reverberant Environments. In IEEE WASPAA, Vol. , pp. . External Links: Document Cited by: §1, §3.1.
  • [18] M. Neri, A. Politis, D. A. Krause, M. Carli, and T. Virtanen (2024) Speaker Distance Estimation in Enclosures From Single-Channel Audio. IEEE/ACM Transactions on Audio, Speech, and Language Processing 32 (), pp. 2242–2254. External Links: Document Cited by: §1, §3.1, §3.3, §4.2.
  • [19] J. Lin, G. Georg, H. S. Llopis, H. Hafsteinsson, S. Guðjónsson, D. G. Nielsen, F. Pind, P. Smaragdis, D. Manocha, J. Hershey, T. Kristjansson, and M. Kim (2025) Generative data augmentation challenge: synthesis of room acoustics for speaker distance estimation. In IEEE ICASSPW, Cited by: §1.
  • [20] B. Efron and C. Morris (1973) Stein’s estimation rule and its competitors—an empirical bayes approach. Journal of the American Statistical Association 68 (341), pp. 117–130. Cited by: §2.3.
  • [21] J. S. Abel and P. Huang (2006) A simple, robust measure of reverberation echo density. In Audio Engineering Society Convention 121, Cited by: §3.1.
  • [22] D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le (2019) SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition. In Interspeech, pp. 2613–2617. External Links: Document Cited by: §3.1.
  • [23] G. Wichern, J. Antognini, M. Flynn, L. R. Zhu, E. McQuinn, D. Crow, E. Manilow, and J. Le Roux (2019) WHAM!: Extending Speech Separation to Noisy Environments. In Interspeech, Cited by: §3.1.
  • [24] J. Richter, Y. Wu, S. Krenn, S. Welker, B. Lay, S. Watanabe, A. Richard, and T. Gerkmann (2024) EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation. In Interspeech, pp. . Cited by: §3.2.
  • [25] R. Scheibler, E. Bezzam, and I. Dokmanić (2018) Pyroomacoustics: A python package for audio room simulation and array processing algorithms. In IEEE ICASSP, pp. . Cited by: §3.2.
  • [26] N. Bertin, E. Camberlein, R. Lebarbenchon, E. Vincent, S. Sivasankaran, I. Illina, and F. Bimbot (2019) VoiceHome-2, an extended corpus for multichannel speech processing in real homes. Speech Communication 106, pp. 68–78. External Links: ISSN 0167-6393, Document Cited by: §3.3, §4.2.
  • [27] K. Shimada, A. Politis, P. Sudarsanam, D. A. Krause, K. Uchida, S. Adavanne, A. Hakala, Y. Koyama, N. Takahashi, S. Takahashi, et al. (2023) STARSS23: An audio-visual dataset of spatial recordings of real scenes with spatiotemporal annotations of sound events. Advances in neural information processing systems 36, pp. 72931–72957. Cited by: §3.3, §4.2.