- MAE
- mean absolute error
- CNN
- convolutional neural network
- CRNN
- convolutional recurrent neural network
- DoA
- direction of arrival
- DNN
- deep neural network
- DRR
- direct-to-reverberant ratio
- ELU
- exponential linear unit
- GMM
- Gaussian mixture model
- GRU
- gated recurrent unit
- IID
- interaural intensity difference
- ITD
- interaural time difference
- LSTM
- long short-term memory
- LPC
- linear predictive coding
- MSE
- mean squared error
- MLP
- multilayer perceptron
- RIR
- room impulse response
- RNN
- recurrent neural network
- SELD
- sound event localization and detection
- SNR
- signal-to-noise ratio
- SVM
- support vector machines
- STFT
- short-time fourier transform
Few-Shot Calibration for Sim-to-Real Single-Channel Speaker Distance Estimation
Abstract
Speaker distance estimators are trained almost exclusively on simulated room acoustics, because real recordings annotated with the true talker-to-microphone distance are scarce. We show that models trained this way transfer poorly. On three real corpora we evaluate, simply predicting the average distance of the corpus is more accurate than any learned model. Then, we ask how few labelled real utterances are needed to make a frozen, synthetic-trained estimator useful, and study post-hoc calibration maps that rescale its output without gradients or retraining. An analysis of the achievable error shows that what the calibration is not limited by the absolute accuracy of the estimator, but how well it orders utterances by distance, since a constant bias or a wrong output scale is removed exactly by the calibration itself. Balancing this against the cost of estimating each coefficient from few samples yields a criterion that accounts for which map wins on which corpus and at which annotation budget, together with a shrinkage variant that requires no hard decision. Our findings suggest selecting synthetic checkpoints by linear correlation with true distances rather than by absolute error. Code, datasets, and analysis are available at https://github.com/michaelneri/audio-distance-estimation.
Index Terms:
Acoustics, Speaker Distance Estimation, Room Impulse Response, Multi-task learning, Sim-to-Real Transfer.{michael.neri, archontis.politis, tuomas.virtanen}@tuni.fi
1 Introduction
Estimating the distance between a talker and a microphone from a single microphone is of practical interests for several applications, ranging from hearing aid devices [1], hands-free communication [2], and speech recognition [3]. The cue is intrinsically acoustic: the direct-to-reverberant ratio, the early-reflection pattern and the spectral envelope all vary systematically with distance [4, 5, 6]. When these pieces of information are degraded, both human and machines distance perception deteriorate considerably [7, 8].
Early approaches to the more generic source distance estimation focused on binaural recordings and hand-crafted features derived from the DRR (DRR) [9] or from inter-channel correlations. These features were used to fit GMM and SVM (SVM) classifiers to distinguish distance between discrete bins [10, 11, 12]. More recent studies revisited the problem with DNN, yet most remain limited to binary far/near classification [13, 14] or to coarse distance bins over a narrow range [8, 15, 16]. Yet, these methods typically required careful hyperparameter tuning and generalised poorly across different acoustic environments.
In fact, the practical obstacle is the availability of audio data with labeled distances. The research community has therefore converged on training with synthetic data. The work in [17] was the first to frame speaker distance as a continuous regression problem, demonstrating that a DNN operating on STFT (STFT) phase features can achieve centimeter-level accuracy in simulated environments. It is worth noting that the centimeter-level accuracy reported in [18, 17] was obtained on simulated data with time- and level-calibration. Similarly, in [19] the authors proposed a data augmentation approach for generating realistic room-acoustic conditions to support speaker distance estimation using the CRNN (CRNN) model in [18]. Most directly relevant to this work, in [4] it was shown that learning-based single-channel distance estimators rely predominantly on early reflections rather than other RIR (RIR) components, but it did not study the sim-to-real transfer for speaker distance estimation without the knowledge of the time calibration. The only result available in this scope is in [18], which provided near-chance results on zero-shot experiments on three real corpora with a simulation-based estimator.
To mitigate this sim-to-real gap, the contributions of this work are: (i) we carried out zero-shot speaker distance estimation analysis, analyzing how much mismatch is present when training on synthetic audios and testing on real data. (ii) We devised a few-shot calibration that adapts a frozen simulation-based estimator to a new corpus with labelled utterances by fitting a two-parameter affine map, requiring no gradients and no retraining. (iii) Based on [4], we modified the head of the CRNN to predict acoustic-based features of the enclosure. While this leaves simulation-based MAE (MAE) largely unchanged, it substantially improves correlation with true distance labels under domain shift, which is what governs how well the estimator can be calibrated.
2 Few-shot calibration
2.1 Problem statement
Let be a distance estimator trained exclusively on synthetic data with learned parameters and kept frozen. Deployed on a target corpus it predicts an uncalibrated distance , nominally in meters from the recording with samples. We assume access to a small calibration set drawn from the training split of that corpus, with of the order of a few tens of samples, and find a map producing a calibrated estimate , requiring neither gradients, nor access to , nor retraining.
2.2 Foundations of distance calibration
Two systematic effects displace from . First, a regressor trained under a squared-error objective shrinks towards the prior mean of , whereas concentrates its distances elsewhere, causing an offset. Second, since the amplitude cue contributes negligibly [4], the estimator maps early-reflection structure to distance under the room-geometry and absorption statistics seen in training; a target enclosure whose reverberant statistics differ therefore induces a multiplicative scale error. Both can be corrected by the affine map
| (1) |
with estimated by least squares on , without any domain adaptation or fine-tuning.
The constant predictor is the reference any calibration must beat: it consumes the same labels, estimates the average , and ignores entirely. Constraining either coefficient of (1) gives two one-parameter members of the same family, offset-only (), which corrects the prior shift alone, and scale-only (), which corrects the scale alone. Both trade modelling bias for estimation variance.
2.3 Estimating the affine mapping
Let and denote the target-corpus label and raw predicted distance statistics, and the Pearson correlation between and . Theoretically with unlimited number of samples (), the coefficients of (1) that minimise the population risk are obtained by solving and , yielding the optimal coefficients
| (2) |
Substituting back to yields after some calculations
| (3) |
It is worth noting that is invariant to affine transformations of , so a constant bias or a wrong output scale leaves (3) unchanged. What remains is how well the model correlates with distances within the target corpus. A model with large zero-shot MAE and high calibrates better than an accurate but unordered one. With finite , are themselves estimated from , and their sampling error adds to the test risk. An affine map ( free parameters) has expected test risk
| (4) |
The constant predictor is the special case , so the same reasoning applies with ,
| (5) |
Having both the risk of the affine and the constant calibrations, we can estimate when it is worth fitting a slope, i.e., when , which yields the inequality
| (6) |
Using this inequality which includes the number of calibrating samples and the correlation between ground-truth and predicted distances of the target corpus, we can set the threshold in function of the cardinality of the calibration set, i.e. at , at and at . When the inequality does not hold, the least-squares slope is dominated by its own sampling noise.
Offset-only admits an analogous test. Shifting the estimator without rescaling it leaves the score spread intact, so its risk is . As it estimates a single parameter, like the constant predictor, both carry the same variance penalty and the factors cancel; the comparison therefore involves no at all and reduces to
| (7) |
which holds only when the estimator under-disperses relative to the target labels.
It is worth noting that the correlation between the estimator and the calibration set is itself estimated from few samples, and thus noisy at small . We denote this estimate . Then, we propose an additional mapping that avoids a hard threshold on (6) and instead shrinks the slope in proportion to the evidence for it [20],
3 Materials
3.1 Distance estimator
We use as distance estimator the convolutional-recurrent regressor of [17] as the baseline. It processes STFT log-magnitude and sine/cosine phase features, a three-block CNN (CNN) with max/average pooling, a two-layer bidirectional GRU (GRU), and a per-frame distance head pooled over time. Input is s of mono audio at kHz. We additionally evaluate a full-stack variant that adds, over the baseline, (i) auxiliary regression heads that predict , , and from the mean-pooled recurrent features, adding roughly k learnable parameters. Here is the mixing time, the boundary between early and late reflections in the RIR, obtained from the echo-density measure of [21]. These parameters are chosen because they summarise the reverberant statistics that carry the distance cue; (ii) SpecAugment [22] ( time masks, frequency masks) applied after feature extraction; and (iii) waveform augmentation (polarity inversion, dB gain, ms shift, each with probability ). Both variants are trained with Adam at under two regimes: clean-trained, and noise-trained with WHAM! [23] noise mixed in at randomised SNR (SNR) spanning uniformly dB. Following [18], both the clip-level prediction and the per-frame prediction are supervised, which encourages the recurrent layers to produce sharper per-frame estimates, giving
| (9) |
where is the MSE (MSE), collects the auxiliary targets with predictions and weight , and depending on whether the log-distance target is enabled. Auxiliary terms are dropped when the corresponding head is disabled in an ablation. All variants use five-fold cross-validation, giving five checkpoints each.
3.2 Synthetic dataset
We employ the same uncalibrated dataset (neither time- nor amplitude-calibration) as in [4]. Specifically, anechoic speech recordings obtained from the EARS dataset [24] are convolved with the simulated RIR from pyroomacoustics [25]. The experiments include audio files of s duration at kHz. The samples are randomly assigned to folds to assess the performance in a -fold cross-validation fashion. By doing so, each cross-fold iteration assigns , , and audios to training, validation, and testing sets, respectively. We follow the five-fold cross-validation protocol of [4]: for each fold , fold is held out for testing, fold for validation, and the remaining three folds are used for training. The simulated rooms span a wide acoustic range: source-to-microphone distances from to m (mean m), reverberation times from to s (mean s), room volumes from to (mean ), and mixing times from to ms (mean ms).
3.3 Real datasets
We evaluate sim-to-real transfer on three corpora spanning two levels of realism. VoiceHome2 [26] which encompasses smart-home commands recorded by twelve speakers in twelve rooms across four houses, under quiet and noisy conditions with uncontrolled domestic interferers (competing talkers, TV, appliances) and no SNR annotation. Five source positions per room, standing and sitting, with a fixed 8-microphone MEMS array on a cubic baffle where we use the first channel only. In total, the dataset encompasses recordings of s, with distances spanning - m (mean m). STARSS23 [27]. Multi-speaker interaction scenes recorded at Tampere University and Sony in eleven rooms with an Eigenmike array, from which a single omnidirectional channel is extracted. We use single-speech excerpts that do not overlap with other annotated directional sources. The corpus is the most challenging of the three: speakers move and change orientation, and diffuse and directional ambient noise is present at significant levels. Distances span – m (mean 2.19 m). QMUL-TIMIT [18]. Measured omnidirectional RIR captured in three rooms at Queen Mary University of London: a 7.5×9×3.5 m classroom ( m3, RIR), the Octagon, an eight-walled Victorian hall with a m domed ceiling ( m3, RIR), and the Great Hall ( RIRs over a 12×12 m region). Each RIR is convolved with anechoic TIMIT utterances, yielding recordings, with RIR split // into training/validation/testing. We report clean and dB conditions with WHAM! noise, matching the synthetic protocol.
4 Results
4.1 Results on synthetic data
Table 1 reports the MAE at the clean and dB extremes. Training the baseline with noise already recovers most of the dB degradation ( m); on top of the noise-aware regime the multi-task and augmentation stack gives a smaller further gain ( m) while also improving the clean condition ( m). The per-component contributions are not monotonic: the log-distance target alone significantly increases dB MAE under clean training (), and neither the multi-task heads nor SpecAugment alone produce a significant change. A plausible explanation is that the log-distance loss reweights gradients toward small distances, which sharpens the target but offers no benefit without the implicit regularisation that noise diversity provides; under noise-aware training the same target is harmless. Only the full stack yields a significant improvement in both regimes.
Clean-trained Noise-trained Model clean dB clean dB Baseline [4] 1.30 0.08 1.83 0.17 1.29 0.06 1.47 0.11 + log-distance target 1.38 0.13 2.01 0.14∗ 1.31 0.09 1.46 0.13 + multi-task heads 1.35 0.06 1.92 0.19 1.34 0.12 1.51 0.10 + SpecAugment 1.32 0.13 1.78 0.20 1.29 0.12 1.49 0.11 + speech aug. (full stack) 1.23 0.13 1.69 0.13∗∗ 1.23 0.09 1.41 0.11∗ Random 2.47 0.06
4.2 Zero-shot real data analysis
We revisit zero-shot analysis proposed in [18] using the same two real corpora (VoiceHome2 [26] and STARSS23 [27]) and the hybrid QMUL-TIMIT [18] to analyze the sim-to-real gap without any fine-tuning. For the comparison, we report two baselines that draw the label-prior floor. The first, Random (synth.), predicts the synthetic training-set mean for every sample, without requiring any real-data labels. The second, Random (real), predicts the training-split label mean of each target corpus. However, it is worth noting that the Random (real) baseline is not strictly zero-shot as it requires label access to the target domain.
Model
VoiceHome2
STARSS23
QMUL-TIMIT
clean
0 dB
Baseline, clean-trained
Baseline, noise-trained
Full stack, clean-trained
Full stack, noise-trained
Random(synth)⋆
Random(real)⋆⋆
⋆Predicts the synthetic training-set mean (6.0 m).
⋆⋆Predicts the training-split label mean of each corpus ( m, m, m for VoiceHome2, STARSS23, QMUL-TIMIT). Requires real-data label access; included as a label-prior lower bound.
Results of the zero-shot analysis are shown in Table 2. On VoiceHome2 and STARSS23 every learned model outperforms the synthetic constant predictor ( m and m), confirming that some transferable distance cue survives the domain shift. On QMUL-TIMIT the ordering reverses: Random(synth) reaches m while the learned models score – m, despite correlating distances far better. On VoiceHome2 the correlation never exceeds and is significant in only one of four configurations, so the MAE gain over the constant predictor reflects a distribution shift rather than genuine distance correlation. This is not a range effect: STARSS23 spans a strictly narrower interval ( m against m) yet reaches . On STARSS23 all models achieve significant correlations (–, ), the full-stack noise-trained model attaining alongside the lowest MAE of m, so even in this regime the full-stack architecture extracts a weak but reliable distance signal. The top row of Fig. 1 shows what this looks like: predictions form near-vertical smears with little dependence on the true distance.
Calibration map VoiceHome2 Uncalibrated () 1.79 Constant 0.85 0.82 0.81 0.80 Offset-only () 1.69 1.62 1.58 1.56 Scale-only () 1.17 1.12 1.09 1.07 Affine (1) 1.00 0.87 0.83 0.81 Shrunk affine (8) 0.92 0.85 0.82 0.80 STARSS23 Uncalibrated () 1.51 Constant 0.47 0.45 0.45 0.45 Offset-only () 2.01 1.93 1.90 1.88 Scale-only () 1.10 1.11 1.12 1.12 Affine (1) 0.51 0.46 0.44∗ 0.44∗∗ Shrunk affine (8) 0.49 0.45 0.44∗ 0.44∗∗ QMUL-TIMIT, clean Uncalibrated () 8.32 Constant 3.20 3.08 3.05 3.02 Offset-only () 2.88∗∗ 2.78∗∗ 2.71∗∗ 2.69∗∗ Scale-only () 2.79∗∗ 2.60∗∗ 2.52∗∗ 2.47∗∗ Affine (1) 3.01∗ 2.54∗∗∗ 2.39∗∗∗ 2.30∗∗∗ Shrunk affine (8) 2.91∗∗ 2.59∗∗∗ 2.46∗∗∗ 2.33∗∗∗ QMUL-TIMIT, dB Uncalibrated () 5.62 Constant 3.19 3.09 3.05 3.02 Offset-only () 2.88∗∗∗ 2.76∗∗∗ 2.70∗∗∗ 2.67∗∗∗ Scale-only () 3.17 2.99∗∗ 2.88∗∗ 2.82∗∗∗ Affine (1) 3.78 3.07 2.82∗∗ 2.70∗∗ Shrunk affine (8) 3.37 2.99 2.82∗∗ 2.72∗∗
On QMUL-TIMIT the evaluation separates correlation from absolute calibration, as the bottom row of Fig. 1 illustrates. Under clean conditions the full-stack model achieves (), far above the baseline (), despite a much higher absolute MAE ( m vs. m): the predictions form a tight, clearly ordered band lying well below the identity line, so the model is badly miscalibrated in scale relative to QMUL-TIMIT’s distance distribution while preserving the correct ordering. At dB the contrast persists, as the baselines retain essentially no correlation () while the full-stack model reaches (), showing that the full-stack recipe is markedly more robust under domain shift.
4.3 Few-shot calibration
We draw samples from the training split of each corpus, fit every map of Section 2 by least squares, and evaluate on the full test split; Table 3 reports the mean over draws for the full-stack noise-trained model. No single map dominates, and the winner is set by the variance budget of (6) together with the spread ratio of (7). On VoiceHome2 and STARSS23 the estimator over-disperses ( and ) while , so both conditions fail: on VoiceHome2 every map is significantly worse than the constant predictor at every (), and on STARSS23 the affine map only overtakes it from and then by m. Offset-only, which retains the inflated spread, is the worst variant on both ( and m at ). On QMUL-TIMIT the estimator under-disperses ( clean, at dB) and the picture reverses: offset-only significantly beats the constant predictor at every in both conditions, winning every column at dB and beating the affine map even at ( vs. m). The affine map leads only where is large enough to pay for its second parameter, on QMUL-TIMIT clean from onwards ( m against m, ). The shrunk estimator (8) follows the same rationale without committing to a hard decision. It tracks the better of the affine and constant maps within for , inherits the significant gains on QMUL-TIMIT in both conditions, and on the corpora where the slope carries no information it degrades gracefully towards the constant predictor rather than towards the affine map, costing m on VoiceHome2 at and matching it thereafter.
5 Conclusion
We studied few-shot calibration of a frozen, synthetic-trained single-channel distance estimator. Zero-shot transfer to real corpora is poor enough that a constant predictor at the corpus mean beats every learned model on all three datasets. Analysing the population risk of the affine map and its one-parameter restrictions shows that the calibrated error depends on how well the estimator correlates with real distances, not its absolute error. Based on the finite-sample cost of each fitted coefficient turns this into a condition, , that accounts for which map wins on which corpus and at which calibration budget, and into a shrinkage estimator that interpolates between the affine and constant maps without a hard threshold. The practical consequence is that the optimal distance estimator should be selected on linear correlation rather than MAE, since calibration repairs scale but cannot repair uncorrelated predictions.
References
- [1] (2005) Signal processing in high-end hearing aids: State of the art, challenges, and future trends. EURASIP Journal on Advances in Signal Processing 2005 (18), pp. 1–15. Cited by: §1.
- [2] (1992) Hands-free voice communication in an automobile with a microphone array. In IEEE ICASSP, Vol. , pp. . Cited by: §1.
- [3] (1998) Environmental conditions and acoustic transduction in hands-free speech recognition. Speech Communication 25 (1-3), pp. 75–95. Cited by: §1.
- [4] (2026) Dependence on Early and Late Reverberation of Single-Channel Speaker Distance Estimation . In IWAENC, Vol. , pp. . External Links: Document Cited by: §1, §1, §1, §2.2, §3.2, Table 1.
- [5] (2005) Auditory distance perception in humans: a summary of past and present research. ACTA Acustica united with Acustica 91 (3), pp. 409–420. Cited by: §1.
- [6] (2016) Modeling the perception of audiovisual distance: bayesian causal inference and other models. PloS one 11 (12), pp. e0165391. Cited by: §1.
- [7] (1953) Some experiments on the recognition of speech, with one and with two ears. The Journal of the Acoustical Society of America 25 (5), pp. 975–979. Cited by: §1.
- [8] (2011) Speaker Distance Detection Using a Single Microphone. IEEE Transactions on Audio, Speech, and Language Processing 19 (7), pp. 1949–1961. External Links: Document Cited by: §1, §1.
- [9] (2010) Binaural Estimation of Sound Source Distance via the Direct-to-Reverberant Energy Ratio for Static and Moving Sources. IEEE Transactions on Audio, Speech, and Language Processing 18 (7), pp. 1793–1805. External Links: Document Cited by: §1.
- [10] (2007) Sound Source Distance Learning Based on Binaural Signals. In IEEE WASPAA, Vol. , pp. . Cited by: §1.
- [11] (2009) Binaural Sound Source Distance Learning in Rooms. IEEE Transactions on Audio, Speech, and Language Processing 17 (8), pp. 1498–1507. External Links: Document Cited by: §1.
- [12] (2013) Sound Source Distance Estimation in Rooms based on Statistical Properties of Binaural Signals. IEEE Transactions on Audio, Speech, and Language Processing 21 (8), pp. 1727–1741. External Links: Document Cited by: §1.
- [13] (2022) Distance-Based Sound Separation. In Interspeech, Cited by: §1.
- [14] (2021) Joint direction and proximity classification of overlapping sound events from binaural audio. In IEEE WASPAA, pp. . Cited by: §1.
- [15] (2019) Sound source distance estimation using deep learning: An image classification approach. Sensors 20 (1), pp. 172. Cited by: §1.
- [16] (2024) A Few-Shot Learning Approach for Sound Source Distance Estimation Using Relation Networks. In ICMLA, Vol. , pp. . External Links: Document Cited by: §1.
- [17] (2023) Single-Channel Speaker Distance Estimation in Reverberant Environments. In IEEE WASPAA, Vol. , pp. . External Links: Document Cited by: §1, §3.1.
- [18] (2024) Speaker Distance Estimation in Enclosures From Single-Channel Audio. IEEE/ACM Transactions on Audio, Speech, and Language Processing 32 (), pp. 2242–2254. External Links: Document Cited by: §1, §3.1, §3.3, §4.2.
- [19] (2025) Generative data augmentation challenge: synthesis of room acoustics for speaker distance estimation. In IEEE ICASSPW, Cited by: §1.
- [20] (1973) Stein’s estimation rule and its competitors—an empirical bayes approach. Journal of the American Statistical Association 68 (341), pp. 117–130. Cited by: §2.3.
- [21] (2006) A simple, robust measure of reverberation echo density. In Audio Engineering Society Convention 121, Cited by: §3.1.
- [22] (2019) SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition. In Interspeech, pp. 2613–2617. External Links: Document Cited by: §3.1.
- [23] (2019) WHAM!: Extending Speech Separation to Noisy Environments. In Interspeech, Cited by: §3.1.
- [24] (2024) EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation. In Interspeech, pp. . Cited by: §3.2.
- [25] (2018) Pyroomacoustics: A python package for audio room simulation and array processing algorithms. In IEEE ICASSP, pp. . Cited by: §3.2.
- [26] (2019) VoiceHome-2, an extended corpus for multichannel speech processing in real homes. Speech Communication 106, pp. 68–78. External Links: ISSN 0167-6393, Document Cited by: §3.3, §4.2.
- [27] (2023) STARSS23: An audio-visual dataset of spatial recordings of real scenes with spatiotemporal annotations of sound events. Advances in neural information processing systems 36, pp. 72931–72957. Cited by: §3.3, §4.2.