BioSync: Transformer-Based Cross-Modal Fusion for a Multimodal Physiological Digital Biomarker
Abstract
Cardiac, neural, behavioral, and speech measurements from wearable and mobile devices provide partial, noise-sensitive views of physiological state. BioSync combines these measurements into the BioSync Index (BSI), a continuous composite digital biomarker defined under the BEST framework. The model applies multi-head self-attention to modality tokens and adds a linear branch whose hypothesis class includes standard feature concatenation. This architecture is motivated by latent-variable measurement theory and by the possibility that joint observations contain information unavailable from individual modalities. We evaluated BioSync on two literature-informed synthetic cohorts: a four-modality cognitive-decline cohort using HRV, EEG, actigraphy, and speech, and a metabolic-autonomic cohort structured around the public AI-READI wearable schema. In the cognitive cohort, BioSync and concatenation obtained AUCs of 0.928 and 0.926, respectively. In the metabolic cohort, BioSync obtained accuracy/F1 of 0.764/0.766, compared with 0.756/0.758 for concatenation. The BSI correlated with latent severity in both cohorts ( and ). A pure-attention ablation obtained cognitive-cohort AUC 0.911, locating the increase to 0.928 in the combined wide-and-deep architecture. With matched modality-dropout training, BioSync led concatenation at five of six cognitive-cohort corruption rates and at the highest metabolic-cohort rate. Its cognitive-cohort AUC was also higher than five published digital-biomarker reference values, although differences in datasets and tasks preclude a controlled benchmark claim. Comparison with single-modality, early-fusion, and late-fusion designs across six prespecified criteria identifies the model’s computational properties; validation on real cohorts remains necessary.
Keywords:
digital biomarker , multimodal data fusion , transformer , self-attention , wearable sensors , mild cognitive impairment , information theory , robustness1 Introduction
Dementia affects more than 55 million people worldwide, with Alzheimer’s disease (AD) accounting for most cases; prevalence is projected to increase as the global population ages [25, 3]. Clinical assessment combines neuropsychological testing with biomarkers such as cerebrospinal fluid assays, amyloid positron emission tomography, and structural imaging. These procedures are not well suited to frequent longitudinal monitoring of mild cognitive impairment (MCI) in daily life [18, 8]. Digital biomarkers—characteristics measured through digital health technologies—offer a complementary route. Wearables, smartphones, and portable EEG systems can repeatedly collect physiological and behavioral measures outside the clinic [38, 21, 30]. In a review of 431 studies, mean AUC was 0.887 for AI-based AD models and 0.821 for MCI models; fewer than 3% of studies performed external validation. Multimodal biomarkers appeared in only 24 studies and were less common than gait-, speech-, or eye-tracking-based measures [31].
A meta-analysis of digital-biomarker technologies for MCI and pre-frailty screening reported pooled sensitivity and specificity of approximately 80% [37]. This result provides a reference point for single-modality screening, although it does not establish a universal performance ceiling. Consumer- and research-grade devices can measure several signals relevant to cognitive decline. Heart-rate variability (HRV) provides a non-invasive measure of autonomic function and has been associated with early AD-related autonomic dysfunction, although studies disagree on the direction and magnitude of the association [6, 26, 24].
Quantitative EEG contributes neural measures: the theta-to-alpha power ratio and reduced spectral complexity are replicated correlates of amnestic MCI, and dry-electrode headbands can record these measures at home [7, 19]. Wrist actigraphy measures rest–activity fragmentation and circadian amplitude, both of which have been linked to cognitive status in longitudinal cohorts [14, 36]. Together, HRV, EEG, and actigraphy sample cardiac-autonomic, neural, and behavioral/circadian pathways. Speech adds acoustic and linguistic measures that have also been used for cognitive screening. Prior multimodal studies support combining signals across these pathways [22, 42, 43].
Most multimodal digital-biomarker pipelines concatenate features before classification (early fusion) or combine independently trained modality-specific classifiers by majority or soft voting (late fusion) [31]. Neither design directly produces subject- and session-specific modality weights that can respond to signal quality. Cross-modal attention can assign each modality a data-dependent contribution to a shared representation. Attention-based fusion has outperformed static fusion in physiological-signal tasks including affect and emotion recognition and cardiovascular classification [12, 23, 20]. Applications to wearable cognitive-decline biomarkers remain limited [7], and few studies evaluate the computational constraints of wearable or edge deployment.
BioSync treats each modality as a token and uses multi-head self-attention to estimate subject-level modality contributions. Its output, the continuous BioSync Index (BSI), is defined as a composite digital biomarker under the FDA-NIH BEST framework [15, 8]. The model also includes a parallel linear branch so that its hypothesis class retains the decision functions available to feature concatenation (Section 4.4). Section 3 relates this design to latent-variable measurement and to information that may arise only from the joint distribution of modalities. Figure 1 summarizes the paper’s high-level workflow, from multimodal measurements through BioSync fusion to the BSI and the four evaluation components.
We test the same encoder in two domains, changing only the number of modality tokens. The cognitive-decline application uses HRV, EEG, actigraphy, and speech. The metabolic-autonomic application follows the public AI-READI wearable-activity-monitor schema [1], which documents Garmin Vivosmart 5 data from an NIH Bridge2AI type 2 diabetes (T2D) program; cohort assumptions are drawn from the diabetic cardiac autonomic neuropathy literature. Independent modality encoders also make the architecture compatible with federated training across clinics, device cohorts, or datasets without pooling raw physiological data [28, 9, 2, 29]. No patient or raw AI-READI data are analyzed. The two literature-seeded synthetic cohorts test whether the fusion pipeline learns cross-modal weights, whether the BSI follows latent severity, and whether performance degrades gradually when a modality is corrupted. These experiments characterize the architecture under simulation and do not support clinical inference. The next step is validation with AI-READI records and with the prospective wearable cohort under collection at our laboratory.
2 Related Work
2.1 Digital biomarkers for Alzheimer’s disease and MCI
The FDA-NIH BEST resource classifies biomarkers as diagnostic, monitoring, response/pharmacodynamic, predictive, prognostic, safety, or susceptibility/risk measures. A measure in any of these categories can be considered a digital biomarker when it is collected through digital health technology [15, 8]. The V3 framework organizes fitness-for-purpose evidence into verification, analytical validation, and clinical validation, spanning sensor performance through comparison of the derived measure with a reference clinical outcome [16]. BioSync has not yet undergone these stages (Section 7).
Kourtis et al. catalogued how mobile and wearable devices could operationalize the BEST categories for AD. Their candidate channels included gait and speech as well as physiological signals [21]. Piau et al. reviewed home-based monitoring technologies for MCI and early AD, calling for measurement that is both longitudinal and ecologically valid [30]. Qi et al. reviewed 431 studies and 86 AI models. Only 24 studies evaluated multimodal digital biomarkers, and external validation and calibration were rarely reported [31]. Xu et al. reported 91.8% accuracy for cognitive-impairment screening with hard-voting late fusion of language-derived digital biomarkers [42]. The result supports multimodal screening, but the fusion rule remains fixed across participants.
2.2 Cardiac-autonomic, neural, and behavioral digital signals
HRV reflects sympathetic and parasympathetic activity through the central autonomic network, which may be disrupted early in the AD pathological cascade [6, 24]. The reported direction and strength of the association vary across studies. Bateman et al. reported preliminary differences in HRV metrics during routine cognitive testing [6], whereas Marcolini et al. found limited evidence that HRV alone predicts cognitive and pathophysiological outcomes [26]. This disagreement supports treating HRV as one signal among several rather than as a stand-alone biomarker. Elevated theta-to-alpha ratios and reduced spectral or Lempel-Ziv complexity are replicated EEG correlates of amnestic MCI. Portable dry-electrode hardware can capture these measures in community or home settings [19, 7]. In the study by Boudaya et al., a joint EEG–HRV feature set outperformed either signal alone for MCI detection [7].
Actigraphy measures rest–activity fragmentation, circadian amplitude, and sleep efficiency. These measures have been linked to cognitive status and biological aging, including in cohorts of more than 80,000 participants used to validate the wearable-derived “CosinorAge” biomarker [36, 14]. Such cohort sizes illustrate the scalability of wrist actigraphy relative to EEG and HRV acquisition.
2.3 Multimodal fusion architectures
Early and late fusion are the two most common multimodal designs in the digital-biomarker literature [31, 22]. Early fusion concatenates features before classification, whereas late fusion combines modality-specific classifiers by voting or averaging. A published multimodal MCI framework uses weighted soft voting across cognitive-test and physiological features [22]. Another protocol combines cognitive and wearable physiological features recorded during virtual-reality speech interaction [41]. Beyond cognitive decline, cross-modal and multi-head attention have outperformed static fusion in several physiological-signal and clinical time-series tasks. Applications include EEG-peripheral fusion for emotion recognition [12] and progressive fusion of ECG with phonocardiogram signals for cardiovascular disease detection [23].
Bidirectional attention between EEG connectivity and ECG/HRV features has been used for cognitive-state recognition in flight safety [43], and self-supervised cross-temporal alignment has been used for asynchronous EEG–peripheral streams [20]. Kernel-based discriminant correlation fusion provides a non-attention example in autism-spectrum diagnosis [40]. Related applications include explainable electrophysiology fusion for sleep-stage classification [13] and outcome prediction in disorders of consciousness [4]. Collectively, these studies motivate testing subject-specific attention for cardiac, neural, and behavioral signals against static concatenation and voting.
2.4 Efficient and privacy-preserving on-device learning
Wearable and edge deployment constrain both computation and the transfer of physiological data.
Low-rank reparameterization was developed to adapt Transformer models with fewer trainable parameters [39, 17, 34]; the same subspace principle could reduce the parameter count of modality-specific encoders, as discussed in Section 7. Federated learning addresses data transfer by allowing institutions or device cohorts to train a shared model without centralizing raw data [28]. Federated learning has been applied to stress detection from smart-band heart-activity data [9], and reviews describe its use in privacy-preserving smart healthcare [2, 29]. Section 4 outlines how BioSync could use the same training arrangement.
3 Theoretical Rationale
A fused multimodal signal can contain information that is absent from any single modality. Two complementary arguments motivate this claim.
3.1 A latent-variable measurement model
Consider an unobserved physiological state , such as the degree of central-autonomic-network disruption associated with cognitive decline or the degree of cardiac autonomic neuropathy associated with a metabolic complication. No wearable sensor measures directly. Each modality instead provides a noisy, partial projection,
| (1) |
where is a modality-specific, possibly nonlinear measurement function and is the modality’s noise level. This noise may vary by subject and session, as with motion artifact in EEG or a loose wristband during actigraphy. In a factor-analytic or structural-equation view, an individual modality may estimate consistently but inefficiently. If the are conditionally independent given , classical estimation theory combines noisy estimates through a precision-weighted (inverse-variance-weighted) average,
| (2) |
which has lower variance than an individual when multiple modalities carry information about under the stated assumptions. Kalman-filter sensor fusion uses the same principle. Section 4.3 treats the softmax-normalized attention weights as a learned analogue of : the model can assign more weight to a modality when its representation is more informative for a particular recording. This analogy motivates the architecture but does not make attention weights calibrated estimates of measurement precision.
3.2 Information-theoretic complementarity and synergy
The measurement model justifies fusion when modalities provide independently noisy views of the same latent cause. Multivariate information theory addresses a second case in which information is present only in a joint observation.
Partial information decomposition (PID) separates the mutual information that a set of modalities jointly carries about an outcome into redundant, unique, and synergistic components. Synergistic information appears only in the joint configuration and cannot be recovered from a single modality or from the marginals. For example, an elevated LF/HF ratio may inform cognitive status only conditional on a simultaneous reduction in EEG spectral complexity. Together, the measurements may indicate autonomic dysregulation driven by the same central process that affects cortical dynamics rather than by an unrelated cardiovascular cause. HRV alone, EEG alone, or a parallel analysis without their interaction would miss this component. Fusion can therefore use information that emerges from the joint distribution. Late voting does not model feature-level cross-modal interactions, while linear early fusion represents only additive effects unless interaction terms are supplied explicitly [31, 22].
3.3 Why a Transformer specifically
These arguments motivate a fusion mechanism that can represent non-additive interactions and reweight modalities by subject and recording when signal reliability varies. For the applications considered here, the mechanism should also accept different numbers of modality tokens without changing its core encoder.
Self-attention provides these operations [39]. Each modality embedding becomes a token; multi-head attention computes pairwise token interactions, and softmax-normalized weights produce data-dependent reweighting. Attention operates on a token sequence rather than a fixed concatenated vector, so the same encoder can process the two- and four-token configurations evaluated here. Adding a new modality still requires a corresponding input encoder and retraining. These specific properties, rather than the general approximation capacity of Transformers, motivate BioSync’s small encoder.
3.4 A multi-criteria definition of what makes a composite biomarker good
We assess fusion with six criteria derived from the BEST and V3 frameworks [15, 8, 16] and from the multimodal and information-theoretic arguments above. The criteria include properties not captured by classification accuracy:
- 1.
- 2.
Multimodal information capture: the biomarker should draw on more of the available, complementary information than any single channel alone, including synergistic information recoverable only from the joint distribution of modalities (Section 3).
- 3.
Adaptive interpretability: the contribution of each input channel to the biomarker’s value should be recoverable per subject rather than only as a global coefficient fit across the cohort, so a clinician or researcher can inspect why the model produced a score for a person.
- 4.
Graceful degradation: performance should decline slowly rather than catastrophically when one channel is missing or unreliable, including through artifact corruption, because wearable deployments lose channels through battery, fit, or motion problems [16].
- 5.
Architectural generality: the fusion encoder should accommodate other modality sets and application domains without redesign of its core attention block, because a method restricted to one sensor combination has limited translational value [5].
- 6.
Table 2 compares four designs: single-modality biomarkers, early fusion by concatenation, late fusion by voting as used in published multimodal MCI-screening methods [22, 42], and BioSync (Section 5.9). This comparison is central because accuracy alone does not measure adaptive interpretability, graceful degradation, or architectural generality, and it does not consistently favor BioSync in the synthetic experiments.
4 Proposed BioSync Framework
4.1 Problem formulation
Let a participant be represented by a set of modality-specific feature vectors , extracted from a fixed-length recording window using standard signal-processing pipelines (Section 4.2). BioSync learns a function
| (3) |
that outputs a continuous risk score , defined as the BioSync Index (BSI), and an attention-weight vector , , summarizing each modality’s contribution to the fused token for that subject. Under the BEST taxonomy, the BSI is intended as a susceptibility/risk and monitoring biomarker: a composite, continuously valued indicator of physiological status rather than a diagnostic label [15, 8]. We evaluate this formulation with for cognitive decline and for metabolic-autonomic risk.
4.2 Two application domains and their modality-specific features
Domain 1: cognitive decline ().
HRV. Inter-beat interval series derived from wearable photoplethysmography (PPG) or ECG provide six features: the time-domain measures SDNN, RMSSD, and pNN50; LF and HF power; and the LF/HF ratio. These measures are used in HRV–cognition studies [6, 24].
EEG. Portable dry-electrode EEG recordings provide six features: the theta-to-alpha power ratio, gamma-band spectral entropy, frontal theta power, parietal alpha power, beta-band power, and a Lempel-Ziv complexity index. The feature set follows the qEEG biomarker literature on early cognitive decline [19, 7].
Actigraphy. Wrist accelerometry provides six features: a rest–activity fragmentation index, circadian amplitude, sleep efficiency, step-count coefficient of variation, interdaily stability, and intradaily variability. These measures are used in actigraphy-biomarker studies [36, 14].
Speech. Brief spontaneous-speech recordings provide four acoustic-linguistic features: pause-to-speech ratio, articulation rate, pitch (fundamental-frequency) variability, and semantic density.
Speech digital-biomarker studies associate cognitive decline with increased pause time, slower articulation, reduced prosodic variability, and lower semantic or idea density [10, 42]. A smartphone or wearable microphone can capture these features separately from the cardiac, neural, and behavioral channels. Video-derived motor or gait features could be represented by an additional token, as discussed in Section 7.
Domain 2: metabolic-autonomic risk, structured around the AI-READI schema (). This domain uses variables documented for the AI-READI wearable-activity-monitor data domain [1].
AI-READI is an NIH Bridge2AI project developing a FAIR multimodal dataset. Its public documentation describes more than 1,000 enrolled adults, a target of 4,000, and T2D severity ranging from no T2D to insulin-managed disease. Participants are recruited at three U.S. sites and wear a Garmin Vivosmart 5 during a 10-day home-monitoring period with 5-second sampling [1]. The documentation lists seven channels: heart rate (bpm), oxygen saturation (SpO2%, during sleep only), step count, physical-activity calories, respiration rate derived from HRV, sleep duration, and a proprietary 0–100 device stress score computed from heart rate and HRV. Records use an Open mHealth-derived JSON schema. We group five of these variables into two tokens. The Cardiac-Autonomic token contains resting heart rate, respiration rate, and the device stress score; the Behavioral token contains step count and sleep duration.
The diabetic cardiac autonomic neuropathy (CAN) literature motivates this grouping: resting heart rate rises while HRV-derived measures fall with autonomic damage from chronic hyperglycemia, and daily activity and sleep duration decline with disease burden and complications. This study uses AI-READI’s public documentation rather than its controlled-access raw data. The synthetic metabolic cohort in Section 5 reproduces the documented schema through these two tokens and literature-seeded effect directions. The same feature grouping can be evaluated on raw AI-READI records after data-use approval, as described in Section 7.
Multimodal learning does not require modalities to share a sensor type, sampling rate, or physiological system when each carries information about the same outcome [5] (Section 3). Within each cross-validation fold, every feature is standardized to zero mean and unit variance using training-partition statistics only.
4.3 Transformer fusion encoder
Figure 2 details the attention path from modality-specific measurements and features to the fusion token, BSI, and modality weights. The parallel linear path is introduced separately in Section 4.4 and is included in the high-level overview in Figure 1.
Each modality’s standardized feature vector is linearly projected into a shared -dimensional embedding, treated as one token:
| (4) |
A learned fusion token , analogous to the [CLS] token in a text Transformer, is prepended to the modality sequence, giving . One multi-head self-attention block updates each token by attending to all tokens in the sequence:
| (5) |
computed independently for heads of dimension and recombined with an output projection . A residual connection and layer normalization follow. The block then applies a position-wise feed-forward network with a GELU nonlinearity before the second residual connection and layer normalization. Its fusion-token output row, , supplies the fused representation for the BSI and the attention interpretation:
| (6) |
where is the post-softmax weight from the fusion token to modality token in head . Because the fusion-token query also attends to its own key, the raw modality weights sum to slightly less than one. We therefore report , renormalized so that , as a per-subject modality-relevance score. The construction computes token similarity after learned projections and represents the pairwise interactions described in Section 3.3. The same attention block processes the four-token cognitive-decline domain and the two-token metabolic-autonomic domain; each domain retains its own modality-specific input projections.
4.4 A wide-and-deep hybrid containing the concatenation baseline
Sections 3 and 3.3 motivate attention for cross-modal interactions, but they do not imply that pure attention is the most direct way to recover the unique and redundant information captured by a linear model. A small Transformer must learn a linear decision boundary that logistic regression represents explicitly. We therefore use the wide-and-deep design introduced by Cheng et al. for recommender systems [11]. A linear (“wide”) path processes the concatenated standardized features in parallel with the Transformer (“deep”) path, and the two contributions are summed before the sigmoid,
| (7) |
where gradient descent jointly learns , , , and the Transformer parameters. Setting recovers the concatenation model, so the wide-and-deep hypothesis class contains the baseline’s linear decision functions. Setting recovers the pure-attention model in Section 4.3. This containment concerns representational capacity; finite-sample optimization and regularization can still yield lower held-out performance. The experiments use BioSync to denote the full wide-and-deep architecture and compare it with both components (Sections 5.4 and 5.7).
4.5 Toward federated, privacy-preserving training
Because each modality encoder processes a local feature vector, BioSync is compatible in principle with federated training. Clinics or wearable cohorts could train local model copies and share parameter updates with an aggregator under federated averaging [28], as demonstrated in wearable and IoT biomedical monitoring [9, 2, 29].
Low-rank reparameterization could reduce the trainable parameter count of the modality encoders and attention projections. This proposed extension follows subspace adaptation for larger Transformer models [17, 33, 34] and is discussed in Section 7.
Federated training is not implemented or evaluated in this study. A future experiment could train across AI-READI’s controlled-access records, the cognitive-decline cohort under collection at our laboratory, and other wearable cohorts without centralizing raw physiological recordings.
5 Experimental Validation
5.1 Rationale and scope
The experiments use two synthetic, literature-informed cohorts to examine the six criteria in Section 3.4. Section 5.5 evaluates graded validity; Sections 5.4–5.5 evaluate multimodal information capture and modality-level interpretability; and Section 5.6 evaluates graceful degradation. Using the same attention block for four- and two-modality domains tests architectural generality. On-device and federated feasibility are assessed from model structure rather than deployment measurements (Section 6). No patient data were used, and the reported metrics do not estimate clinical diagnostic performance.
5.2 Cohort construction
Cognitive cohort (). We generated profiles: 180 healthy-leaning and 180 MCI-leaning. Each profile received a continuous latent severity value , drawn from Beta(2,6) for the healthy-leaning group and Beta(4,2.2) for the MCI-leaning group. The overlapping distributions avoid trivial class separation. Twenty-two features—six HRV, six EEG, six actigraphy, and four speech—were generated as linear functions of severity plus independent Gaussian noise. The feature definitions and literature-seeded effect directions are given in Section 4.2.
Metabolic cohort, AI-READI schema (). An independent set of profiles contained 180 lower-severity and 180 higher-severity T2D-leaning cases. A latent glycemic/autonomic severity value was drawn from the same Beta(2,6) and Beta(4,2.2) distributions.
Five features were generated as linear functions of severity plus Gaussian noise and grouped into two AI-READI-schema tokens: Cardiac-Autonomic (resting heart rate, respiration rate, and device stress score) and Behavioral (step count and sleep duration). Literature-seeded directions specify increasing resting heart rate, respiration rate, and stress score, and decreasing step count and sleep duration, with increasing severity.
In both cohorts, the continuous severity value is not used as a training target; it is reserved for the correlation analysis in Section 5.5. Features are standardized within each cross-validation fold using training-partition statistics only.
5.3 Experimental setup
We used 5-fold stratified cross-validation in each cohort. The evaluated models were (1) separate logistic-regression baselines for each modality or token; (2) early fusion by logistic regression on all concatenated standardized features; and (3) the BioSync wide-and-deep encoder with embedding width , attention heads, and one encoder block (Section 4.4). BioSync was implemented in JAX with automatic differentiation. We compared Adam with plain gradient descent and found the latter more stable at this dataset and model scale after tuning the learning rate. All reported BioSync results therefore use plain gradient descent for 500 epochs with weight decay. Accuracy, F1-score, and AUC are reported as means across the five folds. The severity analysis uses pooled out-of-fold BSI values. We also evaluated BioSync without the wide linear path and swept embedding width and head count (Section 5.7).
5.4 Main results
Table 1 reports the cross-validated results. In the cognitive cohort, single-modality accuracy ranged from 0.753 to 0.822 and AUC from 0.836 to 0.918. Concatenation exceeded every single-modality model, with accuracy 0.850 and AUC 0.926. BioSync had the highest AUC (0.928), while its accuracy (0.842) and F1-score (0.839) were numerically below concatenation (0.850 and 0.847) and differed by less than one fold-level standard deviation. In the metabolic AI-READI-schema cohort, the Cardiac-Autonomic token (accuracy 0.728, AUC 0.799) outperformed the Behavioral token (accuracy 0.639, AUC 0.696). Concatenation reached accuracy 0.756 and AUC 0.823. BioSync had the highest accuracy (0.764) and F1-score (0.766), whereas its AUC (0.814) was lower than concatenation’s (0.823). Figure 3 shows the pooled out-of-fold ROC curves.
Across the six cohort–metric combinations, BioSync ranks first for cognitive-cohort AUC and for metabolic-cohort accuracy and F1-score. The remaining differences are within or near one cross-validation standard deviation; these samples of do not support claims of large effects.
We use AUC as the primary discrimination metric because the BSI is continuous and AUC does not depend on a selected classification threshold. AUC is also commonly reported in the digital-biomarker literature reviewed in Section 2 [31]. On this metric, BioSync ranks first in the cognitive cohort; in the metabolic cohort, the fold-level variation does not support a clear difference from concatenation.
BioSync without the linear path did not exceed concatenation (Section 5.7). On cognitive-cohort AUC, adding the wide path increased performance from 0.911 for pure attention to 0.928 for the full model, compared with 0.926 for concatenation.
| Cohort | Model | Accuracy | F1-score | AUC |
| Cognitive () | HRV only | 0.800 (0.045) | 0.798 (0.043) | 0.884 (0.044) |
| EEG only | 0.817 (0.034) | 0.813 (0.037) | 0.897 (0.034) | |
| Actigraphy only | 0.822 (0.034) | 0.821 (0.037) | 0.918 (0.019) | |
| Speech only | 0.753 (0.028) | 0.750 (0.034) | 0.836 (0.035) | |
| Early fusion (concatenation) | 0.850 (0.034) | 0.847 (0.036) | 0.926 (0.030) | |
| BioSync (wide-and-deep) | 0.842 (0.044) | 0.839 (0.044) | 0.928 (0.028) | |
| Metabolic, AI-READI schema () | Cardiac-Autonomic only | 0.728 (0.047) | 0.721 (0.058) | 0.799 (0.038) |
| Behavioral only | 0.639 (0.035) | 0.633 (0.041) | 0.696 (0.038) | |
| Early fusion (concatenation) | 0.756 (0.032) | 0.758 (0.030) | 0.823 (0.026) | |
| BioSync (wide-and-deep) | 0.764 (0.018) | 0.766 (0.018) | 0.814 (0.035) |
5.5 BSI and simulated severity
Pooled out-of-fold BSI values correlated with latent severity, which was not used as a training target: () in the cognitive cohort and () in the metabolic AI-READI-schema cohort (Figure 4). These correlations assess the graded validity criterion in Section 3.4. A linear model trained directly to regress concatenated features on severity achieved higher correlations ( cognitive and metabolic), as expected for a model optimized on that target; the comparison therefore does not favor BioSync. The renormalized fusion-token attention weights in Figure 5 are close to uniform: approximately 0.25 per modality in the cognitive cohort and 0.50 per token in the metabolic cohort. This pattern is consistent with the synthetic design, in which modality reliability is comparable and does not vary systematically by subject. The absence of weight collapse provides a clean-data reference for the heterogeneous-reliability experiment in Section 5.6, but uniform mean weights alone do not demonstrate adaptive interpretability.
5.6 Robustness to heterogeneous modality reliability
The rationale in Section 3.3 predicts that data-dependent weighting should be most useful when modality reliability varies across recordings. To test this condition, we replace one randomly selected modality with high-variance noise for each test subject with probability , termed the corruption rate. This procedure approximates complete channel failure rather than the detailed temporal structure of device slippage or motion artifact. During training, both BioSync and concatenation receive identical modality-dropout augmentation, with each modality corrupted independently with probability 0.3. Test-time differences therefore follow from how the two models use the same corruption exposure.
Figure 6 reports AUC for . In the cognitive cohort, the models are similar at (0.929 BioSync vs. 0.931 concatenation). BioSync leads at every nonzero rate, reaching 0.891 versus 0.865 at , and its margin increases monotonically. Results in the metabolic cohort are less stable. Concatenation leads at (0.824 vs. 0.794) and (0.771 vs. 0.766); the models are effectively tied at (0.775 vs. 0.774); BioSync leads at (0.765 vs. 0.741); concatenation leads at (0.738 vs. 0.722); and BioSync leads at (0.663 vs. 0.657). Corrupting one of two tokens affects a larger fraction of the metabolic input than corrupting one of four cognitive tokens, and the sample does not resolve each small difference. Thus, the cognitive cohort supports the predicted attention advantage at five of six rates, whereas the metabolic cohort shows no monotonic ordering. Wearable channels can be lost through battery depletion, poor skin contact, motion artifact, or device removal [16]; the matched experiment measures one simplified form of the “graceful degradation” criterion in Table 2.
5.7 Architecture ablation
The first ablation isolates the linear path introduced in Section 4.4. In the cognitive cohort, pure-attention BioSync obtained accuracy 0.847 and AUC 0.911, while the wide-and-deep model obtained 0.842 and 0.928. The full model therefore exceeded both pure attention (0.911) and concatenation (0.926) on AUC, the primary discrimination metric, but not on accuracy. Fold-to-fold variation does not support a large accuracy difference among the three models.
We also varied cognitive-cohort embedding width and attention heads , subject to (Figure 7). Across the grid, accuracy ranged from 0.831 to 0.858 and AUC from 0.915 to 0.932; the prespecified model obtained 0.842/0.928 in Table 1. A 3-fold sweep produced the highest point estimates at (accuracy 0.858, AUC 0.932). We retain because the robustness experiment preceded the sweep and changing the reported model afterward would constitute post hoc selection. Every wide-and-deep grid configuration exceeded the pure-attention AUC, indicating that the result is not confined to one width or head count.
5.8 Positioning against published digital biomarkers
The main results and ablations above provide the controlled comparisons within this study. Published digital-biomarker methods use different cohorts, tasks, feature sets, and performance measures, so they can serve only as descriptive reference points.
Figure 8 places BioSync’s cognitive-cohort AUC of 0.928 beside five values from the literature reviewed in Section 2. Teh et al. report pooled sensitivity and specificity of approximately 0.80 in a meta-analysis of MCI and pre-frailty digital biomarkers [37]. Qi et al. report mean AUCs of 0.821 across 45 MCI-focused AI models and 0.887 across 21 AD-focused AI models in a review of 431 studies [31].
Li et al. report 0.889 accuracy for late fusion of digital cognitive tests with wearable measures [22], and Xu et al. report 0.918 accuracy for late fusion of language digital biomarkers [42]. BioSync’s synthetic-cohort AUC is numerically higher than these five plotted values. Because the references include AUC, accuracy, sensitivity, and specificity from different datasets, the ordering is descriptive rather than inferential.
This comparison does not establish superiority over a published method; it defines reference values for subsequent real-cohort evaluation.
5.9 Comparison against alternative fusion designs
Table 2 compares single-modality biomarkers, early fusion, late fusion as used in published multimodal designs [22, 42], and BioSync on the six criteria from Section 3.4. BioSync is the only design marked as satisfying all six within the scope of the present architectural analysis. This rating does not mean that it wins every classification metric: concatenation has higher cognitive-cohort accuracy and F1-score and higher metabolic-cohort AUC. Concatenation and voting also lack an intrinsic per-subject modality-weight output. The graceful-degradation rating is based specifically on the matched corruption test, in which BioSync leads at five of six cognitive-cohort rates; it does not imply an advantage for every failure pattern or for every metabolic-cohort rate.
| Graded | Multimodal | Adaptive | Graceful | Architectural | On-device/ | |
|---|---|---|---|---|---|---|
| validity | info. capture | interpretability | degradation | generality | federated | |
| Single modality | ✗ | ✗ | n/a | ✗ | ✓ | |
| Early fusion (concat.) | ✓ | ✗ | (Sec. 5.6) | ✓ | ||
| Late fusion (voting) | ✗ | ✗ | ✗ | |||
| BioSync (ours) | ✓ | ✓ | ✓ | ✓ (Sec. 5.6) | ✓ (Sec. 4.2) | ✓ (Sec. 4) |
6 Discussion
BioSync’s clean-data result depends on the metric. It ranked first on cognitive-cohort AUC and on metabolic-cohort accuracy and F1-score, while concatenation ranked first on the other three cohort–metric combinations (Sections 5 and 5.4). All differences were small relative to fold-level variation. Pure attention did not exceed concatenation on cognitive-cohort AUC (Section 5.7); adding the linear path raised AUC from 0.911 to 0.928. This result is consistent with the architectural rationale: the linear branch represents the concatenation baseline directly, while the attention branch can model interactions (Section 4.4).
The corruption experiment more directly tests the proposed role of attention under variable modality reliability. BioSync led at five of six cognitive-cohort rates, with an increasing margin as corruption rose. The two-token metabolic result was not monotonic: concatenation led at low corruption, and BioSync led only at selected intermediate rates and at . Both cohorts used the same prespecified corruption protocol, training augmentation, and test rates. Complete loss of a wearable channel can occur through battery depletion, poor skin contact, motion artifact, or device removal [16], but the high-variance replacement used here remains a simplified failure model.
Classification metrics do not assess graded validity or case-level modality relevance. BioSync reports a modality-weight vector for each subject, whereas logistic-regression coefficients are fixed across the cohort (Table 2). The weights should be interpreted as attention-based relevance scores, not causal contributions or calibrated sensor reliabilities. Although BioSync was trained on class labels rather than severity, the BSI correlated with latent severity at and . Direct linear regression on severity produced higher correlations (Section 5.5), so this analysis shows association with the simulated continuum rather than an advantage over a severity-specific model.
Hand-crafted summary features may limit what the attention path can learn. Prior cross-modal studies report larger gains over static fusion when attention receives raw or intermediate temporal representations [12, 23, 20] (Section 2). Replacing the linear modality projections with small 1D-CNN or GRU encoders would allow windowed HRV, EEG, actigraphy, speech, or AI-READI time series to enter the deep path directly (Section 7). The wide branch could continue to process summary features, retaining explicit access to linear effects while the temporal branch models within-channel structure.
Only 24 of the 431 AD digital-biomarker studies reviewed by Qi et al. used multimodal signals, and external validation was uncommon [31]. Within that context, BioSync contributes a stated fusion rationale, an explicit linear baseline within the model, subject-specific attention weights, a continuous index, a matched corruption analysis, and one attention block evaluated with two- and four-token inputs. Table 2 compares these computational properties with other fusion designs. Their clinical value cannot be determined from synthetic data.
7 Limitations and Future Work
Five limitations bound the interpretation of these results. First, all quantitative results come from synthetic cohorts, not patients, raw AI-READI records, or prospective wearable recordings. The accuracy, F1, and AUC values in Table 1 characterize behavior under literature-seeded assumptions rather than diagnostic performance. Second, the modality encoders process hand-crafted summaries instead of raw or windowed signals, excluding within-channel temporal structure from the deep path. Third, clean-data differences from concatenation are small, and several are less than one cross-validation standard deviation. Larger real cohorts are required for statistical comparisons. Fourth, the robustness experiment replaces an entire modality with high-variance noise. Actual failures may be partial, temporally structured, or correlated; EEG motion artifact, for example, is not spectrally equivalent to random noise. Fifth, the proposed federated-training extension has not been implemented or measured.
Prospective validation is planned in two domains. For cognitive decline, Connected Future Labs in Reno, Nevada, operates a collection pipeline using Empatica wristbands and Muse EEG headsets. The wristbands record PPG, electrodermal activity, accelerometry, and skin temperature; the study is designed to distinguish participants with MCI from cognitively healthy participants. An AWS pipeline supports ingestion, quality auditing, and model development. Passive smartphone speech samples would reproduce the four-modality configuration evaluated here. Video-derived gait or motor features could form a fifth token; gait is the most common modality in AD/MCI digital-biomarker research [31].
For metabolic-autonomic risk, the AI-READI wearable-activity-monitor domain documents the seven Garmin Vivosmart 5 channels used to structure the synthetic cohort [1]. Raw wearable records are distributed through FAIRhub under a data-use agreement. This study uses only the public schema and does not analyze the controlled-access JSON records.
The validation plan has six components. Temporal CNN or GRU encoders will process windowed recordings in the deep path while the wide path retains summary features. Public multimodal datasets, including WESAD [35] and MMASH [32], will provide intermediate tests, with published validation evidence guiding use of the Empatica platform [27]. Analytical and clinical validation will compare the BSI with clinician-adjudicated cognitive status and with T2D severity or complication status under the BEST and V3 frameworks [15, 8, 16]. Device logs will support empirical artifact and missingness models in place of high-variance noise. Federated experiments across sessions, sites, and cohorts will measure the accuracy–privacy trade-off. Finally, low-rank deep-path encoders and attention projections will be evaluated for on-device processing of temporal input [17, 34].
8 Conclusion
BioSync maps heterogeneous physiological, behavioral, and speech features to a continuous BioSync Index. Its attention path models interactions among modality tokens, while its linear path contains the decision functions available to feature concatenation (Section 4.4). The design is motivated by precision-weighted latent-variable measurement and by information that may occur only in joint modality configurations.
We evaluated the same attention block in a four-modality cognitive-decline cohort and a two-token metabolic-autonomic cohort structured around the public AI-READI wearable schema. In the cognitive cohort, BioSync produced AUC 0.928 versus 0.926 for concatenation; in the metabolic cohort, it produced accuracy/F1 of 0.764/0.766 versus 0.756/0.758. Concatenation remained higher on cognitive accuracy/F1 and metabolic AUC. The BSI correlated with latent severity at and . With matched modality-dropout training, BioSync led at five of six cognitive-cohort corruption rates and at the highest metabolic-cohort rate, although the metabolic ordering was not monotonic.
These results specify hypotheses for real-cohort validation rather than clinical conclusions. The cognitive-cohort AUC is numerically higher than five published reference values, but differences in datasets, tasks, and metrics prevent a controlled benchmark claim. The next tests are to estimate discrimination, calibration, severity association, modality relevance, and corruption robustness on controlled-access AI-READI records and on the prospective cognitive-decline cohort. Those experiments will determine whether the computational behavior observed in simulation transfers to multimodal monitoring data.
CRediT authorship contribution statement
Seyed Mahmoud Sajjadi Mohammadabadi: Conceptualization, Methodology, Software, Formal analysis, Investigation, Writing – original draft, Writing – review & editing, Visualization.
Declaration of competing interest
The author declares that he has no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
Funding
This research did not receive any specific grant from public or commercial funding agencies or from not-for-profit organizations.
Data availability
The synthetic data-generation code and analysis scripts used to produce Table 1 are available from the corresponding author upon reasonable request. No real patient data were used in this study.
References
- [1] (2024) AI-READI: rethinking AI data collection, preparation and sharing in diabetes research and beyond. Nature Metabolism. Cited by: §1, §4.2, §4.2, §7.
- [2] (2023) Federated learning for privacy preservation in smart healthcare systems: a comprehensive survey. IEEE Journal of Biomedical and Health Informatics 27 (2), pp. 778–789. External Links: Document Cited by: §1, §2.4, §4.5.
- [3] (2023) 2023 Alzheimer’s disease facts and figures. Alzheimers. Dement. 19 (4), pp. 1598–1695. External Links: Document Cited by: §1.
- [4] (2024) Multimodal prediction of 3- and 12-month outcomes in ICU patients with acute disorders of consciousness. Neurocritical Care 40 (2), pp. 718–733. External Links: Document Cited by: §2.3.
- [5] (2019) Multimodal machine learning: a survey and taxonomy. IEEE Transactions on Pattern Analysis and Machine Intelligence 41 (2), pp. 423–443. External Links: Document Cited by: item 5, §4.2.
- [6] (2024) Heart rate variability during routine cognitive testing: preliminary analyses of time and frequency domain metrics. Alzheimers. Dement. 20 (S2). External Links: Document Cited by: §1, §2.2, §4.2.
- [7] (2024) Mild cognitive impairment detection based on EEG and HRV data. Digital Signal Processing 147 (104399), pp. 104399. External Links: Document Cited by: §1, §1, §2.2, §4.2.
- [8] (2018) Biomarker definitions and their applications. Exp. Biol. Med. (Maywood) 243 (3), pp. 213–221. External Links: Document Cited by: §1, §1, §2.1, item 1, §3.4, §4.1, §7.
- [9] (2021) Privacy-preserving federated deep learning for wearable IoT-based biomedical monitoring. ACM Transactions on Internet Technology 21 (1), pp. 1–17. External Links: Document Cited by: §1, §2.4, §4.5.
- [10] (2024) Harnessing speech-derived digital biomarkers to detect and quantify cognitive decline severity in older adults. Gerontology 70 (4), pp. 429–438. External Links: Document Cited by: §4.2.
- [11] (2016) Wide & deep learning for recommender systems. In Proceedings of the 1st Workshop on Deep Learning for Recommender Systems, pp. 7–10. Cited by: §4.4.
- [12] (2025) Multimodal physiological signal emotion recognition based on multi-head cross attention with representation learning. Frontiers in Psychiatry 16 (1713559), pp. 1713559. External Links: Document Cited by: §1, §2.3, §6.
- [13] (2021) Explainable sleep stage classification with multimodal electrophysiology time-series. In Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), pp. 2363–2366. External Links: Document Cited by: §2.3.
- [14] (2025) From physical activity patterns to cognitive status: development and validation of novel digital biomarkers for cognitive assessment in older adults. Int. J. Behav. Nutr. Phys. Act. 22 (1), pp. 11. External Links: Document Cited by: §1, §2.2, §4.2.
- [15] (2016) BEST (Biomarkers, EndpointS, and other Tools) resource. Technical report Food and Drug Administration (US), Silver Spring (MD), U.S. Food and Drug Administration and National Institutes of Health, Silver Spring, MD. Cited by: §1, §2.1, item 6, §3.4, §4.1, §7.
- [16] (2020) Verification, analytical validation, and clinical validation (v3): the foundation of determining fit-for-purpose for biometric monitoring technologies (biomets). NPJ Digital Medicine 3 (1), pp. 55. External Links: Document Cited by: §2.1, item 4, §3.4, §5.6, §6, §7.
- [17] (2021) LoRA: low-rank adaptation of large language models. arXiv preprint arXiv:2106. 09685 abs/2106.09685. External Links: Document Cited by: §2.4, §4.5, §7.
- [18] (2018) NIA-AA research framework: toward a biological definition of Alzheimer’s disease. Alzheimers. Dement. 14 (4), pp. 535–562. External Links: Document Cited by: §1.
- [19] (2023) Neurophysiological markers in community-dwelling older adults with mild cognitive impairment: an EEG study. Alzheimer’s Research & Therapy 15 (1), pp. 217. External Links: Document Cited by: §1, §2.2, §4.2.
- [20] (2026) Cross-temporal attention fusion (CTAF) for multimodal physiological signals in self-supervised learning. arXiv preprint arXiv:2602. 02784. External Links: 2602.02784, Document Cited by: §1, §2.3, §6.
- [21] (2019) Digital biomarkers for Alzheimer’s disease: the mobile/wearable devices opportunity. NPJ Digital Medicine 2 (1), pp. 9. External Links: Document Cited by: §1, §2.1.
- [22] (2023) Synergy through integration of digital cognitive tests and wearable devices for mild cognitive impairment screening. Frontiers in Human Neuroscience 17, pp. 1183457. External Links: Document Cited by: §1, §2.3, §3.2, §3.4, §5.8, §5.9.
- [23] (2025) A progressive attention-based cross-modal fusion network for cardiovascular disease detection using synchronized electrocardiogram and phonocardiogram signals. PeerJ Computer Science 11, pp. e3038. Cited by: §1, §2.3, §6.
- [24] (2014) Inter-modality relationship constrained multi-modality multi-task feature selection for Alzheimer’s disease and mild cognitive impairment identification. NeuroImage 84, pp. 466–475. External Links: Document Cited by: §1, §2.2, §4.2.
- [25] (2020) Dementia prevention, intervention, and care: 2020 report of the Lancet commission. Lancet 396 (10248), pp. 413–446. External Links: Document Cited by: §1.
- [26] (2026) Limited evidence for heart rate variability as a predictor of cognitive and pathophysiological brain markers. Journal of Alzheimer’s Disease 109 (4), pp. 1723–1732. External Links: Document Cited by: §1, §2.2.
- [27] (2016) Validation of the Empatica E4 wristband. In IEEE EMBS International Student Conference (ISC), pp. 1–4. External Links: Document Cited by: §7.
- [28] (2016) Communication-efficient learning of deep networks from decentralized data. In International Conference on Artificial Intelligence and Statistics, pp. 1273–1282. External Links: Document Cited by: §1, §2.4, item 6, §4.5.
- [29] (2025) Blockchain-enabled privacy-preserving second-order federated edge learning in personalized healthcare. IEEE Trans. Consum. Electron. 71 (4), pp. 9983–9992. External Links: Document Cited by: §1, §2.4, §4.5.
- [30] (2019) Current state of digital biomarker technologies for real-life, home-based monitoring of cognitive function for mild cognitive impairment to mild Alzheimer disease and implications for clinical care: systematic review. Journal of Medical Internet Research 21 (8), pp. e12785. External Links: Document Cited by: §1, §2.1.
- [31] (2025) Alzheimer’s disease digital biomarkers multidimensional landscape and AI model scoping review. NPJ Digital Medicine 8 (1), pp. 366. External Links: Document Cited by: §1, §1, §2.1, §2.3, §3.2, §5.4, §5.8, §6, §7.
- [32] (2020) Multilevel monitoring of activity and sleep in healthy people (MMASH). Scientific Data 7, pp. 161. Cited by: §7.
- [33] (2025) A survey of large language models: evolution, architectures, adaptation, benchmarking, applications, challenges, and societal implications. Electronics 14 (18), pp. 3580. External Links: Document Cited by: §4.5.
- [34] (2026) SOLAR: subspace-oriented lightweight adapter reparameterization for scalable PEFT compression. Under review. Cited by: §2.4, §4.5, §7.
- [35] (2018) Introducing WESAD, a multimodal dataset for wearable stress and affect detection. pp. 400–408. External Links: Document Cited by: §7.
- [36] (2024) Circadian rhythm analysis using wearable-based accelerometry as a digital biomarker of aging and healthspan. NPJ Digital Medicine 7 (1), pp. 146. External Links: Document Cited by: §1, §2.2, §4.2.
- [37] (2022) Predictive accuracy of digital biomarker technologies for detection of mild cognitive impairment and pre-frailty amongst older adults: a systematic review and meta-analysis. IEEE Journal of Biomedical and Health Informatics 26 (8), pp. 3638–3648. External Links: Document Cited by: §1, §5.8.
- [38] (2022) Digital biomarkers: convergence of digital health technologies and biomarkers. NPJ Digital Medicine 5 (1), pp. 36. External Links: Document Cited by: §1.
- [39] (2017) Attention is all you need. Advances in Neural Information Processing Systems 30, pp. 5998–6008. External Links: Document Cited by: §2.4, §3.3.
- [40] (2024) Multimodal kernel-based discriminant correlation analysis data-fusion approach: an automated autism spectrum disorder diagnostic system. Physical and Engineering Sciences in Medicine 47 (1), pp. 361–369. External Links: Document Cited by: §2.3.
- [41] (2023) Screening for mild cognitive impairment with speech interaction based on virtual reality and wearable devices. Brain Sciences 13 (8), pp. 1222. External Links: Document Cited by: §2.3.
- [42] (2025) Biomarkers. Alzheimers. Dement. 21 Suppl 2 (S2), pp. e099209. External Links: Document Cited by: §1, §2.1, §3.4, §4.2, §5.8, §5.9.
- [43] (2026) Detection of pilots’ cognitive states based on cross-modal physiological signal fusion. Electronic Measurement Technology 49 (6), pp. 146–155. Cited by: §1, §2.3.