for the article Diphthong Dynamics under Lexical Tone: Cross-Dialectal Evidence for Category-Specific f0--F1 Coupling
Authors: Chenyu Li and Jalal Al-Tamimi from Laboratoire de linguistique formelle, CNRS, Paris 75013, France
The supplementary material follows the main stages of the analysis. S1 documents data screening and f0 trackability. S2 reports the GAMM analyses and the GAMM-based cascade reconstruction. S3 reports the primary pooled 2 × 2 z-scale fPCA analyses, including bridge and relative-configuration results. S4 provides the speaker-centered non-z robustness reruns. S5 reports the full-curve FoF analyses and residualized tone-recovery metrics. S6 reports the supplementary seven-group fPCA analyses and their non-z reruns, including the multivariate check in S6g. S7 gathers reconstruction-oriented visual summaries across methods.
Full GAMM model summaries, diagnostics, and complete compareML outputs are retained in the accompanying analysis notebook: SuppPub2_GAMM_summaries.html.
Forty-eight native speakers participated: 24 from Xi'an (12 female, 12 male; 18–65 years; mean 42.875, SD 13.19) and 24 from Tianjin (12 female, 12 male; 20–65 years; mean 44.08, SD 14.10). All were native speakers of their local Mandarin variety and reported regular daily use of Standard Mandarin (SM). Local-variety background and SM proficiency were assessed using a questionnaire, self-report, and experimenter evaluation.
Each participant produced both SM and their local variety. Within each location, half completed the SM block first and half the local-variety block first. Before each block, participants read several naturalistic sentences in the intended variety to establish the corresponding speech mode. Target items were randomized within each variety block. Each item was repeated twice within the two carrier sentences described below. Misproduced items were repeated; persistently inaccurate tokens were excluded. After recording, two native SM speakers conducted a blind check of tone-category accuracy across diphthongs.
Recordings took place in a quiet sound-insulated environment using a Zoom H6 Handy Recorder and Shure WH20 head-mounted dynamic microphone, at 48 kHz and 24-bit resolution. The onset inventory was restricted to non-aspirated bilabial and alveolar consonants to limit anticipatory coarticulation and aspiration-related glottal effects. The two carrier sentences provided contrasting preceding tonal contexts within each variety and minimized systematic influences from the following syllable. Acoustic alignment and tracking procedures are detailed in S1f.
The stimulus materials were elicited in two carrier sentences: “Please read X twice” (请将X读两遍) and “Please play X eight times” (请播放X八遍) corresponding to IPA forms tɕʰjəŋT3 tɕjaŋT1 X tuT2 ljaŋT3 pjɛnT4 and tɕʰjəŋT3 poT1 faŋT4 X paT1 pjɛnT4. The item space combined two falling diphthongs (/ai/, /au/), four non-aspirated bilabial or alveolar onsets (/m/, /p/, /n/, /t/), and the four tonal categories (T1–T4), yielding 32 theoretically possible onset-by-diphthong-by-tone slots. Because several slots lacked suitable lexical items with attested characters, the realized stimulus set contained 28 items.
The realized items were:
| Tone | mai | mau | nai | nau | pai | pau | tai | tau |
|---|---|---|---|---|---|---|---|---|
| T1 | mai1 麦 | mau1 猫 | NA | nau1 孬 | pai1 掰 | pau1 包 | tai1 呆 | tau1 刀 |
| T2 | mai2 埋 | mau2 毛 | NA | nau2 挠 | pai2 白 | pau2 薄 | NA | NA |
| T3 | mai3 买 | mau3 铆 | nai3 奶 | nau3 脑 | pai3 摆 | pau3 宝 | tai3 歹 | tau3 岛 |
| T4 | mai4 卖 | mau4 帽 | nai4 奈 | nau4 闹 | pai4 败 | pau4 爆 | tai4 带 | tau4 到 |
The functional analyses (fPCA analyses and function-on-function mapping) and the GAMMs were based on overlapping but not identical datasets. Because the functional analyses treated each token as a full trajectory, they required sufficiently continuous f0 curves: tokens with fewer than five observed f0 points were excluded, and any remaining missing values were filled by linear interpolation with edge filling. The GAMMs, by contrast, were fitted to broader observationwise datasets and therefore did not require the same degree of token-level f0 continuity.
| Dataset group | F1 observations in baseline GAMMs | f0 observations in baseline GAMMs | Mechanistic F1 trajectory observations | Total tokens | Retained functional analysis tokens | Excluded tokens |
|---|---|---|---|---|---|---|
| Tianjin-speaker dataset | 29,513 | 23,230 | 23,230 | 2,683 | 2,403 | 280 |
| Xi'an-speaker dataset | 28,501 | 23,087 | 23,087 | 2,591 | 2,357 | 234 |
| Combined total | 58,014 | 46,317 | 46,317 | 5,274 | 4,760 | 514 |
| Dataset group | Variety | Total tokens | Retained tokens for functional analysis | Excluded tokens | Exclusion rate |
|---|---|---|---|---|---|
| Tianjin-speaker dataset | SM | 1,347 | 1,201 | 146 | 0.108 |
| Tianjin-speaker dataset | TJ | 1,336 | 1,202 | 134 | 0.100 |
| Xi'an-speaker dataset | SM | 1,346 | 1,191 | 155 | 0.115 |
| Xi'an-speaker dataset | XA | 1,245 | 1,166 | 79 | 0.063 |
| Dataset group | Variety | Tone | Diphthong | Total tokens | Retained tokens | Excluded tokens | Exclusion rate |
|---|---|---|---|---|---|---|---|
| Tianjin-speaker dataset | SM | 1 | ai | 132 | 132 | 0 | 0 |
| Tianjin-speaker dataset | SM | 1 | au | 185 | 185 | 0 | 0 |
| Tianjin-speaker dataset | SM | 2 | ai | 95 | 95 | 0 | 0 |
| Tianjin-speaker dataset | SM | 2 | au | 150 | 149 | 1 | 0.007 |
| Tianjin-speaker dataset | SM | 3 | ai | 195 | 131 | 64 | 0.328 |
| Tianjin-speaker dataset | SM | 3 | au | 188 | 121 | 67 | 0.356 |
| Tianjin-speaker dataset | SM | 4 | ai | 210 | 203 | 7 | 0.033 |
| Tianjin-speaker dataset | SM | 4 | au | 192 | 185 | 7 | 0.036 |
| Tianjin-speaker dataset | TJ | 1 | ai | 118 | 93 | 25 | 0.212 |
| Tianjin-speaker dataset | TJ | 1 | au | 168 | 147 | 21 | 0.125 |
| Tianjin-speaker dataset | TJ | 2 | ai | 97 | 97 | 0 | 0 |
| Tianjin-speaker dataset | TJ | 2 | au | 160 | 160 | 0 | 0 |
| Tianjin-speaker dataset | TJ | 3 | ai | 199 | 155 | 44 | 0.221 |
| Tianjin-speaker dataset | TJ | 3 | au | 190 | 151 | 39 | 0.205 |
| Tianjin-speaker dataset | TJ | 4 | ai | 217 | 216 | 1 | 0.005 |
| Tianjin-speaker dataset | TJ | 4 | au | 187 | 183 | 4 | 0.021 |
| Xi'an-speaker dataset | SM | 1 | ai | 123 | 123 | 0 | 0 |
| Xi'an-speaker dataset | SM | 1 | au | 188 | 188 | 0 | 0 |
| Xi'an-speaker dataset | SM | 2 | ai | 89 | 89 | 0 | 0 |
| Xi'an-speaker dataset | SM | 2 | au | 148 | 148 | 0 | 0 |
| Xi'an-speaker dataset | SM | 3 | ai | 202 | 120 | 82 | 0.406 |
| Xi'an-speaker dataset | SM | 3 | au | 189 | 132 | 57 | 0.302 |
| Xi'an-speaker dataset | SM | 4 | ai | 214 | 206 | 8 | 0.037 |
| Xi'an-speaker dataset | SM | 4 | au | 193 | 185 | 8 | 0.041 |
| Xi'an-speaker dataset | XA | 1 | ai | 81 | 59 | 22 | 0.272 |
| Xi'an-speaker dataset | XA | 1 | au | 146 | 103 | 43 | 0.295 |
| Xi'an-speaker dataset | XA | 2 | ai | 69 | 69 | 0 | 0 |
| Xi'an-speaker dataset | XA | 2 | au | 164 | 164 | 0 | 0 |
| Xi'an-speaker dataset | XA | 3 | ai | 199 | 191 | 8 | 0.040 |
| Xi'an-speaker dataset | XA | 3 | au | 193 | 187 | 6 | 0.031 |
| Xi'an-speaker dataset | XA | 4 | aj | 195 | 195 | 0 | 0 |
| Xi'an-speaker dataset | XA | 4 | au | 198 | 198 | 0 | 0 |
| Dataset group | Variety | Tone | Diphthong | Total tokens | Retained tokens | Excluded tokens | Exclusion rate |
|---|---|---|---|---|---|---|---|
| Xi'an-speaker dataset | SM | 3 | ai | 202 | 120 | 82 | 0.406 |
| Tianjin-speaker dataset | SM | 3 | au | 188 | 121 | 67 | 0.356 |
| Tianjin-speaker dataset | SM | 3 | ai | 195 | 131 | 64 | 0.328 |
| Xi'an-speaker dataset | SM | 3 | au | 189 | 132 | 57 | 0.302 |
| Xi'an-speaker dataset | XA | 1 | au | 146 | 103 | 43 | 0.295 |
| Xi'an-speaker dataset | XA | 1 | ai | 81 | 59 | 22 | 0.272 |
| Tianjin-speaker dataset | TJ | 3 | ai | 199 | 155 | 44 | 0.221 |
| Tianjin-speaker dataset | TJ | 1 | ai | 118 | 93 | 25 | 0.212 |
Because functional exclusion is triggered by insufficient observable f0 support for interpolation, Table S1c also serves as the operational summary of f0 trackability loss by condition. The strongest exclusion concentrations occur in low-tone cells, especially SM T3 in both datasets.
f0 missingness modelsTo complement the token-level exclusion audit above, we also modelled pointwise f0 missingness directly with binomial GAMMs in the Tianjin-speaker and Xi'an-speaker datasets. These models estimate the population-level probability that f0 is missing at each normalized time point for each ToneVarDiphInt condition. Together with Table S1c, they provide a more direct picture of f0 trackability loss.
Figure S1-1. Predicted probability of missing f0 in the Tianjin-speaker dataset
Population-level missingness curves from the Tianjin-speaker GAMM, with the strongest probabilities concentrated in low-tone cases (SM T3; TJ T1 and T3).
Table S1e-1. Tianjin-speaker missing-f0 summary by condition
| ToneVarDiphInt | mean_p_missing | max_p_missing | time_of_max |
|---|---|---|---|
| 3 SM aj | 0.272 | 0.492 | 10 |
| 3 SM aw | 0.269 | 0.430 | 10 |
| 3 TJ aj | 0.246 | 0.449 | 0 |
| 3 TJ aw | 0.219 | 0.415 | 0 |
| 1 TJ aj | 0.161 | 0.449 | 10 |
| 1 TJ aw | 0.132 | 0.489 | 10 |
| 4 SM aj | 0.095 | 0.428 | 10 |
| 4 SM aw | 0.084 | 0.351 | 10 |
| 4 TJ aj | 0.066 | 0.337 | 10 |
| 4 TJ aw | 0.055 | 0.244 | 10 |
| 2 TJ aj | 0.031 | 0.162 | 0 |
| 2 SM aj | 0.024 | 0.154 | 0 |
| 2 SM aw | 0.022 | 0.112 | 0 |
| 2 TJ aw | 0.018 | 0.105 | 10 |
| 1 SM aj | 0.016 | 0.122 | 0 |
| 1 SM aw | 0.010 | 0.058 | 0 |
Table S1e-2. Tianjin-speaker missing-f0 summary by condition and time region
The time regions were defined over the 11-point normalized trajectory as early (points 0–3), middle (4–7), and late (8–10), and the reported values are mean predicted missing-f0 probabilities within each region.
| ToneVarDiphInt | time_region | mean_p_missing |
|---|---|---|
| 3 TJ aj | early | 0.383 |
| 3 TJ aw | early | 0.348 |
| 1 TJ aj | late | 0.328 |
| 3 SM aw | early | 0.311 |
| 1 TJ aw | late | 0.303 |
| 3 SM aj | early | 0.286 |
| 3 SM aj | late | 0.285 |
| 4 SM aj | late | 0.251 |
| 3 SM aj | middle | 0.250 |
| 3 SM aw | late | 0.248 |
| 3 SM aw | middle | 0.242 |
| 4 SM aw | late | 0.227 |
| 3 TJ aj | middle | 0.219 |
| 4 TJ aj | late | 0.197 |
| 3 TJ aw | middle | 0.184 |
| 4 TJ aw | late | 0.146 |
| 1 TJ aj | middle | 0.112 |
| 3 TJ aj | late | 0.097 |
| 3 TJ aw | late | 0.093 |
| 1 TJ aj | early | 0.086 |
| 1 TJ aw | middle | 0.083 |
| 2 TJ aj | late | 0.059 |
| 1 TJ aw | early | 0.053 |
| 4 SM aj | middle | 0.044 |
| 2 SM aj | early | 0.044 |
| 2 TJ aw | late | 0.043 |
| 4 SM aw | middle | 0.042 |
| 2 TJ aj | early | 0.041 |
| 2 SM aw | early | 0.033 |
| 2 SM aw | late | 0.031 |
| 1 SM aj | early | 0.031 |
| 4 SM aj | early | 0.027 |
| 4 TJ aw | early | 0.021 |
| 2 SM aj | late | 0.020 |
| 4 TJ aw | middle | 0.020 |
| 4 SM aw | early | 0.020 |
| 1 SM aj | late | 0.019 |
| 4 TJ aj | early | 0.017 |
| 1 SM aw | late | 0.017 |
| 2 TJ aw | early | 0.016 |
| 4 TJ aj | middle | 0.016 |
| 1 SM aw | early | 0.015 |
| 2 SM aj | middle | 0.006 |
| 2 SM aw | middle | 0.004 |
| 2 TJ aw | middle | 0.002 |
| 1 SM aw | middle | 0 |
| 2 TJ aj | middle | 0 |
| 1 SM aj | middle | 0 |
Figure S1-2. Predicted probability of missing f0 in the Xi'an-speaker dataset
Population-level missingness curves from the Xi'an-speaker GAMM. The most severe missingness again concentrates in low-tone cases (SM T3 and XA T1).
Table S1e-3. Xi'an-speaker missing-f0 summary by condition
| ToneVarDiphInt | mean_p_missing_XA | max_p_missing_XA | time_of_max_XA |
|---|---|---|---|
| 3 SM aj | 0.341 | 0.741 | 10 |
| 3 SM aw | 0.276 | 0.666 | 10 |
| 1 XA aj | 0.219 | 0.563 | 0 |
| 1 XA aw | 0.205 | 0.479 | 10 |
| 4 SM aw | 0.112 | 0.457 | 10 |
| 4 SM aj | 0.106 | 0.434 | 10 |
| 3 XA aw | 0.102 | 0.350 | 10 |
| 3 XA aj | 0.096 | 0.328 | 10 |
| 2 SM aj | 0.036 | 0.188 | 0 |
| 2 SM aw | 0.029 | 0.137 | 10 |
| 2 XA aw | 0.028 | 0.088 | 10 |
| 2 XA aj | 0.028 | 0.099 | 0 |
| 1 SM aj | 0.023 | 0.160 | 0 |
| 1 SM aw | 0.020 | 0.106 | 0 |
| 4 XA aw | 0.018 | 0.102 | 10 |
| 4 XA aj | 0.015 | 0.082 | 0 |
Table S1e-4. Xi'an-speaker missing-f0 summary by condition and time region
| ToneVarDiphInt | time_region_XA | mean_p_missing_XA |
|---|---|---|
| 3 SM aj | late | 0.582 |
| 3 SM aw | late | 0.495 |
| 1 XA aw | late | 0.348 |
| 1 XA aj | late | 0.345 |
| 3 SM aj | middle | 0.319 |
| 4 SM aw | late | 0.278 |
| 4 SM aj | late | 0.255 |
| 3 SM aw | middle | 0.239 |
| 3 XA aw | late | 0.218 |
| 3 XA aj | late | 0.218 |
| 1 XA aj | early | 0.212 |
| 3 SM aj | early | 0.183 |
| 1 XA aw | early | 0.161 |
| 3 SM aw | early | 0.150 |
| 1 XA aw | middle | 0.142 |
| 1 XA aj | middle | 0.131 |
| 3 XA aw | middle | 0.095 |
| 4 SM aw | middle | 0.077 |
| 4 SM aj | middle | 0.075 |
| 3 XA aj | middle | 0.073 |
| 2 SM aj | early | 0.060 |
| 2 SM aw | late | 0.058 |
| 1 SM aj | early | 0.041 |
| 4 XA aw | late | 0.041 |
| 2 XA aw | late | 0.040 |
| 2 SM aj | late | 0.039 |
| 2 XA aj | early | 0.039 |
| 1 SM aw | late | 0.035 |
| 2 XA aj | late | 0.033 |
| 1 SM aj | late | 0.029 |
| 2 SM aw | early | 0.029 |
| 4 XA aj | late | 0.028 |
| 2 XA aw | early | 0.028 |
| 1 SM aw | early | 0.028 |
| 3 XA aj | early | 0.027 |
| 4 SM aj | early | 0.025 |
| 4 SM aw | early | 0.021 |
| 3 XA aw | early | 0.021 |
| 4 XA aj | early | 0.021 |
| 2 XA aw | middle | 0.020 |
| 4 XA aw | early | 0.019 |
| 2 XA aj | middle | 0.013 |
| 2 SM aj | middle | 0.008 |
| 2 SM aw | middle | 0.008 |
| 1 SM aw | middle | 0.001 |
| 1 SM aj | middle | 0.001 |
| 4 XA aw | middle | 0 |
| 4 XA aj | middle | 0 |
In both datasets, low-tone conditions, especially SM T3, have the highest predicted probability of missing f0, consistent with the token-level exclusion patterns in Table S1c.
Segmental alignment was obtained with the Montreal Forced Aligner and manually corrected. Vowel onset was placed at the emergence of clear formant structure and offset at its disappearance. Boundaries were primarily guided by F2 and cross-checked against F1–F4. Spectrograms were generally inspected with Praat's default 70-dB dynamic range, adjusted to 50 dB when necessary for low-amplitude signals.
Acoustic measurements were extracted in Praat 6.3.02. f0 was estimated with a two-pass autocorrelation procedure: the first pass used a broad 75–600 Hz range, and the first and third quartiles of the resulting track defined an adaptive range for the second pass. The extracted f0 track was smoothed with Praat's built-in 10-Hz bandwidth smoothing function. F1 was extracted with the Burg method followed by Praat's tracking procedure. Five formants were estimated; the maximum formant frequency depended on speaker sex and diphthong as follows.
Table S1f-1. Maximum formant settings
| Speaker sex | /ai/ | /au/ | Number of formants |
|---|---|---|---|
| Female | 5.5 kHz | 4.9 kHz | 5 |
| Male | 5.0 kHz | 4.4 kHz | 5 |
f0 and F1 were sampled at 11 equidistant points across each diphthong. The primary representation used within-speaker z-scores. The non-z reruns used f0 in semitones relative to the speaker's mean and F1 in speaker-centered Bark. Praat returned no valid f0 estimate at some points; GAMMs retained partially observed tokens and omitted observations with missing model-required values. Functional analyses excluded tokens with fewer than five observed f0 points and filled remaining gaps by linear interpolation with edge filling. These trackability-related exclusions are quantified in S1a–S1e; the FoF check without 10-Hz smoothing is reported in S5h.
The baseline GAMMs model tone-conditioned F1 and f0 trajectories over the normalized time axis. The mechanistic GAMMs replace the explicit tone term with time-varying f0, duration, and interaction structure. Categorical predictors (including ToneVarDiphInt and covariates) in these GAMMs were entered as ordered factors using treatment coding, and random smooths were included for speaker and item (character).
For full GAMM summary and gam.check results, please refer to the document "Full GAMM Summaries and gam.check results".
Baseline: F1 ~ s(time) + s(time, by = ToneVarDiphInt) + covariates + random smooths + AR(1)
Mechanistic: F1 ~ s(time) + s(f0) + s(duration) + ti(time, f0, duration) + covariates + random smooths + AR(1)z and speaker-centered non-z scaling| dataset | scale | family | observations | adjusted_R2 | deviance_explained | n_parametric_terms | n_smooth_terms |
|---|---|---|---|---|---|---|---|
| Tianjin | Z | baseline_F1_tone | 29,513 | 0.806 | 0.820 | 22 | 22 |
| Tianjin | Z | baseline_f0_tone | 23,230 | 0.795 | 0.819 | 22 | 22 |
| Tianjin | Z | mechanistic_F1_f0_duration | 23,230 | 0.832 | 0.841 | 10 | 24 |
| Xi'an | Z | baseline_F1_tone | 28,501 | 0.766 | 0.782 | 22 | 22 |
| Xi'an | Z | baseline_f0_tone | 23,087 | 0.821 | 0.838 | 22 | 22 |
| Xi'an | Z | mechanistic_F1_f0_duration | 23,087 | 0.767 | 0.780 | 10 | 24 |
| Tianjin | notZ | baseline_F1_tone_notZ | 29,513 | 0.817 | 0.830 | 22 | 22 |
| Tianjin | notZ | baseline_f0_tone_notZ | 23,230 | 0.812 | 0.835 | 22 | 22 |
| Tianjin | notZ | mechanistic_F1_f0_duration_notZ | 23,230 | 0.836 | 0.845 | 10 | 24 |
| Xi'an | notZ | baseline_F1_tone_notZ | 28,501 | 0.774 | 0.790 | 22 | 22 |
| Xi'an | notZ | baseline_f0_tone_notZ | 23,087 | 0.835 | 0.852 | 22 | 22 |
| Xi'an | notZ | mechanistic_F1_f0_duration_notZ | 23,087 | 0.778 | 0.791 | 10 | 24 |
Figure S2-1. Baseline F1 curves in the Tianjin-speaker dataset
Baseline F1 trajectory panels from the z-scale analysis. Solid lines show the predicted trajectories, and shaded ribbons show the corresponding confidence intervals.
Figure S2-2. Baseline F1 curves in the Xi'an-speaker dataset
Baseline F1 trajectory panels from the z-scale analysis. Solid lines show the predicted trajectories, and shaded ribbons show the corresponding confidence intervals.
Figure S2-3. Baseline f0 curves in the Tianjin-speaker dataset
Baseline f0 curves by tone for the Tianjin-speaker dataset. Solid lines show the predicted trajectories, and shaded ribbons show the corresponding confidence intervals.
Figure S2-4. Baseline f0 curves in the Xi'an-speaker dataset
Baseline f0 curves by tone for the Xi'an-speaker dataset. Solid lines show the predicted trajectories, and shaded ribbons show the corresponding confidence intervals.
Figure S2-5a. Pairwise F1 difference smooths in the Tianjin-speaker dataset
Combined pairwise F1 difference smooths across sex, variety, diphthong, and the six tone contrasts in the Tianjin-speaker dataset. Shaded regions indicate intervals of significant difference.
Figure S2-5b. Pairwise F1 difference smooths in the Xi'an-speaker dataset
Combined pairwise F1 difference smooths across sex, variety, diphthong, and the six tone contrasts in the Xi'an-speaker dataset. Shaded regions indicate intervals of significant difference.
compareMLThe two models in Table S2b differ only in whether tone is allowed to structure the remaining time-varying F1 trajectory after the shared f0- and duration-based terms have been specified. In the no-tone model, diphthong-by-variety (DiphVar.ord) structures the factor-specific time smooth, whereas in the with-tone model this role is taken by the fuller tone-by-variety-by-diphthong factor (ToneVarDiphInt.ord). The f0- and duration-related smooth terms themselves were intentionally kept the same across the two models, because the goal of this comparison was to ask whether tone still explains residual F1 trajectory shape after controlling for shared nonlinear f0 and duration effects, rather than to build a tone-specific mechanistic mapping for those smooth terms.
| Model | Group | Score | Edf | Difference | Df | p.value | Sig. | comparison | AIC difference (with tone - without tone) |
|---|---|---|---|---|---|---|---|---|---|
| compare_A_XA | Xi'an | 20,684 | 52 | without_tone | |||||
| compare_B_XA | Xi'an | 19,772 | 88 | 911.6 | 36.000 | < 2e-16 | *** | with_tone | -1709.80 |
| compare_A_TJ | Tianjin | 16201 | 52 | without_tone | |||||
| compare_B_TJ | Tianjin | 15830 | 88 | 370.8 | 36.000 | < 2e-16 | *** | with_tone | -743.34 |
Figure S2-6a. Mechanistic GAMM heatmap for /au/ in the Tianjin-speaker dataset
Mechanistic /au/ surface from the z-scale analysis.
Figure S2-6b. Mechanistic GAMM heatmap for /au/ in the Xi'an-speaker dataset
Mechanistic /au/ surface from the z-scale analysis.
Figure S2-6c. Mechanistic GAMM heatmap for /ai/ in the Tianjin-speaker dataset
Mechanistic /ai/ surface from the z-scale analysis.
Figure S2-6d. Mechanistic GAMM heatmap for /ai/ in the Xi'an-speaker dataset
Mechanistic /ai/ surface from the z-scale analysis.
| dataset | scale | baseline_F1_R2 | baseline_f0_R2 | mechanistic_F1_R2 | cascade_mean_RMSE | cascade_mean_nonoverlap_width |
|---|---|---|---|---|---|---|
| Xi'an | Z | 0.766 | 0.821 | 0.767 | 0.242 | 0.799 |
| Tianjin | Z | 0.806 | 0.795 | 0.832 | 0.135 | 0.194 |
| Xi'an | notZ | 0.774 | 0.835 | 0.778 | 0.215 | 0.801 |
| Tianjin | notZ | 0.817 | 0.812 | 0.836 | 0.143 | 0.092 |
Tianjin shows lower mean RMSE and narrower mean non-overlap than Xi'an.
Note: In the GAMM cascade summaries, predicted f0 curves were passed through the mechanistic F1 model together with duration values matched at the full condition-cell level (variety × sex × diphthong × tone × left segment × carrier-type × reading-order); error was then averaged equally across the valid matched cells within each condition.
| Variety | Tone | Mean RMSE (z units) |
|---|---|---|
| Tianjin SM | T1 | 0.126 |
| Tianjin SM | T2 | 0.095 |
| Tianjin SM | T3 | 0.225 |
| Tianjin SM | T4 | 0.121 |
| TJ | T1 | 0.188 |
| TJ | T2 | 0.086 |
| TJ | T3 | 0.144 |
| TJ | T4 | 0.098 |
| Xi'an SM | T1 | 0.272 |
| Xi'an SM | T2 | 0.134 |
| Xi'an SM | T3 | 0.386 |
| Xi'an SM | T4 | 0.120 |
| XA | T1 | 0.428 |
| XA | T2 | 0.281 |
| XA | T3 | 0.174 |
| XA | T4 | 0.143 |
RMSEs (z units) were averaged equally over realized onset, carrier-sentence, and reading-order combinations within each sex-by-diphthong condition, then equally across the two sexes and two diphthongs. Each value is a mean of condition-specific RMSEs, not an RMSE calculated from averaged trajectories.
Figure S2-7. Direct-versus-cascade trajectories for /ai/ in the Tianjin-speaker dataset (z scale)
Representative trajectory-level comparison between the direct F1 fit and the cascade prediction for /ai/. These panels are shown across sex and variety, with carrier type fixed at 1, left segment fixed at m, and reading order set to the panel-specific reference level (S for SM, T for TJ, X for XA). Colored lines distinguish the direct F1 fit from the f0-based cascade, and translucent ribbons in matching colors show their confidence intervals.
Figure S2-8. Direct-versus-cascade trajectories for /ai/ in the Xi'an-speaker dataset (z scale)
Representative trajectory-level comparison between the direct F1 fit and the cascade prediction for /ai/. These panels are shown across sex and variety, with carrier type fixed at 1, left segment fixed at m, and reading order set to the panel-specific reference level (S for SM, T for TJ, X for XA). Colored lines distinguish the direct F1 fit from the f0-based cascade, and translucent ribbons in matching colors show their confidence intervals.
Figure S2-9. Direct-versus-cascade trajectories for /au/ in the Tianjin-speaker dataset (z scale)
Representative trajectory-level comparison between the direct F1 fit and the cascade prediction for /au/. These panels are shown across sex and variety, with carrier type fixed at 1, left segment fixed at m, and reading order set to the panel-specific reference level (S for SM, T for TJ, X for XA). Colored lines distinguish the direct F1 fit from the f0-based cascade, and translucent ribbons in matching colors show their confidence intervals.
Figure S2-10. Direct-versus-cascade trajectories for /au/ in the Xi'an-speaker dataset (z scale)
Representative trajectory-level comparison between the direct F1 fit and the cascade prediction for /au/. These panels are shown across sex and variety, with carrier type fixed at 1, left segment fixed at m, and reading order set to the panel-specific reference level (S for SM, T for TJ, X for XA). Colored lines distinguish the direct F1 fit from the f0-based cascade, and translucent ribbons in matching colors show their confidence intervals.
Figure S2-11. Direct-versus-cascade divergence for /ai/ in the Tianjin-speaker dataset (z scale)
Representative band non-overlap diagnostic for /ai/ under the z-scale cascade comparison, using the same fixed reference setting as Figures S2-7 and S2-8. Shaded regions indicate non-overlap intervals.
Figure S2-12. Direct-versus-cascade divergence for /ai/ in the Xi'an-speaker dataset (z scale)
Representative band non-overlap diagnostic for /ai/ under the z-scale cascade comparison, using the same fixed reference setting as Figures S2-7 and S2-8. Shaded regions indicate non-overlap intervals.
Figure S2-13. Direct-versus-cascade divergence for /au/ in the Tianjin-speaker dataset (z scale)
Representative band non-overlap diagnostic for /au/ under the z-scale cascade comparison, using the same fixed reference setting as Figures S2-9 and S2-10. Shaded regions indicate non-overlap intervals.
Figure S2-14. Direct-versus-cascade divergence for /au/ in the Xi'an-speaker dataset (z scale)
Representative band non-overlap diagnostic for /au/ under the z-scale cascade comparison, using the same fixed reference setting as Figures S2-9 and S2-10. Shaded regions indicate non-overlap intervals.
Residual-tone diagnostics quantify the structured variance that remains after fitting the no-tone f0-duration backbone.
| dataset | mean_rmse | mean_mae | mean_abs_bias | tone_p_rmse | tone_variety_p_rmse | tone_p_mae | tone_variety_p_mae | hardest_condition | hardest_rmse | easiest_condition | easiest_rmse |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Xi'an | 0.407 | 0.340 | 0.233 | < 0.001 | 0.014 | < 0.001 | 0.013 | SM M aj 2 | 0.557 | SM M aw 3 | 0.347 |
| Tianjin | 0.349 | 0.290 | 0.186 | 0.232 | 0.001 | 0.389 | 0.034 | SM F aw 1 | 0.415 | TJ F aj 2 | 0.293 |
| dataset | metric | tone_p | tone_variety_p |
|---|---|---|---|
| Xi'an | rmse_resid | < 0.001 | 0.014 |
| Xi'an | mae_resid | < 0.001 | 0.013 |
| Xi'an | bias_resid | 0.642 | < 0.001 |
| Xi'an | maxabs_resid | < 0.001 | < 0.001 |
| Tianjin | rmse_resid | 0.232 | 0.001 |
| Tianjin | mae_resid | 0.329 | 0.034 |
| Tianjin | bias_resid | 0.620 | < 0.001 |
| Tianjin | maxabs_resid | < 0.001 | < 0.001 |
Residual F1 variation shows tone- and variety-dependent structure, with significant tone-by-variety interactions in both datasets.
Figure S2-15. Residual curves by tone in the Xi'an-speaker dataset
Residual curves after the no-tone mechanistic GAMM.
Figure S2-16. Residual curves by tone in the Tianjin-speaker dataset
Residual curves after the no-tone mechanistic GAMM.
Figure S2-17. Residual RMSE heatmap in the Xi'an-speaker dataset
Residual RMSE by condition after the no-tone mechanistic GAMM.
Figure S2-18. Residual RMSE heatmap in the Tianjin-speaker dataset
Residual RMSE by condition after the no-tone mechanistic GAMM.
Trajectories from both speaker origins were pooled to estimate one f0 basis and one F1 basis for the primary /ai/ + /au/ analysis. Speaker origin (Tianjin or Xi'an) and variety status (SM or local) define four cells: Tianjin SM, Tianjin TJ, Xi'an SM, and Xi'an XA. Local status therefore denotes TJ for Tianjin speakers and XA for Xi'an speakers. The diphthong-specific reruns each re-estimate common bases across the same four cells. Origin-specific summaries below use these common bases; the separate seven-group decompositions are reported in S6.
Table S3a-1. Analysis blocks and retained sample sizes
| analysis_id | n_tokens | n_speakers | f0_pcs_retained | f1_pcs_retained |
|---|---|---|---|---|
| pooled_2x2_ai | 2,174 | 48 | 4 | 9 |
| pooled_2x2_au | 2,586 | 48 | 4 | 10 |
| pooled_2x2_ai_au_pooled | 4,760 | 48 | 4 | 10 |
Table S3a-2. Explained variance proportions
| Analysis | Domain | PC1 | PC2 | PC3 | PC1–PC2 cumulative | PC1–PC3 cumulative |
|---|---|---|---|---|---|---|
| ai_au_pooled | f0 | 0.685 | 0.276 | 0.028 | 0.960 | 0.989 |
| ai_au_pooled | f1 | 0.527 | 0.247 | 0.084 | 0.774 | 0.858 |
| ai | f0 | 0.696 | 0.264 | 0.029 | 0.961 | 0.990 |
| ai | f1 | 0.525 | 0.278 | 0.079 | 0.803 | 0.882 |
| au | f0 | 0.678 | 0.282 | 0.028 | 0.960 | 0.988 |
| au | f1 | 0.464 | 0.245 | 0.104 | 0.709 | 0.814 |
In the primary pooled analysis, the first two components account for 96.05% of f0 variance and 77.44% of F1 variance. PC1 primarily captures overall level and PC2 dynamic tilt; PC3 captures finer trajectory shape. Retention counts in Table S3a-1 refer to the decomposition, while the score and bridge models use the first three predictor PCs and the relative-configuration reconstruction uses PC1–PC2.
Figure S3-1. Pooled eigenfunctions (z scale)
Domain-specific eigenfunctions estimated across both origins and both variety statuses.
Figure S3-2. Pooled trajectory reconstructions (z scale)
Tone-wise trajectories reconstructed from mean PC scores in the four origin × variety cells. These are descriptive, unadjusted reconstructions; covariate-adjusted time shapes are shown in Figure S3-3.
The primary mixed models predict F1-PC1 or F1-PC2 from f0-PC1–PC3, standardized duration, tone × origin × variety status, diphthong, sex, reading order, onset, and carrier type, with a speaker random intercept. The duration R² comparison is a companion ordinary least-squares calculation with the same fixed predictors; its R² values are not mixed-model marginal or conditional R². A likelihood-ratio comparison was also conducted between each full mixed model and a reduced mixed model containing f0-PC1–PC3, duration, the non-tonal covariates, and the speaker random intercept, but omitting the entire tone × origin × variety-status block. Both models were fitted by maximum likelihood. Because the reduced model omits origin and variety status as well as tone and their interactions, this comparison tests the joint contribution of the complete 15-parameter condition block rather than a tone-only effect.
Table S3b-1. Incremental contribution of duration
| Analysis | Outcome | R² without duration | R² with duration | ΔR² |
|---|---|---|---|---|
| ai_au_pooled | F1_PC1 | 0.469 | 0.469 | 3.31e-05 |
| ai_au_pooled | F1_PC2 | 0.267 | 0.332 | 0.065 |
Table S3b-2. f0-PC and duration coefficients in the primary mixed models
| Outcome | Term | Estimate | SE | t | p |
|---|---|---|---|---|---|
| F1_PC1 | F0_PC1 | 0.095 | 0.012 | 8.123 | < 0.001 |
| F1_PC1 | F0_PC2 | 0.003 | 0.018 | 0.145 | 0.885 |
| F1_PC1 | F0_PC3 | 0.026 | 0.035 | 0.736 | 0.462 |
| F1_PC1 | duration_z | 0.010 | 0.018 | 0.566 | 0.571 |
| F1_PC2 | F0_PC1 | 0.000258 | 0.009 | 0.030 | 0.976 |
| F1_PC2 | F0_PC2 | 0.043 | 0.013 | 3.298 | < 0.001 |
| F1_PC2 | F0_PC3 | 0.083 | 0.026 | 3.196 | 0.001 |
| F1_PC2 | duration_z | -0.306 | 0.013 | -23.230 | < 0.001 |
Duration contributes chiefly to F1-PC2 (ΔR² = 0.06498), with negligible change for F1-PC1 (ΔR² = 0.0000331). The conditional F1-PC1 coefficient for f0-PC1 is positive in the mixed model; this conditional coefficient should be distinguished from the negative unadjusted score correlations and the inverse relative tone configurations.
Table S3b-3. Unadjusted cross-domain score correlations in the common bases
| Dataset | PC1 score r | PC2 score r | PC3 score r |
|---|---|---|---|
| pooled | -0.255 | -0.133 | 0.074 |
| Tianjin | -0.072 | -0.042 | 0.048 |
| Xi'an | -0.404 | -0.174 | 0.112 |
Table S3b-4. Likelihood-ratio comparison of the full and reduced mixed models
| Outcome | Reduced parameters | Full parameters | Reduced AIC | Full AIC | χ² | df | p |
|---|---|---|---|---|---|---|---|
| F1-PC1 | 15 | 30 | 16119.12 | 14688.46 | 1460.663 | 15 | < 0.001 |
| F1-PC2 | 15 | 30 | 11986.49 | 11692.92 | 323.562 | 15 | < 0.001 |
For both F1-PC1 and F1-PC2, the full model fitted substantially better than the corresponding reduced model. Thus, tone, origin, variety status, and their interactions jointly accounted for F1 score variation beyond the low-dimensional f0 scores, duration, and the non-tonal covariates.
Figure S3-3. Covariate-adjusted pooled PC time shapes (z scale)
Tone-wise PC1 and PC2 contributions reconstructed from mixed-model estimated marginal mean scores across the four origin × variety cells.
Within the common pooled bases, separate linear bridges are fitted for each origin: each F1-PC1–PC3 score is predicted from f0-PC1–PC3. R² describes token-level in-sample fit. Absolute and Euclidean discrepancies are computed between direct and predicted tone × variety mean scores and averaged equally across the eight conditions within each origin; they are not token-level prediction errors.
Table S3c-1. Origin-specific bridge fit and condition-mean discrepancy
| Origin | Tokens | Bridge R² PC1 | Bridge R² PC2 | Bridge R² PC3 | Mean absolute gap PC1 | Mean absolute gap PC2 | Mean absolute gap PC3 | Mean distance PC1–PC2 | Mean distance PC1–PC3 |
|---|---|---|---|---|---|---|---|---|---|
| Tianjin | 2,403 | 0.015 | 0.007 | 0.006 | 0.243 | 0.116 | 0.021 | 0.297 | 0.298 |
| Xi'an | 2,357 | 0.174 | 0.034 | 0.013 | 0.444 | 0.261 | 0.052 | 0.559 | 0.567 |
Xi'an has higher bridge R² for PC1 and PC2 (0.174 and 0.034) than Tianjin (0.015 and 0.007), but a larger PC1–PC2 condition-mean discrepancy (0.559 versus 0.297). Predictability and discrepancy therefore describe different aspects of recoverability.
Separate mixed models for each domain's PC1 and PC2 estimate tone × origin × variety marginal means, adjusting for duration and the non-tonal covariates with a speaker random intercept. Reconstructed trajectories are centered across the four tones at each time point within each origin × variety cell and domain. Spearman correlation measures agreement in tone rank; Pearson correlation measures agreement in centered spacing. The summary is the mean of the 11 pointwise correlations. Negative correlations indicate inverse configurations.
Table S3d-1. Adjusted relative-configuration summary
| Origin × variety cell | Mean rank correlation | Mean level correlation |
|---|---|---|
| Tianjin SM | -0.727 | -0.693 |
| Tianjin local | -0.655 | -0.645 |
| Xi'an SM | -0.909 | -0.930 |
| Xi'an local | -0.855 | -0.910 |
Both Xi'an cells show stronger inverse alignment than the corresponding Tianjin cells.
Figure S3-4. Adjusted relative f0–F1 configuration (z scale)
Centered trajectories reconstructed from adjusted PC1–PC2 scores; each row identifies one origin × variety cell.
Figure S3-5. Adjusted relative similarity over time (z; /ai/ + /au/)
Correlations across the four tones at each of the 11 time points, calculated from covariate-adjusted PC1–PC2 reconstructions centered within origin × variety cell and domain. Negative values indicate inverse ordering or spacing.
For this supplementary paired-score evaluation, one briged is fitted across all four cells, then compares direct and bridged token scores using method × ToneOriginVar, onset, reading order, carrier type, and diphthong, with random intercepts for speaker and token. This pooled bridge differs from the origin-specific bridges summarized in S3c. Estimated marginal means and within-condition contrasts describe the pooled bridge's remaining condition-structured discrepancies. Each condition has one direct-minus-bridged contrast; the reported p values have no adjustment across the complete collection of conditions and PCs.
Table S3e-1. Direct-minus-bridged contrasts for all three F1 PCs
| F1 component | Condition | Direct − bridged | SE | z | p |
|---|---|---|---|---|---|
| F1_PC1 | Tianjin 1 SM | -0.135 | 0.071 | -1.900 | 0.057 |
| F1_PC1 | Tianjin 1 local | -0.150 | 0.081 | -1.843 | 0.065 |
| F1_PC1 | Tianjin 2 SM | -0.240 | 0.081 | -2.978 | 0.003 |
| F1_PC1 | Tianjin 2 local | -0.244 | 0.079 | -3.098 | 0.002 |
| F1_PC1 | Tianjin 3 SM | 0.059 | 0.079 | 0.743 | 0.457 |
| F1_PC1 | Tianjin 3 local | 0.122 | 0.072 | 1.687 | 0.092 |
| F1_PC1 | Tianjin 4 SM | 0.501 | 0.064 | 7.832 | < 0.001 |
| F1_PC1 | Tianjin 4 local | 0.370 | 0.063 | 5.864 | < 0.001 |
| F1_PC1 | Xi'an 1 SM | -0.234 | 0.071 | -3.276 | 0.001 |
| F1_PC1 | Xi'an 1 local | 0.161 | 0.099 | 1.623 | 0.105 |
| F1_PC1 | Xi'an 2 SM | 0.509 | 0.082 | 6.218 | < 0.001 |
| F1_PC1 | Xi'an 2 local | 0.160 | 0.083 | 1.939 | 0.052 |
| F1_PC1 | Xi'an 3 SM | 1.305 | 0.079 | 16.427 | < 0.001 |
| F1_PC1 | Xi'an 3 local | -0.925 | 0.065 | -14.266 | < 0.001 |
| F1_PC1 | Xi'an 4 SM | 0.373 | 0.064 | 5.850 | < 0.001 |
| F1_PC1 | Xi'an 4 local | -1.095 | 0.064 | -17.224 | < 0.001 |
| F1_PC2 | Tianjin 1 SM | -0.176 | 0.052 | -3.373 | < 0.001 |
| F1_PC2 | Tianjin 1 local | -0.004 | 0.060 | -0.061 | 0.951 |
| F1_PC2 | Tianjin 2 SM | -0.185 | 0.060 | -3.104 | 0.002 |
| F1_PC2 | Tianjin 2 local | -0.046 | 0.058 | -0.800 | 0.424 |
| F1_PC2 | Tianjin 3 SM | -0.334 | 0.059 | -5.703 | < 0.001 |
| F1_PC2 | Tianjin 3 local | -0.516 | 0.053 | -9.712 | < 0.001 |
| F1_PC2 | Tianjin 4 SM | -0.234 | 0.047 | -4.962 | < 0.001 |
| F1_PC2 | Tianjin 4 local | -0.249 | 0.047 | -5.359 | < 0.001 |
| F1_PC2 | Xi'an 1 SM | 0.051 | 0.053 | 0.964 | 0.335 |
| F1_PC2 | Xi'an 1 local | 0.614 | 0.073 | 8.410 | < 0.001 |
| F1_PC2 | Xi'an 2 SM | -0.400 | 0.060 | -6.621 | < 0.001 |
| F1_PC2 | Xi'an 2 local | 0.034 | 0.061 | 0.558 | 0.577 |
| F1_PC2 | Xi'an 3 SM | 0.617 | 0.059 | 10.539 | < 0.001 |
| F1_PC2 | Xi'an 3 local | 0.377 | 0.048 | 7.880 | < 0.001 |
| F1_PC2 | Xi'an 4 SM | 0.070 | 0.047 | 1.489 | 0.136 |
| F1_PC2 | Xi'an 4 local | 0.490 | 0.047 | 10.437 | < 0.001 |
| F1_PC3 | Tianjin 1 SM | 0.072 | 0.033 | 2.205 | 0.027 |
| F1_PC3 | Tianjin 1 local | 0.021 | 0.037 | 0.554 | 0.580 |
| F1_PC3 | Tianjin 2 SM | 0.048 | 0.037 | 1.301 | 0.193 |
| F1_PC3 | Tianjin 2 local | 0.045 | 0.036 | 1.242 | 0.214 |
| F1_PC3 | Tianjin 3 SM | -0.045 | 0.037 | -1.226 | 0.220 |
| F1_PC3 | Tianjin 3 local | -0.017 | 0.033 | -0.515 | 0.607 |
| F1_PC3 | Tianjin 4 SM | 0.009 | 0.029 | 0.295 | 0.768 |
| F1_PC3 | Tianjin 4 local | 0.027 | 0.029 | 0.942 | 0.346 |
| F1_PC3 | Xi'an 1 SM | -0.010 | 0.033 | -0.301 | 0.764 |
| F1_PC3 | Xi'an 1 local | -0.194 | 0.046 | -4.267 | < 0.001 |
| F1_PC3 | Xi'an 2 SM | -0.022 | 0.038 | -0.584 | 0.559 |
| F1_PC3 | Xi'an 2 local | -0.031 | 0.038 | -0.829 | 0.407 |
| F1_PC3 | Xi'an 3 SM | 0.029 | 0.037 | 0.804 | 0.421 |
| F1_PC3 | Xi'an 3 local | -0.003 | 0.030 | -0.098 | 0.922 |
| F1_PC3 | Xi'an 4 SM | 0.079 | 0.029 | 2.712 | 0.007 |
| F1_PC3 | Xi'an 4 local | -0.099 | 0.029 | -3.402 | < 0.001 |
Figure S3-6. Pooled bridge marginal means (z scale)
Covariate-adjusted direct and bridged F1-PC1–PC3 scores with 95% confidence intervals. Bridged scores come from the all-origin pooled bridge described in S3e.
Each diphthong is analyzed in its own common f0 and F1 spaces across both origins and variety statuses. The models follow S3b–S3e, omitting diphthong as a covariate. Origin-specific bridge summaries and adjusted relative configurations are reported separately.
Table S3f-1. Incremental contribution of duration
| Analysis | Outcome | R² without duration | R² with duration | ΔR² |
|---|---|---|---|---|
| ai | F1_PC1 | 0.425 | 0.426 | 0.000435 |
| ai | F1_PC2 | 0.327 | 0.405 | 0.078 |
| au | F1_PC1 | 0.383 | 0.383 | 2e-05 |
| au | F1_PC2 | 0.260 | 0.317 | 0.057 |
Table S3f-2. /ai/ origin-specific bridge summary
| Origin | Tokens | Bridge R² PC1 | Bridge R² PC2 | Bridge R² PC3 | Mean absolute gap PC1 | Mean absolute gap PC2 | Mean absolute gap PC3 | Mean distance PC1–PC2 | Mean distance PC1–PC3 |
|---|---|---|---|---|---|---|---|---|---|
| Tianjin | 1,122 | 0.014 | 0.019 | 0.010 | 0.276 | 0.112 | 0.045 | 0.319 | 0.323 |
| Xi'an | 1,052 | 0.240 | 0.056 | 0.041 | 0.500 | 0.398 | 0.056 | 0.694 | 0.700 |
Table S3f-3. /ai/ adjusted relative configuration
| Origin × variety cell | Mean rank correlation | Mean level correlation |
|---|---|---|
| Tianjin SM | -0.727 | -0.753 |
| Tianjin local | -0.727 | -0.741 |
| Xi'an SM | -0.909 | -0.928 |
| Xi'an local | -0.855 | -0.890 |
Figure S3-7. /ai/ pooled bridge marginal means (z scale)
Direct and bridged estimated marginal means with 95% confidence intervals, using the all-origin bridge within this diphthong.
Figure S3-8. /ai/ adjusted relative configuration (z scale)
Adjusted and centered PC1–PC2 reconstructions across the four cells.
Figure S3-9. Adjusted relative similarity over time (z; ai)
Correlations across the four tones at each of the 11 time points, calculated from covariate-adjusted PC1–PC2 reconstructions centered within origin × variety cell and domain. Negative values indicate inverse ordering or spacing.
Table S3f-4. /au/ origin-specific bridge summary
| Origin | Tokens | Bridge R² PC1 | Bridge R² PC2 | Bridge R² PC3 | Mean absolute gap PC1 | Mean absolute gap PC2 | Mean absolute gap PC3 | Mean distance PC1–PC2 | Mean distance PC1–PC3 |
|---|---|---|---|---|---|---|---|---|---|
| Tianjin | 1,281 | 0.024 | 0.004 | 0.005 | 0.144 | 0.121 | 0.031 | 0.213 | 0.218 |
| Xi'an | 1,305 | 0.156 | 0.044 | 0.004 | 0.378 | 0.203 | 0.048 | 0.473 | 0.483 |
Table S3f-5. /au/ adjusted relative configuration
| Origin × variety cell | Mean rank correlation | Mean level correlation |
|---|---|---|
| Tianjin SM | -0.727 | -0.618 |
| Tianjin local | -0.455 | -0.468 |
| Xi'an SM | -0.909 | -0.916 |
| Xi'an local | -0.855 | -0.909 |
Figure S3-10. /au/ pooled bridge marginal means (z scale)
Direct and bridged estimated marginal means with 95% confidence intervals, using the all-origin bridge within this diphthong.
Figure S3-11. /au/ adjusted relative configuration (z scale)
Adjusted and centered PC1–PC2 reconstructions across the four cells.
Figure S3-12. Adjusted relative similarity over time (z; au)
Correlations across the four tones at each of the 11 time points, calculated from covariate-adjusted PC1–PC2 reconstructions centered within origin × variety cell and domain. Negative values indicate inverse ordering or spacing.
The supplementary multivariate f0 + F1 analysis belongs to the separate-group checks and is reported in S6g.
The pooled 2 × 2 fPCA analysis is repeated with f0 in semitones relative to each speaker's mean and F1 in speaker-centered Bark. The primary pooled and diphthong-specific runs retain the same 4,760, 2,174, and 2,586 tokens, respectively. Models, bridge definitions, and adjusted relative-configuration calculations follow S3. Duration remains standardized. The earlier separate-group non-z analyses are retained in S6h–S6i.
Table S4a-1. Incremental contribution of duration
| Analysis | Outcome | R² without duration | R² with duration | ΔR² |
|---|---|---|---|---|
| ai_au_pooled | F1_PC1 | 0.471 | 0.471 | 4.23e-06 |
| ai_au_pooled | F1_PC2 | 0.364 | 0.411 | 0.046 |
Table S4a-2. Non-z origin-specific bridge summary
| Origin | Tokens | Bridge R² PC1 | Bridge R² PC2 | Bridge R² PC3 | Mean absolute gap PC1 | Mean absolute gap PC2 | Mean absolute gap PC3 | Mean distance PC1–PC2 | Mean distance PC1–PC3 |
|---|---|---|---|---|---|---|---|---|---|
| Tianjin | 2,403 | 0.016 | 0.007 | 0.004 | 0.236 | 0.119 | 0.021 | 0.293 | 0.294 |
| Xi'an | 2,357 | 0.173 | 0.035 | 0.005 | 0.373 | 0.189 | 0.064 | 0.455 | 0.466 |
Table S4a-3. Non-z adjusted relative configuration
| Origin × variety cell | Mean rank correlation | Mean level correlation |
|---|---|---|
| Tianjin SM | -0.727 | -0.699 |
| Tianjin local | -0.582 | -0.675 |
| Xi'an SM | -0.909 | -0.943 |
| Xi'an local | -0.855 | -0.913 |
Table S4a-4. Non-z pooled direct-minus-bridged contrasts
| F1 component | Condition | Direct − bridged | SE | z | p |
|---|---|---|---|---|---|
| F1_PC1 | Tianjin 1 SM | -0.107 | 0.065 | -1.644 | 0.100 |
| F1_PC1 | Tianjin 1 local | -0.057 | 0.074 | -0.767 | 0.443 |
| F1_PC1 | Tianjin 2 SM | -0.127 | 0.074 | -1.717 | 0.086 |
| F1_PC1 | Tianjin 2 local | -0.164 | 0.072 | -2.278 | 0.023 |
| F1_PC1 | Tianjin 3 SM | 0.093 | 0.073 | 1.274 | 0.203 |
| F1_PC1 | Tianjin 3 local | 0.174 | 0.066 | 2.638 | 0.008 |
| F1_PC1 | Tianjin 4 SM | 0.575 | 0.059 | 9.813 | < 0.001 |
| F1_PC1 | Tianjin 4 local | 0.404 | 0.058 | 7.003 | < 0.001 |
| F1_PC1 | Xi'an 1 SM | -0.270 | 0.065 | -4.133 | < 0.001 |
| F1_PC1 | Xi'an 1 local | 0.086 | 0.091 | 0.953 | 0.341 |
| F1_PC1 | Xi'an 2 SM | 0.389 | 0.075 | 5.187 | < 0.001 |
| F1_PC1 | Xi'an 2 local | 0.087 | 0.076 | 1.155 | 0.248 |
| F1_PC1 | Xi'an 3 SM | 0.989 | 0.073 | 13.612 | < 0.001 |
| F1_PC1 | Xi'an 3 local | -0.863 | 0.059 | -14.537 | < 0.001 |
| F1_PC1 | Xi'an 4 SM | 0.279 | 0.058 | 4.784 | < 0.001 |
| F1_PC1 | Xi'an 4 local | -1.056 | 0.058 | -18.151 | < 0.001 |
| F1_PC2 | Tianjin 1 SM | -0.387 | 0.057 | -6.848 | < 0.001 |
| F1_PC2 | Tianjin 1 local | -0.169 | 0.065 | -2.599 | 0.009 |
| F1_PC2 | Tianjin 2 SM | -0.354 | 0.064 | -5.494 | < 0.001 |
| F1_PC2 | Tianjin 2 local | -0.216 | 0.063 | -3.438 | < 0.001 |
| F1_PC2 | Tianjin 3 SM | -0.485 | 0.063 | -7.648 | < 0.001 |
| F1_PC2 | Tianjin 3 local | -0.621 | 0.058 | -10.794 | < 0.001 |
| F1_PC2 | Tianjin 4 SM | -0.373 | 0.051 | -7.295 | < 0.001 |
| F1_PC2 | Tianjin 4 local | -0.397 | 0.050 | -7.874 | < 0.001 |
| F1_PC2 | Xi'an 1 SM | 0.223 | 0.057 | 3.901 | < 0.001 |
| F1_PC2 | Xi'an 1 local | 0.724 | 0.079 | 9.162 | < 0.001 |
| F1_PC2 | Xi'an 2 SM | -0.067 | 0.065 | -1.020 | 0.308 |
| F1_PC2 | Xi'an 2 local | 0.284 | 0.066 | 4.308 | < 0.001 |
| F1_PC2 | Xi'an 3 SM | 0.857 | 0.063 | 13.522 | < 0.001 |
| F1_PC2 | Xi'an 3 local | 0.428 | 0.052 | 8.272 | < 0.001 |
| F1_PC2 | Xi'an 4 SM | 0.276 | 0.051 | 5.418 | < 0.001 |
| F1_PC2 | Xi'an 4 local | 0.502 | 0.051 | 9.898 | < 0.001 |
| F1_PC3 | Tianjin 1 SM | 0.044 | 0.032 | 1.371 | 0.170 |
| F1_PC3 | Tianjin 1 local | -0.010 | 0.037 | -0.265 | 0.791 |
| F1_PC3 | Tianjin 2 SM | 0.042 | 0.037 | 1.131 | 0.258 |
| F1_PC3 | Tianjin 2 local | -0.005 | 0.036 | -0.150 | 0.881 |
| F1_PC3 | Tianjin 3 SM | -0.043 | 0.036 | -1.185 | 0.236 |
| F1_PC3 | Tianjin 3 local | -0.009 | 0.033 | -0.289 | 0.773 |
| F1_PC3 | Tianjin 4 SM | -0.029 | 0.029 | -0.994 | 0.320 |
| F1_PC3 | Tianjin 4 local | -0.008 | 0.029 | -0.283 | 0.777 |
| F1_PC3 | Xi'an 1 SM | 0.027 | 0.033 | 0.824 | 0.410 |
| F1_PC3 | Xi'an 1 local | -0.160 | 0.045 | -3.539 | < 0.001 |
| F1_PC3 | Xi'an 2 SM | 0.045 | 0.037 | 1.214 | 0.225 |
| F1_PC3 | Xi'an 2 local | 0.015 | 0.038 | 0.406 | 0.684 |
| F1_PC3 | Xi'an 3 SM | 0.042 | 0.036 | 1.157 | 0.247 |
| F1_PC3 | Xi'an 3 local | -0.005 | 0.030 | -0.166 | 0.868 |
| F1_PC3 | Xi'an 4 SM | 0.099 | 0.029 | 3.402 | < 0.001 |
| F1_PC3 | Xi'an 4 local | -0.092 | 0.029 | -3.192 | 0.001 |
As in S3e, the marginal-mean contrasts evaluate the all-origin pooled bridge, whereas Table S4a-2 summarizes separately fitted origin-specific bridges. Contrast p values are not adjusted across all conditions and PCs. Duration again contributes mainly to PC2 (ΔR² = 0.04640), with negligible change for PC1 (0.00000423). The origin-specific PC1–PC2 discrepancy remains smaller for Tianjin (0.293) than Xi'an (0.455), and the Xi'an relative configurations remain more strongly inverse.
Table S4b-1. Incremental contribution of duration
| Analysis | Outcome | R² without duration | R² with duration | ΔR² |
|---|---|---|---|---|
| ai | F1_PC1 | 0.438 | 0.439 | 0.000225 |
| ai | F1_PC2 | 0.398 | 0.455 | 0.058 |
| au | F1_PC1 | 0.325 | 0.330 | 0.004 |
| au | F1_PC2 | 0.412 | 0.446 | 0.033 |
Table S4b-2. /ai/ non-z origin-specific bridge summary
| Origin | Tokens | Bridge R² PC1 | Bridge R² PC2 | Bridge R² PC3 | Mean absolute gap PC1 | Mean absolute gap PC2 | Mean absolute gap PC3 | Mean distance PC1–PC2 | Mean distance PC1–PC3 |
|---|---|---|---|---|---|---|---|---|---|
| Tianjin | 1,122 | 0.026 | 0.011 | 0.012 | 0.239 | 0.138 | 0.040 | 0.298 | 0.302 |
| Xi'an | 1,052 | 0.241 | 0.033 | 0.029 | 0.433 | 0.288 | 0.074 | 0.565 | 0.574 |
Table S4b-3. /ai/ non-z adjusted relative configuration
| Origin × variety cell | Mean rank correlation | Mean level correlation |
|---|---|---|
| Tianjin SM | -0.764 | -0.782 |
| Tianjin local | -0.727 | -0.789 |
| Xi'an SM | -0.909 | -0.942 |
| Xi'an local | -0.855 | -0.891 |
Table S4b-4. /au/ non-z origin-specific bridge summary
| Origin | Tokens | Bridge R² PC1 | Bridge R² PC2 | Bridge R² PC3 | Mean absolute gap PC1 | Mean absolute gap PC2 | Mean absolute gap PC3 | Mean distance PC1–PC2 | Mean distance PC1–PC3 |
|---|---|---|---|---|---|---|---|---|---|
| Tianjin | 1,281 | 0.028 | 0.005 | 0.001 | 0.130 | 0.142 | 0.036 | 0.214 | 0.221 |
| Xi'an | 1,305 | 0.141 | 0.069 | 0.003 | 0.300 | 0.184 | 0.055 | 0.389 | 0.403 |
Table S4b-5. /au/ non-z adjusted relative configuration
| Origin × variety cell | Mean rank correlation | Mean level correlation |
|---|---|---|
| Tianjin SM | -0.727 | -0.627 |
| Tianjin local | -0.473 | -0.524 |
| Xi'an SM | -0.855 | -0.919 |
| Xi'an local | -0.836 | -0.916 |
z vs not-z)The fPCA rows use the pooled 2 × 2 results in S3c–S3d and S4a. GAMM and FoF entries retain their respective analyses in S2 and S5. Discrepancy and RMSE values have representation-dependent units; compare the cross-origin pattern within each scale.
| dataset | method | metric | Z | notZ |
|---|---|---|---|---|
| Tianjin | GAMM cascade | mean RMSE | 0.135 | 0.143 |
| Tianjin | GAMM cascade | mean nonoverlap width | 0.194 | 0.092 |
| Tianjin | fPCA bridge | mean Euclidean discrepancy PC1-PC2 | 0.297 | 0.293 |
| Tianjin | Relative configuration | mean rank correlation (local panel) | -0.655 | -0.582 |
| Tianjin | FoF | cross-validated R2 | 0.698 | 0.699 |
| Tianjin | FoF | cross-validated RMSE | 0.544 | 0.591 |
| Xi'an | GAMM cascade | mean RMSE | 0.242 | 0.215 |
| Xi'an | GAMM cascade | mean nonoverlap width | 0.799 | 0.801 |
| Xi'an | fPCA bridge | mean Euclidean discrepancy PC1-PC2 | 0.559 | 0.455 |
| Xi'an | Relative configuration | mean rank correlation (local panel) | -0.855 | -0.855 |
| Xi'an | FoF | cross-validated R2 | 0.566 | 0.592 |
| Xi'an | FoF | cross-validated RMSE | 0.646 | 0.568 |
Taken together, the non-z analyses support the same broad interpretation as the main text: the GAMM cascade asymmetry remains, the fPCA bridge asymmetry remains, the stronger Xi'an inverse relative configuration remains, and the FoF direct full-trajectory contrast is partly scale-sensitive across representations.
Figure S4-1A. Non-z cascade reconstruction for /ai/ in the Tianjin-speaker dataset
Representative non-z direct-versus-cascade trajectory comparison for /ai/. These panels are shown across sex and variety, with carrier type fixed at 1, left segment fixed at m, and reading order set to the panel-specific reference level (S for SM, T for TJ, X for XA).
Figure S4-1B. Non-z cascade reconstruction for /au/ in the Tianjin-speaker dataset
Representative non-z direct-versus-cascade trajectory comparison for /au/. These panels are shown across sex and variety, with carrier type fixed at 1, left segment fixed at m, and reading order set to the panel-specific reference level (S for SM, T for TJ, X for XA).
Figure S4-2A. Non-z cascade reconstruction for /ai/ in the Xi'an-speaker dataset
Representative non-z direct-versus-cascade trajectory comparison for /ai/. These panels are shown across sex and variety, with carrier type fixed at 1, left segment fixed at m, and reading order set to the panel-specific reference level (S for SM, T for TJ, X for XA).
Figure S4-2B. Non-z cascade reconstruction for /au/ in the Xi'an-speaker dataset
Representative non-z direct-versus-cascade trajectory comparison for /au/. These panels are shown across sex and variety, with carrier type fixed at 1, left segment fixed at m, and reading order set to the panel-specific reference level (S for SM, T for TJ, X for XA).
Figure S4-3. Pooled bridge marginal means (non-z)
Adjusted direct and bridged F1-PC scores with 95% confidence intervals from the all-origin pooled bridge.
Figure S4-4. Adjusted relative configuration (non-z)
Centered trajectories from covariate-adjusted PC1–PC2 marginal means.
Figure S4-5. Adjusted relative similarity over time (non-z; /ai/ + /au/)
Correlations across the four tones at each of the 11 time points, calculated from covariate-adjusted PC1–PC2 reconstructions centered within origin × variety cell and domain. Negative values indicate inverse ordering or spacing.
Figure S4-6. Pooled adjusted PC time shapes (non-z)
Tone-wise adjusted PC1 and PC2 contributions in the four origin × variety cells.
Figure S4-7. /ai/ pooled bridge marginal means (non-z)
Adjusted direct and bridged scores with 95% confidence intervals.
Figure S4-8. /ai/ adjusted relative configuration (non-z)
Centered reconstructions from adjusted PC1–PC2 scores.
Figure S4-9. Adjusted relative similarity over time (non-z; ai)
Correlations across the four tones at each of the 11 time points, calculated from covariate-adjusted PC1–PC2 reconstructions centered within origin × variety cell and domain. Negative values indicate inverse ordering or spacing.
Figure S4-10. /au/ pooled bridge marginal means (non-z)
Adjusted direct and bridged scores with 95% confidence intervals.
Figure S4-11. /au/ adjusted relative configuration (non-z)
Centered reconstructions from adjusted PC1–PC2 scores.
Figure S4-12. Adjusted relative similarity over time (non-z; au)
Correlations across the four tones at each of the 11 time points, calculated from covariate-adjusted PC1–PC2 reconstructions centered within origin × variety cell and domain. Negative values indicate inverse ordering or spacing.
The FoF analyses treat the normalized f0 contour as a functional predictor and ask whether the whole f0 trajectory predicts the whole F1 trajectory, rather than only the concurrent relation between f0(t) and F1(t) at the same time point. The target model can be written as
where s indexes time on the input f0 curve, t indexes time on the output F1 curve, Z_i contains token-level covariates, and \(\beta(s,t)\) is a smooth coefficient surface describing how input-time deviations in f0 are associated with output-time variation in F1.
In practice, each token is converted from the original long format into an 11-point f0 curve and an 11-point F1 curve. Missing internal points are linearly interpolated, and leading or trailing gaps are filled by nearest-neighbor carry-forward/carry-backward so that each retained token has a complete grid on 0:10. The f0 matrix is then centered at each input time point, so that the FoF term is interpreted as the effect of token-level deviations from the average normalized f0 contour rather than the effect of the grand mean trajectory itself. The integral is approximated by trapezoidal quadrature, so the discretized FoF contribution becomes
The implementation uses the matrix covariates S, Tmat, and L. For each output row F1_i(t_j), S stores the full input-time grid 0:10, Tmat stores the current output time repeated across columns, and L stores the centered and trapezoidally weighted full f0 curve for that token. The FoF term is then fitted in mgcv::bam() as te(S, Tmat, by = L), together with parametric condition effects, smooths of output time and duration, factor-specific output-time smooths, a speaker-specific factor-smooth term, and an AR(1) residual structure. Cross-validation is speaker-blocked, so held-out speakers are never seen during training; predictions exclude the speaker-specific smooth so that evaluation reflects transferable whole-curve structure rather than speaker memorization.
The FoF coefficient surface shown later in this section is a visualization of \(\beta(s,t)\) on the discrete 11-by-11 grid. It is extracted from the full-data refit by probing the te(S,Tmat):L term one input time point at a time and then plotting the resulting grid as a heatmap. We verified this extraction in two formally equivalent ways: a predict(..., type = "terms") probe implementation and an lpmatrix implementation that isolates the same one-hot input location. The term-prediction and lpmatrix methods yielded identical coefficient estimates and standard errors at all 121 grid points in both datasets.
| dataset | scale | cv_RMSE | cv_R2 |
|---|---|---|---|
| Tianjin | Z | 0.544 | 0.698 |
| Xi'an | Z | 0.646 | 0.566 |
| Tianjin | notZ | 0.591 | 0.699 |
| Xi'an | notZ | 0.568 | 0.592 |
| dataset | scale | exact_diagonal_percent | plus_minus_1_band_percent | weighted_mean_distance |
|---|---|---|---|---|
| Tianjin | Z | 9.370 | 22.400 | 4.416 |
| Xi'an | Z | 10.817 | 30.324 | 3.148 |
| Tianjin | notZ | 9.068 | 23.128 | 4.277 |
| Xi'an | notZ | 12.071 | 32.817 | 2.938 |
| dataset | scale | hardest_condition | hardest_RMSE | easiest_condition | easiest_RMSE |
|---|---|---|---|---|---|
| Tianjin | Z | T3 aj SM M | 0.654 | T2 aj SM F | 0.422 |
| Xi'an | Z | T3 aj SM M | 0.860 | T3 aw XA F | 0.478 |
| Tianjin | notZ | T1 aw SM F | 0.646 | T4 aw SM M | 0.432 |
| Xi'an | notZ | T1 aj XA F | 0.625 | T4 aw SM M | 0.367 |
This table compares the reduced FoF model without the tensor-product surface to the corresponding model including te(S, Tmat, by = L). In both datasets, the tensor-product term significantly improved model fit and reduced AIC, indicating that a distributed input-output surface captures structure beyond a simpler non-surface specification.
| Dataset | Score without te | Edf without te | Score with te | Edf with te | Difference | Df | p.value | ΔAIC (with - without) |
|---|---|---|---|---|---|---|---|---|
| Tianjin | 20800.83 | 32 | 20715.28 | 38 | 85.543 | 6 | < 2e-16 | -226.00 |
| Xi'an | 26170.80 | 32 | 24755.23 | 38 | 1415.569 | 6 | < 2e-16 | -2847.98 |
The main text supplements the direct full-trajectory FoF comparison with a residualized evaluation that targets tone-related structure more directly. The goal is to remove the shared diphthong trajectory within each diphthong-variety-by-sex-by-time cell before comparing observed and predicted curves.
The implementation proceeds cell by cell. Let g index a diphthong-variety-by-sex-by-time cell, and let tau index the four tones. For each curve type c in {Observed, Predicted}, a tone-equal shared baseline is computed as:
This baseline is tone-equal rather than token-count-weighted, so the common trajectory is not dominated by tone imbalance in the sample. Residualized values are then defined as:
The residualized tone-structure coefficient of determination is:
This statistic asks how much of the tone-structured F1 variation is recovered after the shared diphthong trajectory has been removed. The associated \(NRMSE_{\text{tone}}\) rescales the residualized RMSE by the observed residual standard deviation:
A complementary tone-separation recovery analysis summarizes how much observed tonal spacing is retained in the predicted residual curves. Within each cell g, the mean residual is first computed for each tone. Pairwise tone separation is then defined as the average of the six absolute pairwise differences among those four tone means, and the pairwise recovery ratio is:
An RMS-based separation ratio is also retained, using the root-mean-square of the four tone-specific residual means within each cell. Higher values in these recovery ratios indicate that predicted curves preserve more of the observed tone separation after removal of the common diphthong trajectory.
Under these residualized FoF measures, the Xi'an-speaker dataset shows stronger recovery of tone-related structure than the Tianjin-speaker dataset. The overall residualized \(R^2_{\text{tone}}\) is 0.023 in Tianjin and 0.124 in Xi'an, and the overall pairwise tone-separation recovery ratio is 0.365 in Tianjin and 0.616 in Xi'an.
| dataset | n | rmse_tone | sd_tone_target | nrmse_tone | r2_tone |
|---|---|---|---|---|---|
| Tianjin | 26,433 | 0.540 | 0.547 | 0.988 | 0.023 |
| Xi'an | 25,927 | 0.647 | 0.692 | 0.936 | 0.124 |
| dataset | condition | n | rmse_tone | sd_tone_target | nrmse_tone | r2_tone |
|---|---|---|---|---|---|---|
| Tianjin | aj SM | 6,171 | 0.565 | 0.574 | 0.986 | 0.029 |
| Tianjin | aj TJ | 6,171 | 0.543 | 0.541 | 1.005 | -0.010 |
| Tianjin | aw SM | 7,040 | 0.532 | 0.540 | 0.986 | 0.027 |
| Tianjin | aw TJ | 7,051 | 0.522 | 0.534 | 0.978 | 0.043 |
| Xi'an | aj SM | 5,918 | 0.704 | 0.764 | 0.921 | 0.151 |
| Xi'an | aj XA | 5,654 | 0.654 | 0.700 | 0.934 | 0.128 |
| Xi'an | aw SM | 7,183 | 0.628 | 0.653 | 0.962 | 0.075 |
| Xi'an | aw XA | 7,172 | 0.611 | 0.651 | 0.940 | 0.117 |
Here, the recovery-ratio columns indicate how fully observed tone separation is preserved (values closer to 1 are better), whereas the remaining columns report the absolute observed and predicted separation magnitudes.
| dataset | pairwise_recovery_ratio | rms_recovery_ratio | mean_obs_pairwise_sep | mean_pred_pairwise_sep | mean_obs_rms_spread | mean_pred_rms_spread |
|---|---|---|---|---|---|---|
| Tianjin | 0.365 | 0.363 | 0.181 | 0.066 | 0.128 | 0.046 |
| Xi'an | 0.616 | 0.613 | 0.466 | 0.287 | 0.325 | 0.199 |
As in Table S5f, values closer to 1 in the recovery-ratio columns indicate better preservation of observed tone separation within each condition, while the remaining columns report the absolute observed and predicted separation magnitudes.
| dataset | condition | pairwise_recovery_ratio | rms_recovery_ratio | mean_obs_pairwise_sep | mean_pred_pairwise_sep |
|---|---|---|---|---|---|
| Tianjin | aj SM | 0.374 | 0.373 | 0.228 | 0.085 |
| Tianjin | aj TJ | 0.307 | 0.310 | 0.190 | 0.058 |
| Tianjin | aw SM | 0.403 | 0.402 | 0.155 | 0.062 |
| Tianjin | aw TJ | 0.384 | 0.374 | 0.149 | 0.057 |
| Xi'an | aj SM | 0.577 | 0.571 | 0.562 | 0.324 |
| Xi'an | aj XA | 0.434 | 0.433 | 0.546 | 0.237 |
| Xi'an | aw SM | 0.857 | 0.851 | 0.368 | 0.315 |
| Xi'an | aw XA | 0.699 | 0.700 | 0.389 | 0.272 |
10Hz smoothing stepAs a supplementary robustness check, we repeated the FoF analyses using the unsmoothed f0 input (f0_ori) instead of the 10Hz-smoothed f0_10 trajectory before within-speaker z normalization. Across both datasets, the main predictive results changed only minimally. In Tianjin, the cross-validated overall RMSE/R2 changed from 0.54423/0.69818 to 0.54405/0.69839; in Xi'an, they changed from 0.64581/0.56604 to 0.64571/0.56616. The residualized tone-structure and tone-separation summaries also showed only very small shifts (Tianjin R^2_tone: 0.02338 -> 0.02422; Xi'an R^2_tone: 0.12422 -> 0.12459), indicating that the main whole-curve conclusions are not materially dependent on the 10Hz smoothing step.
Differences in the estimated coefficient surfaces were most apparent near the temporal boundaries, despite the small changes in prediction metrics.
Figure S5h-1. FoF coefficient surface without the 10Hz smoothing step (Tianjin-speaker dataset)
Coefficient surface from the full-data FoF refit using unsmoothed f0_ori rather than the 10Hz-smoothed f0_10 input.
Figure S5h-2. FoF coefficient surface without the 10Hz smoothing step (Xi'an-speaker dataset)
Coefficient surface from the full-data FoF refit using unsmoothed f0_ori rather than the 10Hz-smoothed f0_10 input.
Figure S5-1. FoF cross-validated mean curves by tone (z scale, Tianjin-speaker dataset)
Observed-versus-predicted mean curves under speaker-blocked FoF cross-validation.
Figure S5-2. FoF cross-validated mean curves by tone (z scale, Xi'an-speaker dataset)
Observed-versus-predicted mean curves under speaker-blocked FoF cross-validation.
Figure S5-3. FoF coefficient surface (z scale, Tianjin-speaker dataset)
Coefficient surface extracted from the full-data FoF refit after the cross-validation procedure.
Figure S5-4. FoF coefficient surface (z scale, Xi'an-speaker dataset)
Coefficient surface extracted from the full-data FoF refit after the cross-validation procedure.
Figure S5-5. FoF cross-validated mean curves by tone (non-z, Tianjin-speaker dataset)
Observed-versus-predicted mean curves from the non-z FoF rerun.
Figure S5-6. FoF cross-validated mean curves by tone (non-z, Xi'an-speaker dataset)
Observed-versus-predicted mean curves from the non-z FoF rerun.
Figure S5-7. FoF coefficient surface (non-z, Tianjin-speaker dataset)
Coefficient surface extracted from the full-data non-z FoF refit after the cross-validation procedure.
Figure S5-8. FoF coefficient surface (non-z, Xi'an-speaker dataset)
Coefficient surface extracted from the full-data non-z FoF refit after the cross-validation procedure.
This section retains the separate-group analyses: two speaker-group datasets, four origin-specific variety subgroups, and a combined SM dataset. These use separately estimated functional bases and supplement the primary common-basis pooled 2 × 2 analysis in S3. The z-scale results appear in S6a–S6g, followed by their non-z counterparts in S6h–S6i. Condition-specific bridge reconstructions from these separate-group analyses are collected in S7b.
| analysis_id | analysis_label | n_tokens | n_speakers | group_var |
|---|---|---|---|---|
| group_tianjin | Group analysis: SM vs TJ | 2,403 | 24 | ToneVar |
| group_xian | Group analysis: SM vs XA | 2,357 | 24 | ToneVar |
| subgroup_tianjin_sm | Subgroup analysis: Tianjin SM | 1,201 | 24 | tone |
| subgroup_tianjin_tj | Subgroup analysis: Tianjin TJ | 1,202 | 24 | tone |
| subgroup_xian_sm | Subgroup analysis: Xi'an SM | 1,191 | 24 | tone |
| subgroup_xian_xa | Subgroup analysis: Xi'an XA | 1,166 | 24 | tone |
| combined_sm | Combined SM analysis: Tianjin SM vs Xi'an SM | 2,392 | 48 | ToneSource |
This separate-group sensitivity table supports the qualitative observation that duration contributes mainly to the dynamic F1 dimension (PC2), with little effect on the global level dimension (PC1).
| analysis_id | F1_PC1_R2_without_duration | F1_PC1_R2_with_duration | delta_R2_from_duration_PC1 | F1_PC2_R2_without_duration | F1_PC2_R2_with_duration | delta_R2_from_duration_PC2 |
|---|---|---|---|---|---|---|
| combined_sm | 0.411 | 0.414 | 0.003 | 0.266 | 0.352 | 0.086 |
| group_tianjin | 0.496 | 0.497 | 0.001 | 0.145 | 0.273 | 0.129 |
| group_xian | 0.476 | 0.477 | 0.001 | 0.304 | 0.342 | 0.038 |
| subgroup_tianjin_sm | 0.494 | 0.494 | 0 | 0.172 | 0.266 | 0.093 |
| subgroup_tianjin_tj | 0.510 | 0.514 | 0.003 | 0.177 | 0.314 | 0.137 |
| subgroup_xian_sm | 0.433 | 0.434 | 0.001 | 0.317 | 0.390 | 0.072 |
| subgroup_xian_xa | 0.423 | 0.428 | 0.005 | 0.293 | 0.307 | 0.014 |
These summaries quantify how well low-dimensional f0 structure recovers low-dimensional F1 structure in the bridge analysis. The component-wise bridge \(R^2\) values indicate how much variance in F1-PC1 to F1-PC3 is explained by f0-PC1 to f0-PC3, while the mean absolute differences show how far the bridged scores remain from the directly observed F1 scores in each component. The Euclidean discrepancy measures summarize this direct-versus-bridged gap jointly in the PC1--PC2 and PC1--PC3 score spaces.
| analysis_group | tone_factor | tone_levels | bridge_r2_pc1 | bridge_r2_pc2 | bridge_r2_pc3 | mean_abs_diff_pc1 | mean_abs_diff_pc2 | mean_abs_diff_pc3 | mean_bridge_euclidean_pc12 | mean_bridge_euclidean_pc123 |
|---|---|---|---|---|---|---|---|---|---|---|
| Group analysis: SM vs TJ | ToneVar | 8 | 0.015 | 0.007 | 0.004 | 0.250 | 0.112 | 0.021 | 0.297 | 0.298 |
| Group analysis: SM vs XA | ToneVar | 8 | 0.173 | 0.037 | 0.013 | 0.437 | 0.271 | 0.053 | 0.559 | 0.567 |
| Subgroup analysis: Tianjin SM | tone | 4 | 0.033 | 0.016 | 0.003 | 0.235 | 0.066 | 0.009 | 0.254 | 0.254 |
| Subgroup analysis: Tianjin TJ | tone | 4 | 0.004 | 0.021 | 0.005 | 0.217 | 0.119 | 0.034 | 0.286 | 0.289 |
| Subgroup analysis: Xi'an SM | tone | 4 | 0.180 | 0.030 | 0.004 | 0.266 | 0.217 | 0.044 | 0.365 | 0.372 |
| Subgroup analysis: Xi'an XA | tone | 4 | 0.179 | 0.057 | 0.035 | 0.260 | 0.173 | 0.054 | 0.316 | 0.321 |
| Combined SM analysis: Tianjin SM vs Xi'an SM | ToneSource | 8 | 0.095 | 0.012 | 0.002 | 0.363 | 0.242 | 0.040 | 0.472 | 0.475 |
| analysis_group | panel | mean_rank_cor | mean_level_cor |
|---|---|---|---|
| Group analysis: SM vs TJ | SM | -0.691 | -0.735 |
| Group analysis: SM vs TJ | TJ | -0.618 | -0.675 |
| Group analysis: SM vs XA | SM | -0.891 | -0.928 |
| Group analysis: SM vs XA | XA | -0.836 | -0.915 |
| Combined SM analysis: Tianjin SM vs Xi'an SM | Tianjin SM | -0.727 | -0.725 |
| Combined SM analysis: Tianjin SM vs Xi'an SM | Xi'an SM | -0.891 | -0.935 |
Figure S6-1. Group-level eigenfunctions (z scale, Tianjin-speaker dataset)
Pooled eigenfunctions for the Tianjin-speaker z-scale fPCA analysis.
Figure S6-2. Group-level eigenfunctions (z scale, Xi'an-speaker dataset)
Pooled eigenfunctions for the Xi'an-speaker z-scale fPCA analysis.
Figure S6-3. Paired f0-F1 PC effects (z scale, Tianjin-speaker dataset)
Pooled paired f0-F1 PC effect plots for the Tianjin-speaker z-scale analysis.
Figure S6-4. Paired f0-F1 PC effects (z scale, Xi'an-speaker dataset)
Pooled paired f0-F1 PC effect plots for the Xi'an-speaker z-scale analysis.
Figure S6-5. Group-level adjusted PC time-shape curves (z scale, Tianjin-speaker dataset)
Covariate-adjusted mixed-model PC reconstructions for the Tianjin-speaker pooled analysis.
Figure S6-6. Group-level adjusted PC time-shape curves (z scale, Xi'an-speaker dataset)
Covariate-adjusted mixed-model PC reconstructions for the Xi'an-speaker pooled analysis.
Figure S6-7. Group-level bridge marginal means (z scale, Tianjin-speaker dataset)
Direct-versus-bridged F1-PC marginal means for the Tianjin-speaker pooled analysis.
Figure S6-8. Group-level bridge marginal means (z scale, Xi'an-speaker dataset)
Direct-versus-bridged F1-PC marginal means for the Xi'an-speaker pooled analysis.
Figure S6-9. Adjusted relative configuration (z scale, Tianjin-speaker dataset)
Model-adjusted relative f0-F1 configuration trajectories.
Figure S6-10. Adjusted relative configuration (z scale, Xi'an-speaker dataset)
Model-adjusted relative f0-F1 configuration trajectories.
Figure S6-11. Adjusted relative similarity by time (z scale, Tianjin-speaker dataset)
By-time rank and level similarity for the adjusted relative configuration analysis.
Figure S6-12. Adjusted relative similarity by time (z scale, Xi'an-speaker dataset)
By-time rank and level similarity for the adjusted relative configuration analysis.
This table shows that the direct-versus-bridged difference varies significantly across pooled group-level conditions for all three F1 principal components, indicating that bridge reconstruction only partially preserves the observed condition-structured score pattern.
| analysis_id | outcome | Chisq | Df | p_value |
|---|---|---|---|---|
| group_tianjin | F1_PC1 | 1,478 | 6 | < 0.001 |
| group_tianjin | F1_PC2 | 217 | 6 | < 0.001 |
| group_tianjin | F1_PC3 | 85.925 | 6 | < 0.001 |
| group_xian | F1_PC1 | 524.5 | 6 | < 0.001 |
| group_xian | F1_PC2 | 313.5 | 6 | < 0.001 |
| group_xian | F1_PC3 | 74.143 | 6 | < 0.001 |
z scale)Table S6f-1. Group-level split bridge summaries
| dataset_group | diphthong | bridge_r2_pc1 | bridge_r2_pc2 | mean_bridge_euclidean_pc12 |
|---|---|---|---|---|
| Tianjin-speaker dataset | ai | 0.006 | 0.032 | 0.316 |
| Xi'an-speaker dataset | ai | 0.241 | 0.051 | 0.692 |
| Tianjin-speaker dataset | au | 0.024 | 0.004 | 0.213 |
| Xi'an-speaker dataset | au | 0.156 | 0.043 | 0.473 |
Table S6f-2. Group-level split relative-configuration summaries
| dataset_group | panel | diphthong | mean_rank_cor | mean_level_cor |
|---|---|---|---|---|
| Tianjin-speaker dataset | SM | ai | -0.782 | -0.793 |
| Tianjin-speaker dataset | TJ | ai | -0.727 | -0.807 |
| Xi'an-speaker dataset | SM | ai | -0.909 | -0.933 |
| Xi'an-speaker dataset | XA | ai | -0.836 | -0.894 |
| Tianjin-speaker dataset | SM | au | -0.727 | -0.726 |
| Tianjin-speaker dataset | TJ | au | -0.418 | -0.552 |
| Xi'an-speaker dataset | SM | au | -0.873 | -0.930 |
| Xi'an-speaker dataset | XA | au | -0.909 | -0.927 |
Figure S6-13. Split-by-diphthong bridge comparison (z scale, Tianjin-speaker /ai/)
Direct-versus-bridged F1-PC summaries from the fully split /ai/ subset of the Tianjin-speaker group.
Figure S6-14. Split-by-diphthong bridge comparison (z scale, Tianjin-speaker /au/)
Direct-versus-bridged F1-PC summaries from the fully split /au/ subset of the Tianjin-speaker group.
Figure S6-15. Split-by-diphthong bridge comparison (z scale, Xi'an-speaker /ai/)
Direct-versus-bridged F1-PC summaries from the fully split /ai/ subset of the Xi'an-speaker group.
Figure S6-16. Split-by-diphthong bridge comparison (z scale, Xi'an-speaker /au/)
Direct-versus-bridged F1-PC summaries from the fully split /au/ subset of the Xi'an-speaker group.
Figure S6-17. Split-by-diphthong relative similarity (z scale, Tianjin-speaker /ai/)
By-time spearman rank and pearson level similarity from the fully split /ai/ subset of the Tianjin-speaker group.
Figure S6-18. Split-by-diphthong relative similarity (z scale, Tianjin-speaker /au/)
By-time spearman rank and pearson level similarity from the fully split /au/ subset of the Tianjin-speaker group.
Figure S6-19. Split-by-diphthong relative similarity (z scale, Xi'an-speaker /ai/)
By-time spearman rank and pearson level similarity from the fully split /ai/ subset of the Xi'an-speaker group.
Figure S6-20. Split-by-diphthong relative similarity (z scale, Xi'an-speaker /au/)
By-time spearman rank and pearson level similarity from the fully split /au/ subset of the Xi'an-speaker group.
f0+F1 robustness analysisAs a convergence check on the separate f0 and F1 decompositions used in the main text, a multivariate FDA was also run on paired f0 and F1 trajectories within each diphthong. This multivariate analysis uses the same 11-point normalized trajectories and the same split-by-diphthong logic as the supporting analyses above, but treats f0 and F1 as two dimensions of a single functional object rather than concatenating them or decomposing them separately.
The multivariate rerun uses f0_10Z_interp and f1Z and retains only tokens with at least five observed points in each signal before interpolation.
Table S6g-1. Split-by-diphthong explained variance in the multivariate f0+F1 analysis
| dataset_group | diphthong | PC1_variance | PC2_variance | PC3_variance | PC1_to_PC3_cumulative |
|---|---|---|---|---|---|
| Tianjin-speaker dataset | /ai/ | 0.559 | 0.181 | 0.101 | 0.840 |
| Tianjin-speaker dataset | /au/ | 0.526 | 0.222 | 0.077 | 0.825 |
| Xi'an-speaker dataset | /ai/ | 0.502 | 0.182 | 0.153 | 0.837 |
| Xi'an-speaker dataset | /au/ | 0.503 | 0.198 | 0.136 | 0.838 |
Across all four split analyses, the first multivariate component accounts for roughly one half of the total joint f0+F1 variance, and the first three components together account for approximately 82--84% of the total variance. These results indicated that the shared f0+F1 structure is low-dimensional even when the two signals are modeled jointly.
Table S6g-2. Incremental contribution of duration in the split multivariate f0+F1 score-space models
| dataset_group | diphthong | MV_PC1_R2_without_duration | MV_PC1_R2_with_duration | MV_PC1_delta_R2_duration | MV_PC2_R2_without_duration | MV_PC2_R2_with_duration | MV_PC2_delta_R2_duration | MV_PC3_R2_without_duration | MV_PC3_R2_with_duration | MV_PC3_delta_R2_duration |
|---|---|---|---|---|---|---|---|---|---|---|
| Tianjin-speaker dataset | /ai/ | 0.697 | 0.698 | 0.001 | 0.555 | 0.557 | 0.002 | 0.120 | 0.168 | 0.048 |
| Tianjin-speaker dataset | /au/ | 0.615 | 0.618 | 0.003 | 0.639 | 0.640 | 0.001 | 0.220 | 0.224 | 0.004 |
| Xi'an-speaker dataset | /ai/ | 0.782 | 0.785 | 0.003 | 0.714 | 0.715 | 0 | 0.227 | 0.229 | 0.001 |
| Xi'an-speaker dataset | /au/ | 0.788 | 0.789 | 0.001 | 0.763 | 0.763 | 0 | 0.222 | 0.223 | 0.001 |
The duration follow-up shows a selective pattern: duration contributes very little to the dominant shared component in most cells, but can contribute more strongly to higher multivariate components (PC3), especially for the Tianjin /ai/ split. This is compatible with the main-text interpretation that the strongest shared f0-F1 structure is not reducible to duration alone.
Figure S6-21A. Multivariate f0+F1 score space for the Tianjin-speaker /ai/ split
Shared multivariate score space for the split-by-diphthong Tianjin /ai/ analysis.
Figure S6-21B. Multivariate f0+F1 score space for the Tianjin-speaker /au/ split
Shared multivariate score space for the split-by-diphthong Tianjin /au/ analysis.
Figure S6-21C. Multivariate f0+F1 score space for the Xi'an-speaker /ai/ split
Shared multivariate score space for the split-by-diphthong Xi'an /ai/ analysis.
Figure S6-21D. Multivariate f0+F1 score space for the Xi'an-speaker /au/ split
Shared multivariate score space for the split-by-diphthong Xi'an /au/ analysis.
Figure S6-22A. Multivariate f0+F1 reconstruction for the Tianjin-speaker /ai/ split
Joint reconstruction of paired f0 and F1 trajectories from the first three multivariate components in the Tianjin /ai/ split.
Figure S6-22B. Multivariate f0+F1 reconstruction for the Tianjin-speaker /au/ split
Joint reconstruction of paired f0 and F1 trajectories from the first three multivariate components in the Tianjin /au/ split.
Figure S6-22C. Multivariate f0+F1 reconstruction for the Xi'an-speaker /ai/ split
Joint reconstruction of paired f0 and F1 trajectories from the first three multivariate components in the Xi'an /ai/ split.
Figure S6-22D. Multivariate f0+F1 reconstruction for the Xi'an-speaker /au/ split
Joint reconstruction of paired f0 and F1 trajectories from the first three multivariate components in the Xi'an /au/ split.
Table S6h-1. Additional contribution of duration in the non-z score-space models
| analysis_group | F1_PC1_R2_without_duration | F1_PC1_R2_with_duration | delta_R2_from_duration_PC1 | F1_PC2_R2_without_duration | F1_PC2_R2_with_duration | delta_R2_from_duration_PC2 |
|---|---|---|---|---|---|---|
| combined_sm | 0.380 | 0.389 | 0.009 | 0.424 | 0.472 | 0.047 |
| group_tianjin | 0.504 | 0.504 | 0 | 0.187 | 0.278 | 0.091 |
| group_xian | 0.463 | 0.464 | 0.001 | 0.415 | 0.447 | 0.032 |
| subgroup_tianjin_sm | 0.527 | 0.528 | 0.001 | 0.198 | 0.250 | 0.052 |
| subgroup_tianjin_tj | 0.501 | 0.504 | 0.003 | 0.214 | 0.328 | 0.115 |
| subgroup_xian_sm | 0.412 | 0.416 | 0.004 | 0.445 | 0.500 | 0.055 |
| subgroup_xian_xa | 0.410 | 0.415 | 0.005 | 0.408 | 0.421 | 0.013 |
Table S6h-2. Non-z bridge reconstruction summaries
| analysis_group | grouping_factor | grouping_level | bridge_r2_pc1 | bridge_r2_pc2 | bridge_r2_pc3 | mean_abs_diff_pc1 | mean_abs_diff_pc2 | mean_abs_diff_pc3 | mean_bridge_euclidean_pc12 | mean_bridge_euclidean_pc123 |
|---|---|---|---|---|---|---|---|---|---|---|
| Group analysis: SM vs TJ | ToneVar | 8 | 0.017 | 0.007 | 0.002 | 0.244 | 0.125 | 0.017 | 0.293 | 0.294 |
| Group analysis: SM vs XA | ToneVar | 8 | 0.169 | 0.039 | 0.004 | 0.357 | 0.219 | 0.062 | 0.455 | 0.466 |
| Subgroup analysis: Tianjin SM | tone | 4 | 0.030 | 0.022 | 0.002 | 0.216 | 0.088 | 0.009 | 0.253 | 0.254 |
| Subgroup analysis: Tianjin TJ | tone | 4 | 0.013 | 0.023 | 0.004 | 0.202 | 0.138 | 0.022 | 0.279 | 0.280 |
| Subgroup analysis: Xi'an SM | tone | 4 | 0.185 | 0.034 | 0.001 | 0.205 | 0.184 | 0.042 | 0.291 | 0.296 |
| Subgroup analysis: Xi'an XA | tone | 4 | 0.164 | 0.059 | 0.014 | 0.220 | 0.133 | 0.066 | 0.264 | 0.272 |
| Combined SM analysis: Tianjin SM vs Xi'an SM | ToneSource | 8 | 0.085 | 0.009 | 0 | 0.336 | 0.352 | 0.035 | 0.521 | 0.522 |
Table S6h-3. Non-z covariate-adjusted relative-configuration summary
| analysis_group | panel | mean_rank_cor | mean_level_cor |
|---|---|---|---|
| Group analysis: SM vs TJ | SM | -0.727 | -0.740 |
| Group analysis: SM vs TJ | TJ | -0.582 | -0.704 |
| Group analysis: SM vs XA | SM | -0.873 | -0.936 |
| Group analysis: SM vs XA | XA | -0.818 | -0.919 |
| Combined SM analysis: Tianjin SM vs Xi'an SM | Tianjin | -0.727 | -0.739 |
| Combined SM analysis: Tianjin SM vs Xi'an SM | Xi'an | -0.891 | -0.947 |
Table S6h-4. Non-z method-context sensitivity checks
| analysis_id | outcome | Chisq | Df | p_value |
|---|---|---|---|---|
| group_tianjin | F1_PC1 | 1,529 | 6 | < 0.001 |
| group_tianjin | F1_PC2 | 141.8 | 6 | < 0.001 |
| group_tianjin | F1_PC3 | 42.037 | 6 | < 0.001 |
| group_xian | F1_PC1 | 498.9 | 6 | < 0.001 |
| group_xian | F1_PC2 | 290.5 | 6 | < 0.001 |
| group_xian | F1_PC3 | 63.663 | 6 | < 0.001 |
z summariesTable S6i-1. Group-level split bridge summaries under non-z scaling
| dataset_group | diphthong | bridge_r2_pc1 | bridge_r2_pc2 | mean_bridge_euclidean_pc12 |
|---|---|---|---|---|
| Tianjin-speaker dataset | ai | 0.003 | 0.038 | 0.297 |
| Xi'an-speaker dataset | ai | 0.233 | 0.046 | 0.562 |
| Tianjin-speaker dataset | au | 0.020 | 0.013 | 0.213 |
| Xi'an-speaker dataset | au | 0.151 | 0.050 | 0.390 |
Table S6i-2. Group-level split relative summaries under non-z scaling
| dataset_group | panel | diphthong | mean_rank_cor | mean_level_cor |
|---|---|---|---|---|
| Tianjin-speaker dataset | SM | ai | -0.745 | -0.762 |
| Tianjin-speaker dataset | TJ | ai | -0.782 | -0.830 |
| Xi'an-speaker dataset | SM | ai | -0.891 | -0.946 |
| Xi'an-speaker dataset | XA | ai | -0.855 | -0.890 |
| Tianjin-speaker dataset | SM | au | -0.691 | -0.709 |
| Tianjin-speaker dataset | TJ | au | -0.455 | -0.594 |
| Xi'an-speaker dataset | SM | au | -0.873 | -0.932 |
| Xi'an-speaker dataset | XA | au | -0.891 | -0.936 |
Figure S6-23. Group-level bridge marginal means (non-z, Tianjin-speaker dataset)
Direct-versus-bridged F1-PC marginal means from the non-z analysis.
Figure S6-24. Group-level bridge marginal means (non-z, Xi'an-speaker dataset)
Direct-versus-bridged F1-PC marginal means from the non-z analysis.
Figure S6-25. Adjusted relative configuration (non-z, Tianjin-speaker dataset)
Model-adjusted relative configuration from the non-z analysis.
Figure S6-26. Adjusted relative configuration (non-z, Xi'an-speaker dataset)
Model-adjusted relative configuration from the non-z analysis.
Figure S6-27. Adjusted relative similarity by time (non-z, Tianjin-speaker dataset)
By-time rank and level similarity from the non-z analysis.
Figure S6-28. Adjusted relative similarity by time (non-z, Xi'an-speaker dataset)
By-time rank and level similarity from the non-z analysis.
This section presents condition-specific F1 reconstructions from the three modelling frameworks. Each panel shows one variety × diphthong condition, with the four tones distinguished by color. The plots provide descriptive comparisons of reconstructed trajectory shapes.
These GAMM panels use the cascade prediction only, under the reference-style setting used in the follow-up script (leftSegment = m, type = 1, with order held at the script's reference value for each variety).
Figure S7-1. Cascade reconstructed F1 curves by condition in the Tianjin-speaker dataset
Group-level cascade reconstructions from the z-scale analysis, shown by variety × diphthong condition with tones overlaid.
Figure S7-2. Cascade reconstructed F1 curves by condition in the Xi'an-speaker dataset
Group-level cascade reconstructions from the z-scale analysis, shown by variety × diphthong condition with tones overlaid.
These figures are reconstructed from the separate-group, split-by-diphthong bridge outputs summarized in S6f; the primary pooled 2 × 2 bridge results are in S3c and S3e. Each panel corresponds to one variety × diphthong condition, so the two varieties within each group analysis are separated.
Figure S7-3. fPCA bridge reconstruction by tone
Bridge-reconstructed F1 curves for the SM /ai/ condition in the Tianjin-speaker group analysis.
Figure S7-4. fPCA bridge reconstruction by tone
Bridge-reconstructed F1 curves for the TJ /ai/ condition in the Tianjin-speaker group analysis.
Figure S7-5. fPCA bridge reconstruction by tone
Bridge-reconstructed F1 curves for the SM /au/ condition in the Tianjin-speaker group analysis.
Figure S7-6. fPCA bridge reconstruction by tone
Bridge-reconstructed F1 curves for the TJ /au/ condition in the Tianjin-speaker group analysis.
Figure S7-7. fPCA bridge reconstruction by tone
Bridge-reconstructed F1 curves for the SM /ai/ condition in the Xi'an-speaker group analysis.
Figure S7-8. fPCA bridge reconstruction by tone
Bridge-reconstructed F1 curves for the XA /ai/ condition in the Xi'an-speaker group analysis.
Figure S7-9. fPCA bridge reconstruction by tone
Bridge-reconstructed F1 curves for the SM /au/ condition in the Xi'an-speaker group analysis.
Figure S7-10. fPCA bridge reconstruction by tone
Bridge-reconstructed F1 curves for the XA /au/ condition in the Xi'an-speaker group analysis.
These FoF panels are based on the cross-validated predicted curves only.
Figure S7-11. FoF predicted-only reconstruction by tone
FoF predicted-only F1 curves for the SM /ai/ condition in the Tianjin-speaker dataset.
Figure S7-12. FoF predicted-only reconstruction by tone
FoF predicted-only F1 curves for the TJ /ai/ condition in the Tianjin-speaker dataset.
Figure S7-13. FoF predicted-only reconstruction by tone
FoF predicted-only F1 curves for the SM /au/ condition in the Tianjin-speaker dataset.
Figure S7-14. FoF predicted-only reconstruction by tone
FoF predicted-only F1 curves for the TJ /au/ condition in the Tianjin-speaker dataset.
Figure S7-15. FoF predicted-only reconstruction by tone
FoF predicted-only F1 curves for the SM /ai/ condition in the Xi'an-speaker dataset.
Figure S7-16. FoF predicted-only reconstruction by tone
FoF predicted-only F1 curves for the XA /ai/ condition in the Xi'an-speaker dataset.
Figure S7-17. FoF predicted-only reconstruction by tone
FoF predicted-only F1 curves for the SM /au/ condition in the Xi'an-speaker dataset.
Figure S7-18. FoF predicted-only reconstruction by tone
FoF predicted-only F1 curves for the XA /au/ condition in the Xi'an-speaker dataset.