Supplementary Material

Supplementary Material

for the article Diphthong Dynamics under Lexical Tone: Cross-Dialectal Evidence for Category-Specific f0--F1 Coupling

Authors: Chenyu Li and Jalal Al-Tamimi from Laboratoire de linguistique formelle, CNRS, Paris 75013, France

S0. Full analysis pipeline and reading guide

The supplementary material follows the main stages of the analysis. S1 documents data screening and f0 trackability. S2 reports the GAMM analyses and the GAMM-based cascade reconstruction. S3 reports the primary pooled 2 × 2 z-scale fPCA analyses, including bridge and relative-configuration results. S4 provides the speaker-centered non-z robustness reruns. S5 reports the full-curve FoF analyses and residualized tone-recovery metrics. S6 reports the supplementary seven-group fPCA analyses and their non-z reruns, including the multivariate check in S6g. S7 gathers reconstruction-oriented visual summaries across methods.

Full GAMM model summaries, diagnostics, and complete compareML outputs are retained in the accompanying analysis notebook: SuppPub2_GAMM_summaries.html.

S0a. Participants, variety assessment, experimental procedure, and item inventory

Forty-eight native speakers participated: 24 from Xi'an (12 female, 12 male; 18–65 years; mean 42.875, SD 13.19) and 24 from Tianjin (12 female, 12 male; 20–65 years; mean 44.08, SD 14.10). All were native speakers of their local Mandarin variety and reported regular daily use of Standard Mandarin (SM). Local-variety background and SM proficiency were assessed using a questionnaire, self-report, and experimenter evaluation.

Each participant produced both SM and their local variety. Within each location, half completed the SM block first and half the local-variety block first. Before each block, participants read several naturalistic sentences in the intended variety to establish the corresponding speech mode. Target items were randomized within each variety block. Each item was repeated twice within the two carrier sentences described below. Misproduced items were repeated; persistently inaccurate tokens were excluded. After recording, two native SM speakers conducted a blind check of tone-category accuracy across diphthongs.

Recordings took place in a quiet sound-insulated environment using a Zoom H6 Handy Recorder and Shure WH20 head-mounted dynamic microphone, at 48 kHz and 24-bit resolution. The onset inventory was restricted to non-aspirated bilabial and alveolar consonants to limit anticipatory coarticulation and aspiration-related glottal effects. The two carrier sentences provided contrasting preceding tonal contexts within each variety and minimized systematic influences from the following syllable. Acoustic alignment and tracking procedures are detailed in S1f.

The stimulus materials were elicited in two carrier sentences: “Please read X twice” (请将X读两遍) and “Please play X eight times” (请播放X八遍) corresponding to IPA forms tɕʰjəŋT3 tɕjaŋT1 X tuT2 ljaŋT3 pjɛnT4 and tɕʰjəŋT3 poT1 faŋT4 X paT1 pjɛnT4. The item space combined two falling diphthongs (/ai/, /au/), four non-aspirated bilabial or alveolar onsets (/m/, /p/, /n/, /t/), and the four tonal categories (T1T4), yielding 32 theoretically possible onset-by-diphthong-by-tone slots. Because several slots lacked suitable lexical items with attested characters, the realized stimulus set contained 28 items.

The realized items were:

Tonemaimaunainaupaipautaitau
T1mai1 麦mau1 猫NAnau1 孬pai1 掰pau1 包tai1 呆tau1 刀
T2mai2 埋mau2 毛NAnau2 挠pai2 白pau2 薄NANA
T3mai3 买mau3 铆nai3 奶nau3 脑pai3 摆pau3 宝tai3 歹tau3 岛
T4mai4 卖mau4 帽nai4 奈nau4 闹pai4 败pau4 爆tai4 带tau4 到

S1. Data summary and exclusion profile

The functional analyses (fPCA analyses and function-on-function mapping) and the GAMMs were based on overlapping but not identical datasets. Because the functional analyses treated each token as a full trajectory, they required sufficiently continuous f0 curves: tokens with fewer than five observed f0 points were excluded, and any remaining missing values were filled by linear interpolation with edge filling. The GAMMs, by contrast, were fitted to broader observationwise datasets and therefore did not require the same degree of token-level f0 continuity.

Table S1a. Overall sample sizes by analysis type

Dataset groupF1 observations in baseline GAMMsf0 observations in baseline GAMMsMechanistic F1 trajectory observationsTotal tokensRetained functional analysis tokensExcluded tokens
Tianjin-speaker dataset29,51323,23023,2302,6832,403280
Xi'an-speaker dataset28,50123,08723,0872,5912,357234
Combined total58,01446,31746,3175,2744,760514

Table S1b. Exclusion summary by variety

Dataset groupVarietyTotal tokensRetained tokens for functional analysisExcluded tokensExclusion rate
Tianjin-speaker datasetSM1,3471,2011460.108
Tianjin-speaker datasetTJ1,3361,2021340.100
Xi'an-speaker datasetSM1,3461,1911550.115
Xi'an-speaker datasetXA1,2451,166790.063

Table S1c. Exclusion breakdown by tone, variety, and diphthong

Dataset groupVarietyToneDiphthongTotal tokensRetained tokensExcluded tokensExclusion rate
Tianjin-speaker datasetSM1ai13213200
Tianjin-speaker datasetSM1au18518500
Tianjin-speaker datasetSM2ai959500
Tianjin-speaker datasetSM2au15014910.007
Tianjin-speaker datasetSM3ai195131640.328
Tianjin-speaker datasetSM3au188121670.356
Tianjin-speaker datasetSM4ai21020370.033
Tianjin-speaker datasetSM4au19218570.036
Tianjin-speaker datasetTJ1ai11893250.212
Tianjin-speaker datasetTJ1au168147210.125
Tianjin-speaker datasetTJ2ai979700
Tianjin-speaker datasetTJ2au16016000
Tianjin-speaker datasetTJ3ai199155440.221
Tianjin-speaker datasetTJ3au190151390.205
Tianjin-speaker datasetTJ4ai21721610.005
Tianjin-speaker datasetTJ4au18718340.021
Xi'an-speaker datasetSM1ai12312300
Xi'an-speaker datasetSM1au18818800
Xi'an-speaker datasetSM2ai898900
Xi'an-speaker datasetSM2au14814800
Xi'an-speaker datasetSM3ai202120820.406
Xi'an-speaker datasetSM3au189132570.302
Xi'an-speaker datasetSM4ai21420680.037
Xi'an-speaker datasetSM4au19318580.041
Xi'an-speaker datasetXA1ai8159220.272
Xi'an-speaker datasetXA1au146103430.295
Xi'an-speaker datasetXA2ai696900
Xi'an-speaker datasetXA2au16416400
Xi'an-speaker datasetXA3ai19919180.040
Xi'an-speaker datasetXA3au19318760.031
Xi'an-speaker datasetXA4aj19519500
Xi'an-speaker datasetXA4au19819800

Table S1d. Highest-exclusion cells

Dataset groupVarietyToneDiphthongTotal tokensRetained tokensExcluded tokensExclusion rate
Xi'an-speaker datasetSM3ai202120820.406
Tianjin-speaker datasetSM3au188121670.356
Tianjin-speaker datasetSM3ai195131640.328
Xi'an-speaker datasetSM3au189132570.302
Xi'an-speaker datasetXA1au146103430.295
Xi'an-speaker datasetXA1ai8159220.272
Tianjin-speaker datasetTJ3ai199155440.221
Tianjin-speaker datasetTJ1ai11893250.212

Because functional exclusion is triggered by insufficient observable f0 support for interpolation, Table S1c also serves as the operational summary of f0 trackability loss by condition. The strongest exclusion concentrations occur in low-tone cells, especially SM T3 in both datasets.

S1e. Pointwise f0 missingness models

To complement the token-level exclusion audit above, we also modelled pointwise f0 missingness directly with binomial GAMMs in the Tianjin-speaker and Xi'an-speaker datasets. These models estimate the population-level probability that f0 is missing at each normalized time point for each ToneVarDiphInt condition. Together with Table S1c, they provide a more direct picture of f0 trackability loss.

Figure S1-1. Predicted probability of missing f0 in the Tianjin-speaker dataset

Population-level missingness curves from the Tianjin-speaker GAMM, with the strongest probabilities concentrated in low-tone cases (SM T3; TJ T1 and T3).

Figure S1-1. Predicted probability of missing `f0` in the Tianjin-speaker dataset
Figure S1-1. Predicted probability of missing `f0` in the Tianjin-speaker dataset

Table S1e-1. Tianjin-speaker missing-f0 summary by condition

ToneVarDiphIntmean_p_missingmax_p_missingtime_of_max
3 SM aj0.2720.49210
3 SM aw0.2690.43010
3 TJ aj0.2460.4490
3 TJ aw0.2190.4150
1 TJ aj0.1610.44910
1 TJ aw0.1320.48910
4 SM aj0.0950.42810
4 SM aw0.0840.35110
4 TJ aj0.0660.33710
4 TJ aw0.0550.24410
2 TJ aj0.0310.1620
2 SM aj0.0240.1540
2 SM aw0.0220.1120
2 TJ aw0.0180.10510
1 SM aj0.0160.1220
1 SM aw0.0100.0580

Table S1e-2. Tianjin-speaker missing-f0 summary by condition and time region

The time regions were defined over the 11-point normalized trajectory as early (points 0–3), middle (4–7), and late (8–10), and the reported values are mean predicted missing-f0 probabilities within each region.

ToneVarDiphInttime_regionmean_p_missing
3 TJ ajearly0.383
3 TJ awearly0.348
1 TJ ajlate0.328
3 SM awearly0.311
1 TJ awlate0.303
3 SM ajearly0.286
3 SM ajlate0.285
4 SM ajlate0.251
3 SM ajmiddle0.250
3 SM awlate0.248
3 SM awmiddle0.242
4 SM awlate0.227
3 TJ ajmiddle0.219
4 TJ ajlate0.197
3 TJ awmiddle0.184
4 TJ awlate0.146
1 TJ ajmiddle0.112
3 TJ ajlate0.097
3 TJ awlate0.093
1 TJ ajearly0.086
1 TJ awmiddle0.083
2 TJ ajlate0.059
1 TJ awearly0.053
4 SM ajmiddle0.044
2 SM ajearly0.044
2 TJ awlate0.043
4 SM awmiddle0.042
2 TJ ajearly0.041
2 SM awearly0.033
2 SM awlate0.031
1 SM ajearly0.031
4 SM ajearly0.027
4 TJ awearly0.021
2 SM ajlate0.020
4 TJ awmiddle0.020
4 SM awearly0.020
1 SM ajlate0.019
4 TJ ajearly0.017
1 SM awlate0.017
2 TJ awearly0.016
4 TJ ajmiddle0.016
1 SM awearly0.015
2 SM ajmiddle0.006
2 SM awmiddle0.004
2 TJ awmiddle0.002
1 SM awmiddle0
2 TJ ajmiddle0
1 SM ajmiddle0

Figure S1-2. Predicted probability of missing f0 in the Xi'an-speaker dataset

Population-level missingness curves from the Xi'an-speaker GAMM. The most severe missingness again concentrates in low-tone cases (SM T3 and XA T1).

Figure S1-2. Predicted probability of missing `f0` in the Xi'an-speaker dataset
Figure S1-2. Predicted probability of missing `f0` in the Xi'an-speaker dataset

Table S1e-3. Xi'an-speaker missing-f0 summary by condition

ToneVarDiphIntmean_p_missing_XAmax_p_missing_XAtime_of_max_XA
3 SM aj0.3410.74110
3 SM aw0.2760.66610
1 XA aj0.2190.5630
1 XA aw0.2050.47910
4 SM aw0.1120.45710
4 SM aj0.1060.43410
3 XA aw0.1020.35010
3 XA aj0.0960.32810
2 SM aj0.0360.1880
2 SM aw0.0290.13710
2 XA aw0.0280.08810
2 XA aj0.0280.0990
1 SM aj0.0230.1600
1 SM aw0.0200.1060
4 XA aw0.0180.10210
4 XA aj0.0150.0820

Table S1e-4. Xi'an-speaker missing-f0 summary by condition and time region

ToneVarDiphInttime_region_XAmean_p_missing_XA
3 SM ajlate0.582
3 SM awlate0.495
1 XA awlate0.348
1 XA ajlate0.345
3 SM ajmiddle0.319
4 SM awlate0.278
4 SM ajlate0.255
3 SM awmiddle0.239
3 XA awlate0.218
3 XA ajlate0.218
1 XA ajearly0.212
3 SM ajearly0.183
1 XA awearly0.161
3 SM awearly0.150
1 XA awmiddle0.142
1 XA ajmiddle0.131
3 XA awmiddle0.095
4 SM awmiddle0.077
4 SM ajmiddle0.075
3 XA ajmiddle0.073
2 SM ajearly0.060
2 SM awlate0.058
1 SM ajearly0.041
4 XA awlate0.041
2 XA awlate0.040
2 SM ajlate0.039
2 XA ajearly0.039
1 SM awlate0.035
2 XA ajlate0.033
1 SM ajlate0.029
2 SM awearly0.029
4 XA ajlate0.028
2 XA awearly0.028
1 SM awearly0.028
3 XA ajearly0.027
4 SM ajearly0.025
4 SM awearly0.021
3 XA awearly0.021
4 XA ajearly0.021
2 XA awmiddle0.020
4 XA awearly0.019
2 XA ajmiddle0.013
2 SM ajmiddle0.008
2 SM awmiddle0.008
1 SM awmiddle0.001
1 SM ajmiddle0.001
4 XA awmiddle0
4 XA ajmiddle0

In both datasets, low-tone conditions, especially SM T3, have the highest predicted probability of missing f0, consistent with the token-level exclusion patterns in Table S1c.

S1f. Segment boundaries, f0 extraction, and formant tracking

Segmental alignment was obtained with the Montreal Forced Aligner and manually corrected. Vowel onset was placed at the emergence of clear formant structure and offset at its disappearance. Boundaries were primarily guided by F2 and cross-checked against F1–F4. Spectrograms were generally inspected with Praat's default 70-dB dynamic range, adjusted to 50 dB when necessary for low-amplitude signals.

Acoustic measurements were extracted in Praat 6.3.02. f0 was estimated with a two-pass autocorrelation procedure: the first pass used a broad 75–600 Hz range, and the first and third quartiles of the resulting track defined an adaptive range for the second pass. The extracted f0 track was smoothed with Praat's built-in 10-Hz bandwidth smoothing function. F1 was extracted with the Burg method followed by Praat's tracking procedure. Five formants were estimated; the maximum formant frequency depended on speaker sex and diphthong as follows.

Table S1f-1. Maximum formant settings

Speaker sex/ai//au/Number of formants
Female5.5 kHz4.9 kHz5
Male5.0 kHz4.4 kHz5

f0 and F1 were sampled at 11 equidistant points across each diphthong. The primary representation used within-speaker z-scores. The non-z reruns used f0 in semitones relative to the speaker's mean and F1 in speaker-centered Bark. Praat returned no valid f0 estimate at some points; GAMMs retained partially observed tokens and omitted observations with missing model-required values. Functional analyses excluded tokens with fewer than five observed f0 points and filled remaining gaps by linear interpolation with edge filling. These trackability-related exclusions are quantified in S1a–S1e; the FoF check without 10-Hz smoothing is reported in S5h.

S2. GAMM analyses

S2a. Baseline and mechanistic GAMM fit summaries

The baseline GAMMs model tone-conditioned F1 and f0 trajectories over the normalized time axis. The mechanistic GAMMs replace the explicit tone term with time-varying f0, duration, and interaction structure. Categorical predictors (including ToneVarDiphInt and covariates) in these GAMMs were entered as ordered factors using treatment coding, and random smooths were included for speaker and item (character).

For full GAMM summary and gam.check results, please refer to the document "Full GAMM Summaries and gam.check results".

Baseline:    F1 ~ s(time) + s(time, by = ToneVarDiphInt) + covariates + random smooths + AR(1)
Mechanistic: F1 ~ s(time) + s(f0) + s(duration) + ti(time, f0, duration) + covariates + random smooths + AR(1)

Table S2a. Main GAMM fit indices under z and speaker-centered non-z scaling

datasetscalefamilyobservationsadjusted_R2deviance_explainedn_parametric_termsn_smooth_terms
TianjinZbaseline_F1_tone29,5130.8060.8202222
TianjinZbaseline_f0_tone23,2300.7950.8192222
TianjinZmechanistic_F1_f0_duration23,2300.8320.8411024
Xi'anZbaseline_F1_tone28,5010.7660.7822222
Xi'anZbaseline_f0_tone23,0870.8210.8382222
Xi'anZmechanistic_F1_f0_duration23,0870.7670.7801024
TianjinnotZbaseline_F1_tone_notZ29,5130.8170.8302222
TianjinnotZbaseline_f0_tone_notZ23,2300.8120.8352222
TianjinnotZmechanistic_F1_f0_duration_notZ23,2300.8360.8451024
Xi'annotZbaseline_F1_tone_notZ28,5010.7740.7902222
Xi'annotZbaseline_f0_tone_notZ23,0870.8350.8522222
Xi'annotZmechanistic_F1_f0_duration_notZ23,0870.7780.7911024

Figure S2-1. Baseline F1 curves in the Tianjin-speaker dataset

Baseline F1 trajectory panels from the z-scale analysis. Solid lines show the predicted trajectories, and shaded ribbons show the corresponding confidence intervals.

Figure S2-1. Baseline F1 curves in the Tianjin-speaker dataset (`z` scale, female panels)
Figure S2-1. Baseline F1 curves in the Tianjin-speaker dataset (`z` scale, female panels)
Figure S2-1B. Baseline F1 curves in the Tianjin-speaker dataset (`z` scale, male panels)
Figure S2-1B. Baseline F1 curves in the Tianjin-speaker dataset (`z` scale, male panels)

Figure S2-2. Baseline F1 curves in the Xi'an-speaker dataset

Baseline F1 trajectory panels from the z-scale analysis. Solid lines show the predicted trajectories, and shaded ribbons show the corresponding confidence intervals.

Figure S2-2. Baseline F1 curves in the Xi'an-speaker dataset (`z` scale, female panels)
Figure S2-2. Baseline F1 curves in the Xi'an-speaker dataset (`z` scale, female panels)
Figure S2-2B. Baseline F1 curves in the Xi'an-speaker dataset (`z` scale, male panels)
Figure S2-2B. Baseline F1 curves in the Xi'an-speaker dataset (`z` scale, male panels)

Figure S2-3. Baseline f0 curves in the Tianjin-speaker dataset

Baseline f0 curves by tone for the Tianjin-speaker dataset. Solid lines show the predicted trajectories, and shaded ribbons show the corresponding confidence intervals.

Figure S2-3. Baseline `f0` curves in the Tianjin-speaker dataset (`z` scale, female panels)
Figure S2-3. Baseline `f0` curves in the Tianjin-speaker dataset (`z` scale, female panels)
Figure S2-3B. Baseline `f0` curves in the Tianjin-speaker dataset (`z` scale, male panels)
Figure S2-3B. Baseline `f0` curves in the Tianjin-speaker dataset (`z` scale, male panels)

Figure S2-4. Baseline f0 curves in the Xi'an-speaker dataset

Baseline f0 curves by tone for the Xi'an-speaker dataset. Solid lines show the predicted trajectories, and shaded ribbons show the corresponding confidence intervals.

Figure S2-4. Baseline `f0` curves in the Xi'an-speaker dataset (`z` scale, female panels)
Figure S2-4. Baseline `f0` curves in the Xi'an-speaker dataset (`z` scale, female panels)
Figure S2-4B. Baseline `f0` curves in the Xi'an-speaker dataset (`z` scale, male panels)
Figure S2-4B. Baseline `f0` curves in the Xi'an-speaker dataset (`z` scale, male panels)

Figure S2-5a. Pairwise F1 difference smooths in the Tianjin-speaker dataset

Combined pairwise F1 difference smooths across sex, variety, diphthong, and the six tone contrasts in the Tianjin-speaker dataset. Shaded regions indicate intervals of significant difference.

Figure S2-5a. Pairwise F1 difference smooths in the Tianjin-speaker dataset (`z` scale, combined panels)
Figure S2-5a. Pairwise F1 difference smooths in the Tianjin-speaker dataset (`z` scale, combined panels)

Figure S2-5b. Pairwise F1 difference smooths in the Xi'an-speaker dataset

Combined pairwise F1 difference smooths across sex, variety, diphthong, and the six tone contrasts in the Xi'an-speaker dataset. Shaded regions indicate intervals of significant difference.

Figure S2-5b. Pairwise F1 difference smooths in the Xi'an-speaker dataset (`z` scale, combined panels)
Figure S2-5b. Pairwise F1 difference smooths in the Xi'an-speaker dataset (`z` scale, combined panels)

Table S2b. Mechanistic tone-vs-no-tone comparison by compareML

The two models in Table S2b differ only in whether tone is allowed to structure the remaining time-varying F1 trajectory after the shared f0- and duration-based terms have been specified. In the no-tone model, diphthong-by-variety (DiphVar.ord) structures the factor-specific time smooth, whereas in the with-tone model this role is taken by the fuller tone-by-variety-by-diphthong factor (ToneVarDiphInt.ord). The f0- and duration-related smooth terms themselves were intentionally kept the same across the two models, because the goal of this comparison was to ask whether tone still explains residual F1 trajectory shape after controlling for shared nonlinear f0 and duration effects, rather than to build a tone-specific mechanistic mapping for those smooth terms.

ModelGroupScoreEdfDifferenceDfp.valueSig.comparisonAIC difference (with tone - without tone)
compare_A_XAXi'an20,68452without_tone
compare_B_XAXi'an19,77288911.636.000< 2e-16***with_tone-1709.80
compare_A_TJTianjin1620152without_tone
compare_B_TJTianjin1583088370.836.000< 2e-16***with_tone-743.34

Figure S2-6a. Mechanistic GAMM heatmap for /au/ in the Tianjin-speaker dataset

Mechanistic /au/ surface from the z-scale analysis.

Figure S2-6a. Mechanistic GAMM heatmap for `/au/` in the Tianjin-speaker dataset
Figure S2-6a. Mechanistic GAMM heatmap for `/au/` in the Tianjin-speaker dataset

Figure S2-6b. Mechanistic GAMM heatmap for /au/ in the Xi'an-speaker dataset

Mechanistic /au/ surface from the z-scale analysis.

Figure S2-6b. Mechanistic GAMM heatmap for `/au/` in the Xi'an-speaker dataset
Figure S2-6b. Mechanistic GAMM heatmap for `/au/` in the Xi'an-speaker dataset

Figure S2-6c. Mechanistic GAMM heatmap for /ai/ in the Tianjin-speaker dataset

Mechanistic /ai/ surface from the z-scale analysis.

Figure S2-6c. Mechanistic GAMM heatmap for `/ai/` in the Tianjin-speaker dataset
Figure S2-6c. Mechanistic GAMM heatmap for `/ai/` in the Tianjin-speaker dataset

Figure S2-6d. Mechanistic GAMM heatmap for /ai/ in the Xi'an-speaker dataset

Mechanistic /ai/ surface from the z-scale analysis.

Figure S2-6d. Mechanistic GAMM heatmap for `/ai/` in the Xi'an-speaker dataset
Figure S2-6d. Mechanistic GAMM heatmap for `/ai/` in the Xi'an-speaker dataset

Table S2c. Cascade summary comparison across scales

datasetscalebaseline_F1_R2baseline_f0_R2mechanistic_F1_R2cascade_mean_RMSEcascade_mean_nonoverlap_width
Xi'anZ0.7660.8210.7670.2420.799
TianjinZ0.8060.7950.8320.1350.194
Xi'annotZ0.7740.8350.7780.2150.801
TianjinnotZ0.8170.8120.8360.1430.092

Tianjin shows lower mean RMSE and narrower mean non-overlap than Xi'an.

Note: In the GAMM cascade summaries, predicted f0 curves were passed through the mechanistic F1 model together with duration values matched at the full condition-cell level (variety × sex × diphthong × tone × left segment × carrier-type × reading-order); error was then averaged equally across the valid matched cells within each condition.

Table S2c-1. Tone-specific GAMM cascade RMSE averaged across sex and diphthong

VarietyToneMean RMSE (z units)
Tianjin SMT10.126
Tianjin SMT20.095
Tianjin SMT30.225
Tianjin SMT40.121
TJT10.188
TJT20.086
TJT30.144
TJT40.098
Xi'an SMT10.272
Xi'an SMT20.134
Xi'an SMT30.386
Xi'an SMT40.120
XAT10.428
XAT20.281
XAT30.174
XAT40.143

RMSEs (z units) were averaged equally over realized onset, carrier-sentence, and reading-order combinations within each sex-by-diphthong condition, then equally across the two sexes and two diphthongs. Each value is a mean of condition-specific RMSEs, not an RMSE calculated from averaged trajectories.

Figure S2-7. Direct-versus-cascade trajectories for /ai/ in the Tianjin-speaker dataset (z scale)

Representative trajectory-level comparison between the direct F1 fit and the cascade prediction for /ai/. These panels are shown across sex and variety, with carrier type fixed at 1, left segment fixed at m, and reading order set to the panel-specific reference level (S for SM, T for TJ, X for XA). Colored lines distinguish the direct F1 fit from the f0-based cascade, and translucent ribbons in matching colors show their confidence intervals.

Figure S2-7. Direct-versus-cascade trajectories for `/ai/` in the Tianjin-speaker dataset (`z` scale)
Figure S2-7. Direct-versus-cascade trajectories for `/ai/` in the Tianjin-speaker dataset (`z` scale)

Figure S2-8. Direct-versus-cascade trajectories for /ai/ in the Xi'an-speaker dataset (z scale)

Representative trajectory-level comparison between the direct F1 fit and the cascade prediction for /ai/. These panels are shown across sex and variety, with carrier type fixed at 1, left segment fixed at m, and reading order set to the panel-specific reference level (S for SM, T for TJ, X for XA). Colored lines distinguish the direct F1 fit from the f0-based cascade, and translucent ribbons in matching colors show their confidence intervals.

Figure S2-8. Direct-versus-cascade trajectories for `/ai/` in the Xi'an-speaker dataset (`z` scale)
Figure S2-8. Direct-versus-cascade trajectories for `/ai/` in the Xi'an-speaker dataset (`z` scale)

Figure S2-9. Direct-versus-cascade trajectories for /au/ in the Tianjin-speaker dataset (z scale)

Representative trajectory-level comparison between the direct F1 fit and the cascade prediction for /au/. These panels are shown across sex and variety, with carrier type fixed at 1, left segment fixed at m, and reading order set to the panel-specific reference level (S for SM, T for TJ, X for XA). Colored lines distinguish the direct F1 fit from the f0-based cascade, and translucent ribbons in matching colors show their confidence intervals.

Figure S2-9. Direct-versus-cascade trajectories for `/au/` in the Tianjin-speaker dataset (`z` scale)
Figure S2-9. Direct-versus-cascade trajectories for `/au/` in the Tianjin-speaker dataset (`z` scale)

Figure S2-10. Direct-versus-cascade trajectories for /au/ in the Xi'an-speaker dataset (z scale)

Representative trajectory-level comparison between the direct F1 fit and the cascade prediction for /au/. These panels are shown across sex and variety, with carrier type fixed at 1, left segment fixed at m, and reading order set to the panel-specific reference level (S for SM, T for TJ, X for XA). Colored lines distinguish the direct F1 fit from the f0-based cascade, and translucent ribbons in matching colors show their confidence intervals.

Figure S2-10. Direct-versus-cascade trajectories for `/au/` in the Xi'an-speaker dataset (`z` scale)
Figure S2-10. Direct-versus-cascade trajectories for `/au/` in the Xi'an-speaker dataset (`z` scale)

Figure S2-11. Direct-versus-cascade divergence for /ai/ in the Tianjin-speaker dataset (z scale)

Representative band non-overlap diagnostic for /ai/ under the z-scale cascade comparison, using the same fixed reference setting as Figures S2-7 and S2-8. Shaded regions indicate non-overlap intervals.

Figure S2-11. Direct-versus-cascade divergence for `/ai/` in the Tianjin-speaker dataset (`z` scale)
Figure S2-11. Direct-versus-cascade divergence for `/ai/` in the Tianjin-speaker dataset (`z` scale)

Figure S2-12. Direct-versus-cascade divergence for /ai/ in the Xi'an-speaker dataset (z scale)

Representative band non-overlap diagnostic for /ai/ under the z-scale cascade comparison, using the same fixed reference setting as Figures S2-7 and S2-8. Shaded regions indicate non-overlap intervals.

Figure S2-12. Direct-versus-cascade divergence for `/ai/` in the Xi'an-speaker dataset (`z` scale)
Figure S2-12. Direct-versus-cascade divergence for `/ai/` in the Xi'an-speaker dataset (`z` scale)

Figure S2-13. Direct-versus-cascade divergence for /au/ in the Tianjin-speaker dataset (z scale)

Representative band non-overlap diagnostic for /au/ under the z-scale cascade comparison, using the same fixed reference setting as Figures S2-9 and S2-10. Shaded regions indicate non-overlap intervals.

Figure S2-13. Direct-versus-cascade divergence for `/au/` in the Tianjin-speaker dataset (`z` scale)
Figure S2-13. Direct-versus-cascade divergence for `/au/` in the Tianjin-speaker dataset (`z` scale)

Figure S2-14. Direct-versus-cascade divergence for /au/ in the Xi'an-speaker dataset (z scale)

Representative band non-overlap diagnostic for /au/ under the z-scale cascade comparison, using the same fixed reference setting as Figures S2-9 and S2-10. Shaded regions indicate non-overlap intervals.

Figure S2-14. Direct-versus-cascade divergence for `/au/` in the Xi'an-speaker dataset (`z` scale)
Figure S2-14. Direct-versus-cascade divergence for `/au/` in the Xi'an-speaker dataset (`z` scale)

S2b. Residual structure after the no-tone mechanistic GAMM

Residual-tone diagnostics quantify the structured variance that remains after fitting the no-tone f0-duration backbone.

Table S2d. Overall residual summary by dataset

datasetmean_rmsemean_maemean_abs_biastone_p_rmsetone_variety_p_rmsetone_p_maetone_variety_p_maehardest_conditionhardest_rmseeasiest_conditioneasiest_rmse
Xi'an0.4070.3400.233< 0.0010.014< 0.0010.013SM M aj 20.557SM M aw 30.347
Tianjin0.3490.2900.1860.2320.0010.3890.034SM F aw 10.415TJ F aj 20.293

Table S2e. Mixed-model residual ANOVA summary

datasetmetrictone_ptone_variety_p
Xi'anrmse_resid< 0.0010.014
Xi'anmae_resid< 0.0010.013
Xi'anbias_resid0.642< 0.001
Xi'anmaxabs_resid< 0.001< 0.001
Tianjinrmse_resid0.2320.001
Tianjinmae_resid0.3290.034
Tianjinbias_resid0.620< 0.001
Tianjinmaxabs_resid< 0.001< 0.001

Residual F1 variation shows tone- and variety-dependent structure, with significant tone-by-variety interactions in both datasets.

Figure S2-15. Residual curves by tone in the Xi'an-speaker dataset

Residual curves after the no-tone mechanistic GAMM.

Figure S2-15. Residual curves by tone in the Xi'an-speaker dataset
Figure S2-15. Residual curves by tone in the Xi'an-speaker dataset

Figure S2-16. Residual curves by tone in the Tianjin-speaker dataset

Residual curves after the no-tone mechanistic GAMM.

Figure S2-16. Residual curves by tone in the Tianjin-speaker dataset
Figure S2-16. Residual curves by tone in the Tianjin-speaker dataset

Figure S2-17. Residual RMSE heatmap in the Xi'an-speaker dataset

Residual RMSE by condition after the no-tone mechanistic GAMM.

Figure S2-17. Residual RMSE heatmap in the Xi'an-speaker dataset
Figure S2-17. Residual RMSE heatmap in the Xi'an-speaker dataset

Figure S2-18. Residual RMSE heatmap in the Tianjin-speaker dataset

Residual RMSE by condition after the no-tone mechanistic GAMM.

Figure S2-18. Residual RMSE heatmap in the Tianjin-speaker dataset
Figure S2-18. Residual RMSE heatmap in the Tianjin-speaker dataset

S3. Primary pooled 2 × 2 fPCA analyses under z scaling

S3a. Common functional spaces, analysis blocks, and explained variance

Trajectories from both speaker origins were pooled to estimate one f0 basis and one F1 basis for the primary /ai/ + /au/ analysis. Speaker origin (Tianjin or Xi'an) and variety status (SM or local) define four cells: Tianjin SM, Tianjin TJ, Xi'an SM, and Xi'an XA. Local status therefore denotes TJ for Tianjin speakers and XA for Xi'an speakers. The diphthong-specific reruns each re-estimate common bases across the same four cells. Origin-specific summaries below use these common bases; the separate seven-group decompositions are reported in S6.

Table S3a-1. Analysis blocks and retained sample sizes

analysis_idn_tokensn_speakersf0_pcs_retainedf1_pcs_retained
pooled_2x2_ai2,1744849
pooled_2x2_au2,58648410
pooled_2x2_ai_au_pooled4,76048410

Table S3a-2. Explained variance proportions

AnalysisDomainPC1PC2PC3PC1–PC2 cumulativePC1–PC3 cumulative
ai_au_pooledf00.6850.2760.0280.9600.989
ai_au_pooledf10.5270.2470.0840.7740.858
aif00.6960.2640.0290.9610.990
aif10.5250.2780.0790.8030.882
auf00.6780.2820.0280.9600.988
auf10.4640.2450.1040.7090.814

In the primary pooled analysis, the first two components account for 96.05% of f0 variance and 77.44% of F1 variance. PC1 primarily captures overall level and PC2 dynamic tilt; PC3 captures finer trajectory shape. Retention counts in Table S3a-1 refer to the decomposition, while the score and bridge models use the first three predictor PCs and the relative-configuration reconstruction uses PC1–PC2.

Figure S3-1. Pooled eigenfunctions (z scale)

Domain-specific eigenfunctions estimated across both origins and both variety statuses.

Figure S3-1. Pooled eigenfunctions (z scale)
Figure S3-1. Pooled eigenfunctions (z scale)

Figure S3-2. Pooled trajectory reconstructions (z scale)

Tone-wise trajectories reconstructed from mean PC scores in the four origin × variety cells. These are descriptive, unadjusted reconstructions; covariate-adjusted time shapes are shown in Figure S3-3.

Figure S3-2. Pooled trajectory reconstructions (z scale)
Figure S3-2. Pooled trajectory reconstructions (z scale)

S3b. Score-space models, duration, and condition-block comparison

The primary mixed models predict F1-PC1 or F1-PC2 from f0-PC1–PC3, standardized duration, tone × origin × variety status, diphthong, sex, reading order, onset, and carrier type, with a speaker random intercept. The duration R² comparison is a companion ordinary least-squares calculation with the same fixed predictors; its R² values are not mixed-model marginal or conditional R². A likelihood-ratio comparison was also conducted between each full mixed model and a reduced mixed model containing f0-PC1–PC3, duration, the non-tonal covariates, and the speaker random intercept, but omitting the entire tone × origin × variety-status block. Both models were fitted by maximum likelihood. Because the reduced model omits origin and variety status as well as tone and their interactions, this comparison tests the joint contribution of the complete 15-parameter condition block rather than a tone-only effect.

Table S3b-1. Incremental contribution of duration

AnalysisOutcomeR² without durationR² with durationΔR²
ai_au_pooledF1_PC10.4690.4693.31e-05
ai_au_pooledF1_PC20.2670.3320.065

Table S3b-2. f0-PC and duration coefficients in the primary mixed models

OutcomeTermEstimateSEtp
F1_PC1F0_PC10.0950.0128.123< 0.001
F1_PC1F0_PC20.0030.0180.1450.885
F1_PC1F0_PC30.0260.0350.7360.462
F1_PC1duration_z0.0100.0180.5660.571
F1_PC2F0_PC10.0002580.0090.0300.976
F1_PC2F0_PC20.0430.0133.298< 0.001
F1_PC2F0_PC30.0830.0263.1960.001
F1_PC2duration_z-0.3060.013-23.230< 0.001

Duration contributes chiefly to F1-PC2 (ΔR² = 0.06498), with negligible change for F1-PC1 (ΔR² = 0.0000331). The conditional F1-PC1 coefficient for f0-PC1 is positive in the mixed model; this conditional coefficient should be distinguished from the negative unadjusted score correlations and the inverse relative tone configurations.

Table S3b-3. Unadjusted cross-domain score correlations in the common bases

DatasetPC1 score rPC2 score rPC3 score r
pooled-0.255-0.1330.074
Tianjin-0.072-0.0420.048
Xi'an-0.404-0.1740.112

Table S3b-4. Likelihood-ratio comparison of the full and reduced mixed models

OutcomeReduced parametersFull parametersReduced AICFull AICχ²dfp
F1-PC1153016119.1214688.461460.66315< 0.001
F1-PC2153011986.4911692.92323.56215< 0.001

For both F1-PC1 and F1-PC2, the full model fitted substantially better than the corresponding reduced model. Thus, tone, origin, variety status, and their interactions jointly accounted for F1 score variation beyond the low-dimensional f0 scores, duration, and the non-tonal covariates.

Figure S3-3. Covariate-adjusted pooled PC time shapes (z scale)

Tone-wise PC1 and PC2 contributions reconstructed from mixed-model estimated marginal mean scores across the four origin × variety cells.

Figure S3-3. Covariate-adjusted pooled PC time shapes (z scale)
Figure S3-3. Covariate-adjusted pooled PC time shapes (z scale)

S3c. Origin-specific bridge reconstruction summaries

Within the common pooled bases, separate linear bridges are fitted for each origin: each F1-PC1–PC3 score is predicted from f0-PC1–PC3. R² describes token-level in-sample fit. Absolute and Euclidean discrepancies are computed between direct and predicted tone × variety mean scores and averaged equally across the eight conditions within each origin; they are not token-level prediction errors.

Table S3c-1. Origin-specific bridge fit and condition-mean discrepancy

OriginTokensBridge R² PC1Bridge R² PC2Bridge R² PC3Mean absolute gap PC1Mean absolute gap PC2Mean absolute gap PC3Mean distance PC1–PC2Mean distance PC1–PC3
Tianjin2,4030.0150.0070.0060.2430.1160.0210.2970.298
Xi'an2,3570.1740.0340.0130.4440.2610.0520.5590.567

Xi'an has higher bridge R² for PC1 and PC2 (0.174 and 0.034) than Tianjin (0.015 and 0.007), but a larger PC1–PC2 condition-mean discrepancy (0.559 versus 0.297). Predictability and discrepancy therefore describe different aspects of recoverability.

S3d. Covariate-adjusted relative tonal configuration

Separate mixed models for each domain's PC1 and PC2 estimate tone × origin × variety marginal means, adjusting for duration and the non-tonal covariates with a speaker random intercept. Reconstructed trajectories are centered across the four tones at each time point within each origin × variety cell and domain. Spearman correlation measures agreement in tone rank; Pearson correlation measures agreement in centered spacing. The summary is the mean of the 11 pointwise correlations. Negative correlations indicate inverse configurations.

Table S3d-1. Adjusted relative-configuration summary

Origin × variety cellMean rank correlationMean level correlation
Tianjin SM-0.727-0.693
Tianjin local-0.655-0.645
Xi'an SM-0.909-0.930
Xi'an local-0.855-0.910

Both Xi'an cells show stronger inverse alignment than the corresponding Tianjin cells.

Figure S3-4. Adjusted relative f0–F1 configuration (z scale)

Centered trajectories reconstructed from adjusted PC1–PC2 scores; each row identifies one origin × variety cell.

Figure S3-4. Adjusted relative f0–F1 configuration (z scale)
Figure S3-4. Adjusted relative f0–F1 configuration (z scale)

Figure S3-5. Adjusted relative similarity over time (z; /ai/ + /au/)

Correlations across the four tones at each of the 11 time points, calculated from covariate-adjusted PC1–PC2 reconstructions centered within origin × variety cell and domain. Negative values indicate inverse ordering or spacing.

Figure S3-5. Adjusted relative similarity over time (z; /ai/ + /au/)
Figure S3-5. Adjusted relative similarity over time (z; /ai/ + /au/)

S3e. Covariate-adjusted direct-versus-bridged comparisons

For this supplementary paired-score evaluation, one briged is fitted across all four cells, then compares direct and bridged token scores using method × ToneOriginVar, onset, reading order, carrier type, and diphthong, with random intercepts for speaker and token. This pooled bridge differs from the origin-specific bridges summarized in S3c. Estimated marginal means and within-condition contrasts describe the pooled bridge's remaining condition-structured discrepancies. Each condition has one direct-minus-bridged contrast; the reported p values have no adjustment across the complete collection of conditions and PCs.

Table S3e-1. Direct-minus-bridged contrasts for all three F1 PCs

F1 componentConditionDirect − bridgedSEzp
F1_PC1Tianjin 1 SM-0.1350.071-1.9000.057
F1_PC1Tianjin 1 local-0.1500.081-1.8430.065
F1_PC1Tianjin 2 SM-0.2400.081-2.9780.003
F1_PC1Tianjin 2 local-0.2440.079-3.0980.002
F1_PC1Tianjin 3 SM0.0590.0790.7430.457
F1_PC1Tianjin 3 local0.1220.0721.6870.092
F1_PC1Tianjin 4 SM0.5010.0647.832< 0.001
F1_PC1Tianjin 4 local0.3700.0635.864< 0.001
F1_PC1Xi'an 1 SM-0.2340.071-3.2760.001
F1_PC1Xi'an 1 local0.1610.0991.6230.105
F1_PC1Xi'an 2 SM0.5090.0826.218< 0.001
F1_PC1Xi'an 2 local0.1600.0831.9390.052
F1_PC1Xi'an 3 SM1.3050.07916.427< 0.001
F1_PC1Xi'an 3 local-0.9250.065-14.266< 0.001
F1_PC1Xi'an 4 SM0.3730.0645.850< 0.001
F1_PC1Xi'an 4 local-1.0950.064-17.224< 0.001
F1_PC2Tianjin 1 SM-0.1760.052-3.373< 0.001
F1_PC2Tianjin 1 local-0.0040.060-0.0610.951
F1_PC2Tianjin 2 SM-0.1850.060-3.1040.002
F1_PC2Tianjin 2 local-0.0460.058-0.8000.424
F1_PC2Tianjin 3 SM-0.3340.059-5.703< 0.001
F1_PC2Tianjin 3 local-0.5160.053-9.712< 0.001
F1_PC2Tianjin 4 SM-0.2340.047-4.962< 0.001
F1_PC2Tianjin 4 local-0.2490.047-5.359< 0.001
F1_PC2Xi'an 1 SM0.0510.0530.9640.335
F1_PC2Xi'an 1 local0.6140.0738.410< 0.001
F1_PC2Xi'an 2 SM-0.4000.060-6.621< 0.001
F1_PC2Xi'an 2 local0.0340.0610.5580.577
F1_PC2Xi'an 3 SM0.6170.05910.539< 0.001
F1_PC2Xi'an 3 local0.3770.0487.880< 0.001
F1_PC2Xi'an 4 SM0.0700.0471.4890.136
F1_PC2Xi'an 4 local0.4900.04710.437< 0.001
F1_PC3Tianjin 1 SM0.0720.0332.2050.027
F1_PC3Tianjin 1 local0.0210.0370.5540.580
F1_PC3Tianjin 2 SM0.0480.0371.3010.193
F1_PC3Tianjin 2 local0.0450.0361.2420.214
F1_PC3Tianjin 3 SM-0.0450.037-1.2260.220
F1_PC3Tianjin 3 local-0.0170.033-0.5150.607
F1_PC3Tianjin 4 SM0.0090.0290.2950.768
F1_PC3Tianjin 4 local0.0270.0290.9420.346
F1_PC3Xi'an 1 SM-0.0100.033-0.3010.764
F1_PC3Xi'an 1 local-0.1940.046-4.267< 0.001
F1_PC3Xi'an 2 SM-0.0220.038-0.5840.559
F1_PC3Xi'an 2 local-0.0310.038-0.8290.407
F1_PC3Xi'an 3 SM0.0290.0370.8040.421
F1_PC3Xi'an 3 local-0.0030.030-0.0980.922
F1_PC3Xi'an 4 SM0.0790.0292.7120.007
F1_PC3Xi'an 4 local-0.0990.029-3.402< 0.001

Figure S3-6. Pooled bridge marginal means (z scale)

Covariate-adjusted direct and bridged F1-PC1–PC3 scores with 95% confidence intervals. Bridged scores come from the all-origin pooled bridge described in S3e.

Figure S3-6. Pooled bridge marginal means (z scale)
Figure S3-6. Pooled bridge marginal means (z scale)

S3f. Split-by-diphthong pooled 2 × 2 checks

Each diphthong is analyzed in its own common f0 and F1 spaces across both origins and variety statuses. The models follow S3b–S3e, omitting diphthong as a covariate. Origin-specific bridge summaries and adjusted relative configurations are reported separately.

Table S3f-1. Incremental contribution of duration

AnalysisOutcomeR² without durationR² with durationΔR²
aiF1_PC10.4250.4260.000435
aiF1_PC20.3270.4050.078
auF1_PC10.3830.3832e-05
auF1_PC20.2600.3170.057

Table S3f-2. /ai/ origin-specific bridge summary

OriginTokensBridge R² PC1Bridge R² PC2Bridge R² PC3Mean absolute gap PC1Mean absolute gap PC2Mean absolute gap PC3Mean distance PC1–PC2Mean distance PC1–PC3
Tianjin1,1220.0140.0190.0100.2760.1120.0450.3190.323
Xi'an1,0520.2400.0560.0410.5000.3980.0560.6940.700

Table S3f-3. /ai/ adjusted relative configuration

Origin × variety cellMean rank correlationMean level correlation
Tianjin SM-0.727-0.753
Tianjin local-0.727-0.741
Xi'an SM-0.909-0.928
Xi'an local-0.855-0.890

Figure S3-7. /ai/ pooled bridge marginal means (z scale)

Direct and bridged estimated marginal means with 95% confidence intervals, using the all-origin bridge within this diphthong.

Figure S3-7. /ai/ pooled bridge marginal means (z scale)
Figure S3-7. /ai/ pooled bridge marginal means (z scale)

Figure S3-8. /ai/ adjusted relative configuration (z scale)

Adjusted and centered PC1–PC2 reconstructions across the four cells.

Figure S3-8. /ai/ adjusted relative configuration (z scale)
Figure S3-8. /ai/ adjusted relative configuration (z scale)

Figure S3-9. Adjusted relative similarity over time (z; ai)

Correlations across the four tones at each of the 11 time points, calculated from covariate-adjusted PC1–PC2 reconstructions centered within origin × variety cell and domain. Negative values indicate inverse ordering or spacing.

Figure S3-9. Adjusted relative similarity over time (z; ai)
Figure S3-9. Adjusted relative similarity over time (z; ai)

Table S3f-4. /au/ origin-specific bridge summary

OriginTokensBridge R² PC1Bridge R² PC2Bridge R² PC3Mean absolute gap PC1Mean absolute gap PC2Mean absolute gap PC3Mean distance PC1–PC2Mean distance PC1–PC3
Tianjin1,2810.0240.0040.0050.1440.1210.0310.2130.218
Xi'an1,3050.1560.0440.0040.3780.2030.0480.4730.483

Table S3f-5. /au/ adjusted relative configuration

Origin × variety cellMean rank correlationMean level correlation
Tianjin SM-0.727-0.618
Tianjin local-0.455-0.468
Xi'an SM-0.909-0.916
Xi'an local-0.855-0.909

Figure S3-10. /au/ pooled bridge marginal means (z scale)

Direct and bridged estimated marginal means with 95% confidence intervals, using the all-origin bridge within this diphthong.

Figure S3-10. /au/ pooled bridge marginal means (z scale)
Figure S3-10. /au/ pooled bridge marginal means (z scale)

Figure S3-11. /au/ adjusted relative configuration (z scale)

Adjusted and centered PC1–PC2 reconstructions across the four cells.

Figure S3-11. /au/ adjusted relative configuration (z scale)
Figure S3-11. /au/ adjusted relative configuration (z scale)

Figure S3-12. Adjusted relative similarity over time (z; au)

Correlations across the four tones at each of the 11 time points, calculated from covariate-adjusted PC1–PC2 reconstructions centered within origin × variety cell and domain. Negative values indicate inverse ordering or spacing.

Figure S3-12. Adjusted relative similarity over time (z; au)
Figure S3-12. Adjusted relative similarity over time (z; au)

The supplementary multivariate f0 + F1 analysis belongs to the separate-group checks and is reported in S6g.

S4. Speaker-centered non-z robustness analyses

The pooled 2 × 2 fPCA analysis is repeated with f0 in semitones relative to each speaker's mean and F1 in speaker-centered Bark. The primary pooled and diphthong-specific runs retain the same 4,760, 2,174, and 2,586 tokens, respectively. Models, bridge definitions, and adjusted relative-configuration calculations follow S3. Duration remains standardized. The earlier separate-group non-z analyses are retained in S6h–S6i.

S4a. Pooled non-z score, bridge, and relative-configuration summaries

Table S4a-1. Incremental contribution of duration

AnalysisOutcomeR² without durationR² with durationΔR²
ai_au_pooledF1_PC10.4710.4714.23e-06
ai_au_pooledF1_PC20.3640.4110.046

Table S4a-2. Non-z origin-specific bridge summary

OriginTokensBridge R² PC1Bridge R² PC2Bridge R² PC3Mean absolute gap PC1Mean absolute gap PC2Mean absolute gap PC3Mean distance PC1–PC2Mean distance PC1–PC3
Tianjin2,4030.0160.0070.0040.2360.1190.0210.2930.294
Xi'an2,3570.1730.0350.0050.3730.1890.0640.4550.466

Table S4a-3. Non-z adjusted relative configuration

Origin × variety cellMean rank correlationMean level correlation
Tianjin SM-0.727-0.699
Tianjin local-0.582-0.675
Xi'an SM-0.909-0.943
Xi'an local-0.855-0.913

Table S4a-4. Non-z pooled direct-minus-bridged contrasts

F1 componentConditionDirect − bridgedSEzp
F1_PC1Tianjin 1 SM-0.1070.065-1.6440.100
F1_PC1Tianjin 1 local-0.0570.074-0.7670.443
F1_PC1Tianjin 2 SM-0.1270.074-1.7170.086
F1_PC1Tianjin 2 local-0.1640.072-2.2780.023
F1_PC1Tianjin 3 SM0.0930.0731.2740.203
F1_PC1Tianjin 3 local0.1740.0662.6380.008
F1_PC1Tianjin 4 SM0.5750.0599.813< 0.001
F1_PC1Tianjin 4 local0.4040.0587.003< 0.001
F1_PC1Xi'an 1 SM-0.2700.065-4.133< 0.001
F1_PC1Xi'an 1 local0.0860.0910.9530.341
F1_PC1Xi'an 2 SM0.3890.0755.187< 0.001
F1_PC1Xi'an 2 local0.0870.0761.1550.248
F1_PC1Xi'an 3 SM0.9890.07313.612< 0.001
F1_PC1Xi'an 3 local-0.8630.059-14.537< 0.001
F1_PC1Xi'an 4 SM0.2790.0584.784< 0.001
F1_PC1Xi'an 4 local-1.0560.058-18.151< 0.001
F1_PC2Tianjin 1 SM-0.3870.057-6.848< 0.001
F1_PC2Tianjin 1 local-0.1690.065-2.5990.009
F1_PC2Tianjin 2 SM-0.3540.064-5.494< 0.001
F1_PC2Tianjin 2 local-0.2160.063-3.438< 0.001
F1_PC2Tianjin 3 SM-0.4850.063-7.648< 0.001
F1_PC2Tianjin 3 local-0.6210.058-10.794< 0.001
F1_PC2Tianjin 4 SM-0.3730.051-7.295< 0.001
F1_PC2Tianjin 4 local-0.3970.050-7.874< 0.001
F1_PC2Xi'an 1 SM0.2230.0573.901< 0.001
F1_PC2Xi'an 1 local0.7240.0799.162< 0.001
F1_PC2Xi'an 2 SM-0.0670.065-1.0200.308
F1_PC2Xi'an 2 local0.2840.0664.308< 0.001
F1_PC2Xi'an 3 SM0.8570.06313.522< 0.001
F1_PC2Xi'an 3 local0.4280.0528.272< 0.001
F1_PC2Xi'an 4 SM0.2760.0515.418< 0.001
F1_PC2Xi'an 4 local0.5020.0519.898< 0.001
F1_PC3Tianjin 1 SM0.0440.0321.3710.170
F1_PC3Tianjin 1 local-0.0100.037-0.2650.791
F1_PC3Tianjin 2 SM0.0420.0371.1310.258
F1_PC3Tianjin 2 local-0.0050.036-0.1500.881
F1_PC3Tianjin 3 SM-0.0430.036-1.1850.236
F1_PC3Tianjin 3 local-0.0090.033-0.2890.773
F1_PC3Tianjin 4 SM-0.0290.029-0.9940.320
F1_PC3Tianjin 4 local-0.0080.029-0.2830.777
F1_PC3Xi'an 1 SM0.0270.0330.8240.410
F1_PC3Xi'an 1 local-0.1600.045-3.539< 0.001
F1_PC3Xi'an 2 SM0.0450.0371.2140.225
F1_PC3Xi'an 2 local0.0150.0380.4060.684
F1_PC3Xi'an 3 SM0.0420.0361.1570.247
F1_PC3Xi'an 3 local-0.0050.030-0.1660.868
F1_PC3Xi'an 4 SM0.0990.0293.402< 0.001
F1_PC3Xi'an 4 local-0.0920.029-3.1920.001

As in S3e, the marginal-mean contrasts evaluate the all-origin pooled bridge, whereas Table S4a-2 summarizes separately fitted origin-specific bridges. Contrast p values are not adjusted across all conditions and PCs. Duration again contributes mainly to PC2 (ΔR² = 0.04640), with negligible change for PC1 (0.00000423). The origin-specific PC1–PC2 discrepancy remains smaller for Tianjin (0.293) than Xi'an (0.455), and the Xi'an relative configurations remain more strongly inverse.

S4b. Split-by-diphthong pooled non-z summaries

Table S4b-1. Incremental contribution of duration

AnalysisOutcomeR² without durationR² with durationΔR²
aiF1_PC10.4380.4390.000225
aiF1_PC20.3980.4550.058
auF1_PC10.3250.3300.004
auF1_PC20.4120.4460.033

Table S4b-2. /ai/ non-z origin-specific bridge summary

OriginTokensBridge R² PC1Bridge R² PC2Bridge R² PC3Mean absolute gap PC1Mean absolute gap PC2Mean absolute gap PC3Mean distance PC1–PC2Mean distance PC1–PC3
Tianjin1,1220.0260.0110.0120.2390.1380.0400.2980.302
Xi'an1,0520.2410.0330.0290.4330.2880.0740.5650.574

Table S4b-3. /ai/ non-z adjusted relative configuration

Origin × variety cellMean rank correlationMean level correlation
Tianjin SM-0.764-0.782
Tianjin local-0.727-0.789
Xi'an SM-0.909-0.942
Xi'an local-0.855-0.891

Table S4b-4. /au/ non-z origin-specific bridge summary

OriginTokensBridge R² PC1Bridge R² PC2Bridge R² PC3Mean absolute gap PC1Mean absolute gap PC2Mean absolute gap PC3Mean distance PC1–PC2Mean distance PC1–PC3
Tianjin1,2810.0280.0050.0010.1300.1420.0360.2140.221
Xi'an1,3050.1410.0690.0030.3000.1840.0550.3890.403

Table S4b-5. /au/ non-z adjusted relative configuration

Origin × variety cellMean rank correlationMean level correlation
Tianjin SM-0.727-0.627
Tianjin local-0.473-0.524
Xi'an SM-0.855-0.919
Xi'an local-0.836-0.916

S4c. Cross-method reconstruction summary across scales (z vs not-z)

The fPCA rows use the pooled 2 × 2 results in S3c–S3d and S4a. GAMM and FoF entries retain their respective analyses in S2 and S5. Discrepancy and RMSE values have representation-dependent units; compare the cross-origin pattern within each scale.

datasetmethodmetricZnotZ
TianjinGAMM cascademean RMSE0.1350.143
TianjinGAMM cascademean nonoverlap width0.1940.092
TianjinfPCA bridgemean Euclidean discrepancy PC1-PC20.2970.293
TianjinRelative configurationmean rank correlation (local panel)-0.655-0.582
TianjinFoFcross-validated R20.6980.699
TianjinFoFcross-validated RMSE0.5440.591
Xi'anGAMM cascademean RMSE0.2420.215
Xi'anGAMM cascademean nonoverlap width0.7990.801
Xi'anfPCA bridgemean Euclidean discrepancy PC1-PC20.5590.455
Xi'anRelative configurationmean rank correlation (local panel)-0.855-0.855
Xi'anFoFcross-validated R20.5660.592
Xi'anFoFcross-validated RMSE0.6460.568

Taken together, the non-z analyses support the same broad interpretation as the main text: the GAMM cascade asymmetry remains, the fPCA bridge asymmetry remains, the stronger Xi'an inverse relative configuration remains, and the FoF direct full-trajectory contrast is partly scale-sensitive across representations.

Figure S4-1A. Non-z cascade reconstruction for /ai/ in the Tianjin-speaker dataset

Representative non-z direct-versus-cascade trajectory comparison for /ai/. These panels are shown across sex and variety, with carrier type fixed at 1, left segment fixed at m, and reading order set to the panel-specific reference level (S for SM, T for TJ, X for XA).

Figure S4-1A. Non-`z` cascade reconstruction for `/ai/` in the Tianjin-speaker dataset
Figure S4-1A. Non-`z` cascade reconstruction for `/ai/` in the Tianjin-speaker dataset

Figure S4-1B. Non-z cascade reconstruction for /au/ in the Tianjin-speaker dataset

Representative non-z direct-versus-cascade trajectory comparison for /au/. These panels are shown across sex and variety, with carrier type fixed at 1, left segment fixed at m, and reading order set to the panel-specific reference level (S for SM, T for TJ, X for XA).

Figure S4-1B. Non-`z` cascade reconstruction for `/au/` in the Tianjin-speaker dataset
Figure S4-1B. Non-`z` cascade reconstruction for `/au/` in the Tianjin-speaker dataset

Figure S4-2A. Non-z cascade reconstruction for /ai/ in the Xi'an-speaker dataset

Representative non-z direct-versus-cascade trajectory comparison for /ai/. These panels are shown across sex and variety, with carrier type fixed at 1, left segment fixed at m, and reading order set to the panel-specific reference level (S for SM, T for TJ, X for XA).

Figure S4-2A. Non-`z` cascade reconstruction for `/ai/` in the Xi'an-speaker dataset
Figure S4-2A. Non-`z` cascade reconstruction for `/ai/` in the Xi'an-speaker dataset

Figure S4-2B. Non-z cascade reconstruction for /au/ in the Xi'an-speaker dataset

Representative non-z direct-versus-cascade trajectory comparison for /au/. These panels are shown across sex and variety, with carrier type fixed at 1, left segment fixed at m, and reading order set to the panel-specific reference level (S for SM, T for TJ, X for XA).

Figure S4-2B. Non-`z` cascade reconstruction for `/au/` in the Xi'an-speaker dataset
Figure S4-2B. Non-`z` cascade reconstruction for `/au/` in the Xi'an-speaker dataset

Figure S4-3. Pooled bridge marginal means (non-z)

Adjusted direct and bridged F1-PC scores with 95% confidence intervals from the all-origin pooled bridge.

Figure S4-3. Pooled bridge marginal means (non-z)
Figure S4-3. Pooled bridge marginal means (non-z)

Figure S4-4. Adjusted relative configuration (non-z)

Centered trajectories from covariate-adjusted PC1–PC2 marginal means.

Figure S4-4. Adjusted relative configuration (non-z)
Figure S4-4. Adjusted relative configuration (non-z)

Figure S4-5. Adjusted relative similarity over time (non-z; /ai/ + /au/)

Correlations across the four tones at each of the 11 time points, calculated from covariate-adjusted PC1–PC2 reconstructions centered within origin × variety cell and domain. Negative values indicate inverse ordering or spacing.

Figure S4-5. Adjusted relative similarity over time (non-z; /ai/ + /au/)
Figure S4-5. Adjusted relative similarity over time (non-z; /ai/ + /au/)

Figure S4-6. Pooled adjusted PC time shapes (non-z)

Tone-wise adjusted PC1 and PC2 contributions in the four origin × variety cells.

Figure S4-6. Pooled adjusted PC time shapes (non-z)
Figure S4-6. Pooled adjusted PC time shapes (non-z)

Figure S4-7. /ai/ pooled bridge marginal means (non-z)

Adjusted direct and bridged scores with 95% confidence intervals.

Figure S4-7. /ai/ pooled bridge marginal means (non-z)
Figure S4-7. /ai/ pooled bridge marginal means (non-z)

Figure S4-8. /ai/ adjusted relative configuration (non-z)

Centered reconstructions from adjusted PC1–PC2 scores.

Figure S4-8. /ai/ adjusted relative configuration (non-z)
Figure S4-8. /ai/ adjusted relative configuration (non-z)

Figure S4-9. Adjusted relative similarity over time (non-z; ai)

Correlations across the four tones at each of the 11 time points, calculated from covariate-adjusted PC1–PC2 reconstructions centered within origin × variety cell and domain. Negative values indicate inverse ordering or spacing.

Figure S4-9. Adjusted relative similarity over time (non-z; ai)
Figure S4-9. Adjusted relative similarity over time (non-z; ai)

Figure S4-10. /au/ pooled bridge marginal means (non-z)

Adjusted direct and bridged scores with 95% confidence intervals.

Figure S4-10. /au/ pooled bridge marginal means (non-z)
Figure S4-10. /au/ pooled bridge marginal means (non-z)

Figure S4-11. /au/ adjusted relative configuration (non-z)

Centered reconstructions from adjusted PC1–PC2 scores.

Figure S4-11. /au/ adjusted relative configuration (non-z)
Figure S4-11. /au/ adjusted relative configuration (non-z)

Figure S4-12. Adjusted relative similarity over time (non-z; au)

Correlations across the four tones at each of the 11 time points, calculated from covariate-adjusted PC1–PC2 reconstructions centered within origin × variety cell and domain. Negative values indicate inverse ordering or spacing.

Figure S4-12. Adjusted relative similarity over time (non-z; au)
Figure S4-12. Adjusted relative similarity over time (non-z; au)

S5. Function-on-function (FoF) mapping

The FoF analyses treat the normalized f0 contour as a functional predictor and ask whether the whole f0 trajectory predicts the whole F1 trajectory, rather than only the concurrent relation between f0(t) and F1(t) at the same time point. The target model can be written as

\[F1_i(t)=\alpha(t)+\int \beta(s,t) f0_i(s)\,ds+\gamma^\top Z_i+b_i(t)+\varepsilon_{it},\]

where s indexes time on the input f0 curve, t indexes time on the output F1 curve, Z_i contains token-level covariates, and \(\beta(s,t)\) is a smooth coefficient surface describing how input-time deviations in f0 are associated with output-time variation in F1.

In practice, each token is converted from the original long format into an 11-point f0 curve and an 11-point F1 curve. Missing internal points are linearly interpolated, and leading or trailing gaps are filled by nearest-neighbor carry-forward/carry-backward so that each retained token has a complete grid on 0:10. The f0 matrix is then centered at each input time point, so that the FoF term is interpreted as the effect of token-level deviations from the average normalized f0 contour rather than the effect of the grand mean trajectory itself. The integral is approximated by trapezoidal quadrature, so the discretized FoF contribution becomes

\[\sum_{k=0}^{10}\beta(s_k,t_j)\,f0_i(s_k)\,w_k.\]

The implementation uses the matrix covariates S, Tmat, and L. For each output row F1_i(t_j), S stores the full input-time grid 0:10, Tmat stores the current output time repeated across columns, and L stores the centered and trapezoidally weighted full f0 curve for that token. The FoF term is then fitted in mgcv::bam() as te(S, Tmat, by = L), together with parametric condition effects, smooths of output time and duration, factor-specific output-time smooths, a speaker-specific factor-smooth term, and an AR(1) residual structure. Cross-validation is speaker-blocked, so held-out speakers are never seen during training; predictions exclude the speaker-specific smooth so that evaluation reflects transferable whole-curve structure rather than speaker memorization.

The FoF coefficient surface shown later in this section is a visualization of \(\beta(s,t)\) on the discrete 11-by-11 grid. It is extracted from the full-data refit by probing the te(S,Tmat):L term one input time point at a time and then plotting the resulting grid as a heatmap. We verified this extraction in two formally equivalent ways: a predict(..., type = "terms") probe implementation and an lpmatrix implementation that isolates the same one-hot input location. The term-prediction and lpmatrix methods yielded identical coefficient estimates and standard errors at all 121 grid points in both datasets.

Table S5a. FoF overall prediction summary

datasetscalecv_RMSEcv_R2
TianjinZ0.5440.698
Xi'anZ0.6460.566
TianjinnotZ0.5910.699
Xi'annotZ0.5680.592

Table S5b-1. Diagonal concentration statistics for the coefficient surfaces

datasetscaleexact_diagonal_percentplus_minus_1_band_percentweighted_mean_distance
TianjinZ9.37022.4004.416
Xi'anZ10.81730.3243.148
TianjinnotZ9.06823.1284.277
Xi'annotZ12.07132.8172.938

Table S5b-2. Hardest and easiest FoF conditions by dataset and scale

datasetscalehardest_conditionhardest_RMSEeasiest_conditioneasiest_RMSE
TianjinZT3 aj SM M0.654T2 aj SM F0.422
Xi'anZT3 aj SM M0.860T3 aw XA F0.478
TianjinnotZT1 aw SM F0.646T4 aw SM M0.432
Xi'annotZT1 aj XA F0.625T4 aw SM M0.367

Table S5c. FoF tensor-surface compareML summary

This table compares the reduced FoF model without the tensor-product surface to the corresponding model including te(S, Tmat, by = L). In both datasets, the tensor-product term significantly improved model fit and reduced AIC, indicating that a distributed input-output surface captures structure beyond a simpler non-surface specification.

DatasetScore without teEdf without teScore with teEdf with teDifferenceDfp.valueΔAIC (with - without)
Tianjin20800.833220715.283885.5436< 2e-16-226.00
Xi'an26170.803224755.23381415.5696< 2e-16-2847.98

The main text supplements the direct full-trajectory FoF comparison with a residualized evaluation that targets tone-related structure more directly. The goal is to remove the shared diphthong trajectory within each diphthong-variety-by-sex-by-time cell before comparing observed and predicted curves.

The implementation proceeds cell by cell. Let g index a diphthong-variety-by-sex-by-time cell, and let tau index the four tones. For each curve type c in {Observed, Predicted}, a tone-equal shared baseline is computed as:

\[b_c(g) = \frac{1}{4} \sum_{\tau} \operatorname{mean}(F1_c \mid \text{tone}=\tau, g)\]

This baseline is tone-equal rather than token-count-weighted, so the common trajectory is not dominated by tone imbalance in the sample. Residualized values are then defined as:

\[r_c = F1_c - b_c(g)\]

The residualized tone-structure coefficient of determination is:

\[R^2_{\text{tone}} = 1 - \frac{\sum (r_{\text{Observed}} - r_{\text{Predicted}})^2} {\sum (r_{\text{Observed}} - \operatorname{mean}(r_{\text{Observed}}))^2}\]

This statistic asks how much of the tone-structured F1 variation is recovered after the shared diphthong trajectory has been removed. The associated \(NRMSE_{\text{tone}}\) rescales the residualized RMSE by the observed residual standard deviation:

\[NRMSE_{\text{tone}} = \frac{RMSE_{\text{tone}}}{SD(r_{\text{Observed}})}\]

A complementary tone-separation recovery analysis summarizes how much observed tonal spacing is retained in the predicted residual curves. Within each cell g, the mean residual is first computed for each tone. Pairwise tone separation is then defined as the average of the six absolute pairwise differences among those four tone means, and the pairwise recovery ratio is:

\[\mathtt{pairwise\_recovery\_ratio} = \frac{\operatorname{mean}_g\!\left(\mathtt{pairwise\_sep}_{\mathrm{Predicted}}(g)\right)} {\operatorname{mean}_g\!\left(\mathtt{pairwise\_sep}_{\mathrm{Observed}}(g)\right)}\]

An RMS-based separation ratio is also retained, using the root-mean-square of the four tone-specific residual means within each cell. Higher values in these recovery ratios indicate that predicted curves preserve more of the observed tone separation after removal of the common diphthong trajectory.

Under these residualized FoF measures, the Xi'an-speaker dataset shows stronger recovery of tone-related structure than the Tianjin-speaker dataset. The overall residualized \(R^2_{\text{tone}}\) is 0.023 in Tianjin and 0.124 in Xi'an, and the overall pairwise tone-separation recovery ratio is 0.365 in Tianjin and 0.616 in Xi'an.

Table S5d. Residualized FoF tone-structure summary (overall)

datasetnrmse_tonesd_tone_targetnrmse_toner2_tone
Tianjin26,4330.5400.5470.9880.023
Xi'an25,9270.6470.6920.9360.124

Table S5e. Residualized FoF tone-structure summary by condition

datasetconditionnrmse_tonesd_tone_targetnrmse_toner2_tone
Tianjinaj SM6,1710.5650.5740.9860.029
Tianjinaj TJ6,1710.5430.5411.005-0.010
Tianjinaw SM7,0400.5320.5400.9860.027
Tianjinaw TJ7,0510.5220.5340.9780.043
Xi'anaj SM5,9180.7040.7640.9210.151
Xi'anaj XA5,6540.6540.7000.9340.128
Xi'anaw SM7,1830.6280.6530.9620.075
Xi'anaw XA7,1720.6110.6510.9400.117

Table S5f. Residualized FoF tone-separation recovery (overall)

Here, the recovery-ratio columns indicate how fully observed tone separation is preserved (values closer to 1 are better), whereas the remaining columns report the absolute observed and predicted separation magnitudes.

datasetpairwise_recovery_ratiorms_recovery_ratiomean_obs_pairwise_sepmean_pred_pairwise_sepmean_obs_rms_spreadmean_pred_rms_spread
Tianjin0.3650.3630.1810.0660.1280.046
Xi'an0.6160.6130.4660.2870.3250.199

Table S5g. Residualized FoF tone-separation recovery by condition

As in Table S5f, values closer to 1 in the recovery-ratio columns indicate better preservation of observed tone separation within each condition, while the remaining columns report the absolute observed and predicted separation magnitudes.

datasetconditionpairwise_recovery_ratiorms_recovery_ratiomean_obs_pairwise_sepmean_pred_pairwise_sep
Tianjinaj SM0.3740.3730.2280.085
Tianjinaj TJ0.3070.3100.1900.058
Tianjinaw SM0.4030.4020.1550.062
Tianjinaw TJ0.3840.3740.1490.057
Xi'anaj SM0.5770.5710.5620.324
Xi'anaj XA0.4340.4330.5460.237
Xi'anaw SM0.8570.8510.3680.315
Xi'anaw XA0.6990.7000.3890.272

S5h. Robustness to omission of the 10Hz smoothing step

As a supplementary robustness check, we repeated the FoF analyses using the unsmoothed f0 input (f0_ori) instead of the 10Hz-smoothed f0_10 trajectory before within-speaker z normalization. Across both datasets, the main predictive results changed only minimally. In Tianjin, the cross-validated overall RMSE/R2 changed from 0.54423/0.69818 to 0.54405/0.69839; in Xi'an, they changed from 0.64581/0.56604 to 0.64571/0.56616. The residualized tone-structure and tone-separation summaries also showed only very small shifts (Tianjin R^2_tone: 0.02338 -> 0.02422; Xi'an R^2_tone: 0.12422 -> 0.12459), indicating that the main whole-curve conclusions are not materially dependent on the 10Hz smoothing step.

Differences in the estimated coefficient surfaces were most apparent near the temporal boundaries, despite the small changes in prediction metrics.

Figure S5h-1. FoF coefficient surface without the 10Hz smoothing step (Tianjin-speaker dataset)

Coefficient surface from the full-data FoF refit using unsmoothed f0_ori rather than the 10Hz-smoothed f0_10 input.

Figure S5h-1. FoF coefficient surface without the `10Hz` smoothing step (Tianjin-speaker dataset)
Figure S5h-1. FoF coefficient surface without the `10Hz` smoothing step (Tianjin-speaker dataset)

Figure S5h-2. FoF coefficient surface without the 10Hz smoothing step (Xi'an-speaker dataset)

Coefficient surface from the full-data FoF refit using unsmoothed f0_ori rather than the 10Hz-smoothed f0_10 input.

Figure S5h-2. FoF coefficient surface without the `10Hz` smoothing step (Xi'an-speaker dataset)
Figure S5h-2. FoF coefficient surface without the `10Hz` smoothing step (Xi'an-speaker dataset)

Figure S5-1. FoF cross-validated mean curves by tone (z scale, Tianjin-speaker dataset)

Observed-versus-predicted mean curves under speaker-blocked FoF cross-validation.

Figure S5-1. FoF cross-validated mean curves by tone (`z` scale, Tianjin-speaker dataset)
Figure S5-1. FoF cross-validated mean curves by tone (`z` scale, Tianjin-speaker dataset)

Figure S5-2. FoF cross-validated mean curves by tone (z scale, Xi'an-speaker dataset)

Observed-versus-predicted mean curves under speaker-blocked FoF cross-validation.

Figure S5-2. FoF cross-validated mean curves by tone (`z` scale, Xi'an-speaker dataset)
Figure S5-2. FoF cross-validated mean curves by tone (`z` scale, Xi'an-speaker dataset)

Figure S5-3. FoF coefficient surface (z scale, Tianjin-speaker dataset)

Coefficient surface extracted from the full-data FoF refit after the cross-validation procedure.

Figure S5-3. FoF coefficient surface (`z` scale, Tianjin-speaker dataset)
Figure S5-3. FoF coefficient surface (`z` scale, Tianjin-speaker dataset)

Figure S5-4. FoF coefficient surface (z scale, Xi'an-speaker dataset)

Coefficient surface extracted from the full-data FoF refit after the cross-validation procedure.

Figure S5-4. FoF coefficient surface (`z` scale, Xi'an-speaker dataset)
Figure S5-4. FoF coefficient surface (`z` scale, Xi'an-speaker dataset)

Figure S5-5. FoF cross-validated mean curves by tone (non-z, Tianjin-speaker dataset)

Observed-versus-predicted mean curves from the non-z FoF rerun.

Figure S5-5. FoF cross-validated mean curves by tone (non-`z`, Tianjin-speaker dataset)
Figure S5-5. FoF cross-validated mean curves by tone (non-`z`, Tianjin-speaker dataset)

Figure S5-6. FoF cross-validated mean curves by tone (non-z, Xi'an-speaker dataset)

Observed-versus-predicted mean curves from the non-z FoF rerun.

Figure S5-6. FoF cross-validated mean curves by tone (non-`z`, Xi'an-speaker dataset)
Figure S5-6. FoF cross-validated mean curves by tone (non-`z`, Xi'an-speaker dataset)

Figure S5-7. FoF coefficient surface (non-z, Tianjin-speaker dataset)

Coefficient surface extracted from the full-data non-z FoF refit after the cross-validation procedure.

Figure S5-7. FoF coefficient surface (non-`z`, Tianjin-speaker dataset)
Figure S5-7. FoF coefficient surface (non-`z`, Tianjin-speaker dataset)

Figure S5-8. FoF coefficient surface (non-z, Xi'an-speaker dataset)

Coefficient surface extracted from the full-data non-z FoF refit after the cross-validation procedure.

Figure S5-8. FoF coefficient surface (non-`z`, Xi'an-speaker dataset)
Figure S5-8. FoF coefficient surface (non-`z`, Xi'an-speaker dataset)

S6. Supplementary seven-group fPCA analyses and non-z reruns

This section retains the separate-group analyses: two speaker-group datasets, four origin-specific variety subgroups, and a combined SM dataset. These use separately estimated functional bases and supplement the primary common-basis pooled 2 × 2 analysis in S3. The z-scale results appear in S6a–S6g, followed by their non-z counterparts in S6h–S6i. Condition-specific bridge reconstructions from these separate-group analyses are collected in S7b.

S6a. Analysis blocks and token counts

analysis_idanalysis_labeln_tokensn_speakersgroup_var
group_tianjinGroup analysis: SM vs TJ2,40324ToneVar
group_xianGroup analysis: SM vs XA2,35724ToneVar
subgroup_tianjin_smSubgroup analysis: Tianjin SM1,20124tone
subgroup_tianjin_tjSubgroup analysis: Tianjin TJ1,20224tone
subgroup_xian_smSubgroup analysis: Xi'an SM1,19124tone
subgroup_xian_xaSubgroup analysis: Xi'an XA1,16624tone
combined_smCombined SM analysis: Tianjin SM vs Xi'an SM2,39248ToneSource

S6b. Change in model fit after adding duration to score-space models

This separate-group sensitivity table supports the qualitative observation that duration contributes mainly to the dynamic F1 dimension (PC2), with little effect on the global level dimension (PC1).

analysis_idF1_PC1_R2_without_durationF1_PC1_R2_with_durationdelta_R2_from_duration_PC1F1_PC2_R2_without_durationF1_PC2_R2_with_durationdelta_R2_from_duration_PC2
combined_sm0.4110.4140.0030.2660.3520.086
group_tianjin0.4960.4970.0010.1450.2730.129
group_xian0.4760.4770.0010.3040.3420.038
subgroup_tianjin_sm0.4940.49400.1720.2660.093
subgroup_tianjin_tj0.5100.5140.0030.1770.3140.137
subgroup_xian_sm0.4330.4340.0010.3170.3900.072
subgroup_xian_xa0.4230.4280.0050.2930.3070.014

S6c. Bridge reconstruction summaries

These summaries quantify how well low-dimensional f0 structure recovers low-dimensional F1 structure in the bridge analysis. The component-wise bridge \(R^2\) values indicate how much variance in F1-PC1 to F1-PC3 is explained by f0-PC1 to f0-PC3, while the mean absolute differences show how far the bridged scores remain from the directly observed F1 scores in each component. The Euclidean discrepancy measures summarize this direct-versus-bridged gap jointly in the PC1--PC2 and PC1--PC3 score spaces.

analysis_grouptone_factortone_levelsbridge_r2_pc1bridge_r2_pc2bridge_r2_pc3mean_abs_diff_pc1mean_abs_diff_pc2mean_abs_diff_pc3mean_bridge_euclidean_pc12mean_bridge_euclidean_pc123
Group analysis: SM vs TJToneVar80.0150.0070.0040.2500.1120.0210.2970.298
Group analysis: SM vs XAToneVar80.1730.0370.0130.4370.2710.0530.5590.567
Subgroup analysis: Tianjin SMtone40.0330.0160.0030.2350.0660.0090.2540.254
Subgroup analysis: Tianjin TJtone40.0040.0210.0050.2170.1190.0340.2860.289
Subgroup analysis: Xi'an SMtone40.1800.0300.0040.2660.2170.0440.3650.372
Subgroup analysis: Xi'an XAtone40.1790.0570.0350.2600.1730.0540.3160.321
Combined SM analysis: Tianjin SM vs Xi'an SMToneSource80.0950.0120.0020.3630.2420.0400.4720.475

S6d. Covariate-adjusted relative-configuration summary

analysis_grouppanelmean_rank_cormean_level_cor
Group analysis: SM vs TJSM-0.691-0.735
Group analysis: SM vs TJTJ-0.618-0.675
Group analysis: SM vs XASM-0.891-0.928
Group analysis: SM vs XAXA-0.836-0.915
Combined SM analysis: Tianjin SM vs Xi'an SMTianjin SM-0.727-0.725
Combined SM analysis: Tianjin SM vs Xi'an SMXi'an SM-0.891-0.935

Figure S6-1. Group-level eigenfunctions (z scale, Tianjin-speaker dataset)

Pooled eigenfunctions for the Tianjin-speaker z-scale fPCA analysis.

Figure S6-1. Group-level eigenfunctions (`z` scale, Tianjin-speaker dataset)
Figure S6-1. Group-level eigenfunctions (`z` scale, Tianjin-speaker dataset)

Figure S6-2. Group-level eigenfunctions (z scale, Xi'an-speaker dataset)

Pooled eigenfunctions for the Xi'an-speaker z-scale fPCA analysis.

Figure S6-2. Group-level eigenfunctions (`z` scale, Xi'an-speaker dataset)
Figure S6-2. Group-level eigenfunctions (`z` scale, Xi'an-speaker dataset)

Figure S6-3. Paired f0-F1 PC effects (z scale, Tianjin-speaker dataset)

Pooled paired f0-F1 PC effect plots for the Tianjin-speaker z-scale analysis.

Figure S6-3. Paired `f0`-F1 PC effects (`z` scale, Tianjin-speaker dataset)
Figure S6-3. Paired `f0`-F1 PC effects (`z` scale, Tianjin-speaker dataset)

Figure S6-4. Paired f0-F1 PC effects (z scale, Xi'an-speaker dataset)

Pooled paired f0-F1 PC effect plots for the Xi'an-speaker z-scale analysis.

Figure S6-4. Paired `f0`-F1 PC effects (`z` scale, Xi'an-speaker dataset)
Figure S6-4. Paired `f0`-F1 PC effects (`z` scale, Xi'an-speaker dataset)

Figure S6-5. Group-level adjusted PC time-shape curves (z scale, Tianjin-speaker dataset)

Covariate-adjusted mixed-model PC reconstructions for the Tianjin-speaker pooled analysis.

Figure S6-5. Group-level adjusted PC time-shape curves (`z` scale, Tianjin-speaker dataset)
Figure S6-5. Group-level adjusted PC time-shape curves (`z` scale, Tianjin-speaker dataset)

Figure S6-6. Group-level adjusted PC time-shape curves (z scale, Xi'an-speaker dataset)

Covariate-adjusted mixed-model PC reconstructions for the Xi'an-speaker pooled analysis.

Figure S6-6. Group-level adjusted PC time-shape curves (`z` scale, Xi'an-speaker dataset)
Figure S6-6. Group-level adjusted PC time-shape curves (`z` scale, Xi'an-speaker dataset)

Figure S6-7. Group-level bridge marginal means (z scale, Tianjin-speaker dataset)

Direct-versus-bridged F1-PC marginal means for the Tianjin-speaker pooled analysis.

Figure S6-7. Group-level bridge marginal means (`z` scale, Tianjin-speaker dataset)
Figure S6-7. Group-level bridge marginal means (`z` scale, Tianjin-speaker dataset)

Figure S6-8. Group-level bridge marginal means (z scale, Xi'an-speaker dataset)

Direct-versus-bridged F1-PC marginal means for the Xi'an-speaker pooled analysis.

Figure S6-8. Group-level bridge marginal means (`z` scale, Xi'an-speaker dataset)
Figure S6-8. Group-level bridge marginal means (`z` scale, Xi'an-speaker dataset)

Figure S6-9. Adjusted relative configuration (z scale, Tianjin-speaker dataset)

Model-adjusted relative f0-F1 configuration trajectories.

Figure S6-9. Adjusted relative configuration (`z` scale, Tianjin-speaker dataset)
Figure S6-9. Adjusted relative configuration (`z` scale, Tianjin-speaker dataset)

Figure S6-10. Adjusted relative configuration (z scale, Xi'an-speaker dataset)

Model-adjusted relative f0-F1 configuration trajectories.

Figure S6-10. Adjusted relative configuration (`z` scale, Xi'an-speaker dataset)
Figure S6-10. Adjusted relative configuration (`z` scale, Xi'an-speaker dataset)

Figure S6-11. Adjusted relative similarity by time (z scale, Tianjin-speaker dataset)

By-time rank and level similarity for the adjusted relative configuration analysis.

Figure S6-11. Adjusted relative similarity by time (`z` scale, Tianjin-speaker dataset)
Figure S6-11. Adjusted relative similarity by time (`z` scale, Tianjin-speaker dataset)

Figure S6-12. Adjusted relative similarity by time (z scale, Xi'an-speaker dataset)

By-time rank and level similarity for the adjusted relative configuration analysis.

Figure S6-12. Adjusted relative similarity by time (`z` scale, Xi'an-speaker dataset)
Figure S6-12. Adjusted relative similarity by time (`z` scale, Xi'an-speaker dataset)

Table S6e. Method-context sensitivity checks for pooled bridge mixed models

This table shows that the direct-versus-bridged difference varies significantly across pooled group-level conditions for all three F1 principal components, indicating that bridge reconstruction only partially preserves the observed condition-structured score pattern.

analysis_idoutcomeChisqDfp_value
group_tianjinF1_PC11,4786< 0.001
group_tianjinF1_PC22176< 0.001
group_tianjinF1_PC385.9256< 0.001
group_xianF1_PC1524.56< 0.001
group_xianF1_PC2313.56< 0.001
group_xianF1_PC374.1436< 0.001

S6f. Split-by-diphthong summaries (z scale)

Table S6f-1. Group-level split bridge summaries

dataset_groupdiphthongbridge_r2_pc1bridge_r2_pc2mean_bridge_euclidean_pc12
Tianjin-speaker datasetai0.0060.0320.316
Xi'an-speaker datasetai0.2410.0510.692
Tianjin-speaker datasetau0.0240.0040.213
Xi'an-speaker datasetau0.1560.0430.473

Table S6f-2. Group-level split relative-configuration summaries

dataset_grouppaneldiphthongmean_rank_cormean_level_cor
Tianjin-speaker datasetSMai-0.782-0.793
Tianjin-speaker datasetTJai-0.727-0.807
Xi'an-speaker datasetSMai-0.909-0.933
Xi'an-speaker datasetXAai-0.836-0.894
Tianjin-speaker datasetSMau-0.727-0.726
Tianjin-speaker datasetTJau-0.418-0.552
Xi'an-speaker datasetSMau-0.873-0.930
Xi'an-speaker datasetXAau-0.909-0.927

Figure S6-13. Split-by-diphthong bridge comparison (z scale, Tianjin-speaker /ai/)

Direct-versus-bridged F1-PC summaries from the fully split /ai/ subset of the Tianjin-speaker group.

Figure S6-13. Split-by-diphthong bridge comparison (`z` scale, Tianjin-speaker `/ai/` rerun)
Figure S6-13. Split-by-diphthong bridge comparison (`z` scale, Tianjin-speaker `/ai/` rerun)

Figure S6-14. Split-by-diphthong bridge comparison (z scale, Tianjin-speaker /au/)

Direct-versus-bridged F1-PC summaries from the fully split /au/ subset of the Tianjin-speaker group.

Figure S6-14. Split-by-diphthong bridge comparison (`z` scale, Tianjin-speaker `/au/` rerun)
Figure S6-14. Split-by-diphthong bridge comparison (`z` scale, Tianjin-speaker `/au/` rerun)

Figure S6-15. Split-by-diphthong bridge comparison (z scale, Xi'an-speaker /ai/)

Direct-versus-bridged F1-PC summaries from the fully split /ai/ subset of the Xi'an-speaker group.

Figure S6-15. Split-by-diphthong bridge comparison (`z` scale, Xi'an-speaker `/ai/` rerun)
Figure S6-15. Split-by-diphthong bridge comparison (`z` scale, Xi'an-speaker `/ai/` rerun)

Figure S6-16. Split-by-diphthong bridge comparison (z scale, Xi'an-speaker /au/)

Direct-versus-bridged F1-PC summaries from the fully split /au/ subset of the Xi'an-speaker group.

Figure S6-16. Split-by-diphthong bridge comparison (`z` scale, Xi'an-speaker `/au/` rerun)
Figure S6-16. Split-by-diphthong bridge comparison (`z` scale, Xi'an-speaker `/au/` rerun)

Figure S6-17. Split-by-diphthong relative similarity (z scale, Tianjin-speaker /ai/)

By-time spearman rank and pearson level similarity from the fully split /ai/ subset of the Tianjin-speaker group.

Figure S6-17. Split-by-diphthong relative similarity (`z` scale, Tianjin-speaker `/ai/` rerun)
Figure S6-17. Split-by-diphthong relative similarity (`z` scale, Tianjin-speaker `/ai/` rerun)

Figure S6-18. Split-by-diphthong relative similarity (z scale, Tianjin-speaker /au/)

By-time spearman rank and pearson level similarity from the fully split /au/ subset of the Tianjin-speaker group.

Figure S6-18. Split-by-diphthong relative similarity (`z` scale, Tianjin-speaker `/au/` rerun)
Figure S6-18. Split-by-diphthong relative similarity (`z` scale, Tianjin-speaker `/au/` rerun)

Figure S6-19. Split-by-diphthong relative similarity (z scale, Xi'an-speaker /ai/)

By-time spearman rank and pearson level similarity from the fully split /ai/ subset of the Xi'an-speaker group.

Figure S6-19. Split-by-diphthong relative similarity (`z` scale, Xi'an-speaker `/ai/` rerun)
Figure S6-19. Split-by-diphthong relative similarity (`z` scale, Xi'an-speaker `/ai/` rerun)

Figure S6-20. Split-by-diphthong relative similarity (z scale, Xi'an-speaker /au/)

By-time spearman rank and pearson level similarity from the fully split /au/ subset of the Xi'an-speaker group.

Figure S6-20. Split-by-diphthong relative similarity (`z` scale, Xi'an-speaker `/au/` rerun)
Figure S6-20. Split-by-diphthong relative similarity (`z` scale, Xi'an-speaker `/au/` rerun)

S6g. Multivariate f0+F1 robustness analysis

As a convergence check on the separate f0 and F1 decompositions used in the main text, a multivariate FDA was also run on paired f0 and F1 trajectories within each diphthong. This multivariate analysis uses the same 11-point normalized trajectories and the same split-by-diphthong logic as the supporting analyses above, but treats f0 and F1 as two dimensions of a single functional object rather than concatenating them or decomposing them separately.

The multivariate rerun uses f0_10Z_interp and f1Z and retains only tokens with at least five observed points in each signal before interpolation.

Table S6g-1. Split-by-diphthong explained variance in the multivariate f0+F1 analysis

dataset_groupdiphthongPC1_variancePC2_variancePC3_variancePC1_to_PC3_cumulative
Tianjin-speaker dataset/ai/0.5590.1810.1010.840
Tianjin-speaker dataset/au/0.5260.2220.0770.825
Xi'an-speaker dataset/ai/0.5020.1820.1530.837
Xi'an-speaker dataset/au/0.5030.1980.1360.838

Across all four split analyses, the first multivariate component accounts for roughly one half of the total joint f0+F1 variance, and the first three components together account for approximately 82--84% of the total variance. These results indicated that the shared f0+F1 structure is low-dimensional even when the two signals are modeled jointly.

Table S6g-2. Incremental contribution of duration in the split multivariate f0+F1 score-space models

dataset_groupdiphthongMV_PC1_R2_without_durationMV_PC1_R2_with_durationMV_PC1_delta_R2_durationMV_PC2_R2_without_durationMV_PC2_R2_with_durationMV_PC2_delta_R2_durationMV_PC3_R2_without_durationMV_PC3_R2_with_durationMV_PC3_delta_R2_duration
Tianjin-speaker dataset/ai/0.6970.6980.0010.5550.5570.0020.1200.1680.048
Tianjin-speaker dataset/au/0.6150.6180.0030.6390.6400.0010.2200.2240.004
Xi'an-speaker dataset/ai/0.7820.7850.0030.7140.71500.2270.2290.001
Xi'an-speaker dataset/au/0.7880.7890.0010.7630.76300.2220.2230.001

The duration follow-up shows a selective pattern: duration contributes very little to the dominant shared component in most cells, but can contribute more strongly to higher multivariate components (PC3), especially for the Tianjin /ai/ split. This is compatible with the main-text interpretation that the strongest shared f0-F1 structure is not reducible to duration alone.

Figure S6-21A. Multivariate f0+F1 score space for the Tianjin-speaker /ai/ split

Shared multivariate score space for the split-by-diphthong Tianjin /ai/ analysis.

Figure S6-21A. Multivariate `f0+F1` score space for the Tianjin-speaker `/ai/` split
Figure S6-21A. Multivariate `f0+F1` score space for the Tianjin-speaker `/ai/` split

Figure S6-21B. Multivariate f0+F1 score space for the Tianjin-speaker /au/ split

Shared multivariate score space for the split-by-diphthong Tianjin /au/ analysis.

Figure S6-21B. Multivariate `f0+F1` score space for the Tianjin-speaker `/au/` split
Figure S6-21B. Multivariate `f0+F1` score space for the Tianjin-speaker `/au/` split

Figure S6-21C. Multivariate f0+F1 score space for the Xi'an-speaker /ai/ split

Shared multivariate score space for the split-by-diphthong Xi'an /ai/ analysis.

Figure S6-21C. Multivariate `f0+F1` score space for the Xi'an-speaker `/ai/` split
Figure S6-21C. Multivariate `f0+F1` score space for the Xi'an-speaker `/ai/` split

Figure S6-21D. Multivariate f0+F1 score space for the Xi'an-speaker /au/ split

Shared multivariate score space for the split-by-diphthong Xi'an /au/ analysis.

Figure S6-21D. Multivariate `f0+F1` score space for the Xi'an-speaker `/au/` split
Figure S6-21D. Multivariate `f0+F1` score space for the Xi'an-speaker `/au/` split

Figure S6-22A. Multivariate f0+F1 reconstruction for the Tianjin-speaker /ai/ split

Joint reconstruction of paired f0 and F1 trajectories from the first three multivariate components in the Tianjin /ai/ split.

Figure S6-22A. Multivariate `f0+F1` reconstruction for the Tianjin-speaker `/ai/` split
Figure S6-22A. Multivariate `f0+F1` reconstruction for the Tianjin-speaker `/ai/` split

Figure S6-22B. Multivariate f0+F1 reconstruction for the Tianjin-speaker /au/ split

Joint reconstruction of paired f0 and F1 trajectories from the first three multivariate components in the Tianjin /au/ split.

Figure S6-22B. Multivariate `f0+F1` reconstruction for the Tianjin-speaker `/au/` split
Figure S6-22B. Multivariate `f0+F1` reconstruction for the Tianjin-speaker `/au/` split

Figure S6-22C. Multivariate f0+F1 reconstruction for the Xi'an-speaker /ai/ split

Joint reconstruction of paired f0 and F1 trajectories from the first three multivariate components in the Xi'an /ai/ split.

Figure S6-22C. Multivariate `f0+F1` reconstruction for the Xi'an-speaker `/ai/` split
Figure S6-22C. Multivariate `f0+F1` reconstruction for the Xi'an-speaker `/ai/` split

Figure S6-22D. Multivariate f0+F1 reconstruction for the Xi'an-speaker /au/ split

Joint reconstruction of paired f0 and F1 trajectories from the first three multivariate components in the Xi'an /au/ split.

Figure S6-22D. Multivariate `f0+F1` reconstruction for the Xi'an-speaker `/au/` split
Figure S6-22D. Multivariate `f0+F1` reconstruction for the Xi'an-speaker `/au/` split

S6h. Seven-group non-z duration, bridge, and relative-configuration summaries

Table S6h-1. Additional contribution of duration in the non-z score-space models

analysis_groupF1_PC1_R2_without_durationF1_PC1_R2_with_durationdelta_R2_from_duration_PC1F1_PC2_R2_without_durationF1_PC2_R2_with_durationdelta_R2_from_duration_PC2
combined_sm0.3800.3890.0090.4240.4720.047
group_tianjin0.5040.50400.1870.2780.091
group_xian0.4630.4640.0010.4150.4470.032
subgroup_tianjin_sm0.5270.5280.0010.1980.2500.052
subgroup_tianjin_tj0.5010.5040.0030.2140.3280.115
subgroup_xian_sm0.4120.4160.0040.4450.5000.055
subgroup_xian_xa0.4100.4150.0050.4080.4210.013

Table S6h-2. Non-z bridge reconstruction summaries

analysis_groupgrouping_factorgrouping_levelbridge_r2_pc1bridge_r2_pc2bridge_r2_pc3mean_abs_diff_pc1mean_abs_diff_pc2mean_abs_diff_pc3mean_bridge_euclidean_pc12mean_bridge_euclidean_pc123
Group analysis: SM vs TJToneVar80.0170.0070.0020.2440.1250.0170.2930.294
Group analysis: SM vs XAToneVar80.1690.0390.0040.3570.2190.0620.4550.466
Subgroup analysis: Tianjin SMtone40.0300.0220.0020.2160.0880.0090.2530.254
Subgroup analysis: Tianjin TJtone40.0130.0230.0040.2020.1380.0220.2790.280
Subgroup analysis: Xi'an SMtone40.1850.0340.0010.2050.1840.0420.2910.296
Subgroup analysis: Xi'an XAtone40.1640.0590.0140.2200.1330.0660.2640.272
Combined SM analysis: Tianjin SM vs Xi'an SMToneSource80.0850.00900.3360.3520.0350.5210.522

Table S6h-3. Non-z covariate-adjusted relative-configuration summary

analysis_grouppanelmean_rank_cormean_level_cor
Group analysis: SM vs TJSM-0.727-0.740
Group analysis: SM vs TJTJ-0.582-0.704
Group analysis: SM vs XASM-0.873-0.936
Group analysis: SM vs XAXA-0.818-0.919
Combined SM analysis: Tianjin SM vs Xi'an SMTianjin-0.727-0.739
Combined SM analysis: Tianjin SM vs Xi'an SMXi'an-0.891-0.947

Table S6h-4. Non-z method-context sensitivity checks

analysis_idoutcomeChisqDfp_value
group_tianjinF1_PC11,5296< 0.001
group_tianjinF1_PC2141.86< 0.001
group_tianjinF1_PC342.0376< 0.001
group_xianF1_PC1498.96< 0.001
group_xianF1_PC2290.56< 0.001
group_xianF1_PC363.6636< 0.001

S6i. Split-by-diphthong non-z summaries

Table S6i-1. Group-level split bridge summaries under non-z scaling

dataset_groupdiphthongbridge_r2_pc1bridge_r2_pc2mean_bridge_euclidean_pc12
Tianjin-speaker datasetai0.0030.0380.297
Xi'an-speaker datasetai0.2330.0460.562
Tianjin-speaker datasetau0.0200.0130.213
Xi'an-speaker datasetau0.1510.0500.390

Table S6i-2. Group-level split relative summaries under non-z scaling

dataset_grouppaneldiphthongmean_rank_cormean_level_cor
Tianjin-speaker datasetSMai-0.745-0.762
Tianjin-speaker datasetTJai-0.782-0.830
Xi'an-speaker datasetSMai-0.891-0.946
Xi'an-speaker datasetXAai-0.855-0.890
Tianjin-speaker datasetSMau-0.691-0.709
Tianjin-speaker datasetTJau-0.455-0.594
Xi'an-speaker datasetSMau-0.873-0.932
Xi'an-speaker datasetXAau-0.891-0.936

Figure S6-23. Group-level bridge marginal means (non-z, Tianjin-speaker dataset)

Direct-versus-bridged F1-PC marginal means from the non-z analysis.

Figure S6-23. Group-level bridge marginal means (non-`z`, Tianjin-speaker dataset)
Figure S6-23. Group-level bridge marginal means (non-`z`, Tianjin-speaker dataset)

Figure S6-24. Group-level bridge marginal means (non-z, Xi'an-speaker dataset)

Direct-versus-bridged F1-PC marginal means from the non-z analysis.

Figure S6-24. Group-level bridge marginal means (non-`z`, Xi'an-speaker dataset)
Figure S6-24. Group-level bridge marginal means (non-`z`, Xi'an-speaker dataset)

Figure S6-25. Adjusted relative configuration (non-z, Tianjin-speaker dataset)

Model-adjusted relative configuration from the non-z analysis.

Figure S6-25. Adjusted relative configuration (non-`z`, Tianjin-speaker dataset)
Figure S6-25. Adjusted relative configuration (non-`z`, Tianjin-speaker dataset)

Figure S6-26. Adjusted relative configuration (non-z, Xi'an-speaker dataset)

Model-adjusted relative configuration from the non-z analysis.

Figure S6-26. Adjusted relative configuration (non-`z`, Xi'an-speaker dataset)
Figure S6-26. Adjusted relative configuration (non-`z`, Xi'an-speaker dataset)

Figure S6-27. Adjusted relative similarity by time (non-z, Tianjin-speaker dataset)

By-time rank and level similarity from the non-z analysis.

Figure S6-27. Adjusted relative similarity by time (non-`z`, Tianjin-speaker dataset)
Figure S6-27. Adjusted relative similarity by time (non-`z`, Tianjin-speaker dataset)

Figure S6-28. Adjusted relative similarity by time (non-z, Xi'an-speaker dataset)

By-time rank and level similarity from the non-z analysis.

Figure S6-28. Adjusted relative similarity by time (non-`z`, Xi'an-speaker dataset)
Figure S6-28. Adjusted relative similarity by time (non-`z`, Xi'an-speaker dataset)

S7. Cross-method reconstructed F1 curves by condition

This section presents condition-specific F1 reconstructions from the three modelling frameworks. Each panel shows one variety × diphthong condition, with the four tones distinguished by color. The plots provide descriptive comparisons of reconstructed trajectory shapes.

S7a. GAMM cascade reconstructions

These GAMM panels use the cascade prediction only, under the reference-style setting used in the follow-up script (leftSegment = m, type = 1, with order held at the script's reference value for each variety).

Figure S7-1. Cascade reconstructed F1 curves by condition in the Tianjin-speaker dataset

Group-level cascade reconstructions from the z-scale analysis, shown by variety × diphthong condition with tones overlaid.

Figure S7-1. Cascade reconstructed F1 curves by condition in the Tianjin-speaker dataset
Figure S7-1. Cascade reconstructed F1 curves by condition in the Tianjin-speaker dataset

Figure S7-2. Cascade reconstructed F1 curves by condition in the Xi'an-speaker dataset

Group-level cascade reconstructions from the z-scale analysis, shown by variety × diphthong condition with tones overlaid.

Figure S7-2. Cascade reconstructed F1 curves by condition in the Xi'an-speaker dataset
Figure S7-2. Cascade reconstructed F1 curves by condition in the Xi'an-speaker dataset

S7b. fPCA bridge reconstructions

These figures are reconstructed from the separate-group, split-by-diphthong bridge outputs summarized in S6f; the primary pooled 2 × 2 bridge results are in S3c and S3e. Each panel corresponds to one variety × diphthong condition, so the two varieties within each group analysis are separated.

Figure S7-3. fPCA bridge reconstruction by tone

Bridge-reconstructed F1 curves for the SM /ai/ condition in the Tianjin-speaker group analysis.

Figure S7-3. fPCA bridge reconstruction by tone
Figure S7-3. fPCA bridge reconstruction by tone

Figure S7-4. fPCA bridge reconstruction by tone

Bridge-reconstructed F1 curves for the TJ /ai/ condition in the Tianjin-speaker group analysis.

Figure S7-4. fPCA bridge reconstruction by tone
Figure S7-4. fPCA bridge reconstruction by tone

Figure S7-5. fPCA bridge reconstruction by tone

Bridge-reconstructed F1 curves for the SM /au/ condition in the Tianjin-speaker group analysis.

Figure S7-5. fPCA bridge reconstruction by tone
Figure S7-5. fPCA bridge reconstruction by tone

Figure S7-6. fPCA bridge reconstruction by tone

Bridge-reconstructed F1 curves for the TJ /au/ condition in the Tianjin-speaker group analysis.

Figure S7-6. fPCA bridge reconstruction by tone
Figure S7-6. fPCA bridge reconstruction by tone

Figure S7-7. fPCA bridge reconstruction by tone

Bridge-reconstructed F1 curves for the SM /ai/ condition in the Xi'an-speaker group analysis.

Figure S7-7. fPCA bridge reconstruction by tone
Figure S7-7. fPCA bridge reconstruction by tone

Figure S7-8. fPCA bridge reconstruction by tone

Bridge-reconstructed F1 curves for the XA /ai/ condition in the Xi'an-speaker group analysis.

Figure S7-8. fPCA bridge reconstruction by tone
Figure S7-8. fPCA bridge reconstruction by tone

Figure S7-9. fPCA bridge reconstruction by tone

Bridge-reconstructed F1 curves for the SM /au/ condition in the Xi'an-speaker group analysis.

Figure S7-9. fPCA bridge reconstruction by tone
Figure S7-9. fPCA bridge reconstruction by tone

Figure S7-10. fPCA bridge reconstruction by tone

Bridge-reconstructed F1 curves for the XA /au/ condition in the Xi'an-speaker group analysis.

Figure S7-10. fPCA bridge reconstruction by tone
Figure S7-10. fPCA bridge reconstruction by tone

S7c. FoF predicted-only reconstructions

These FoF panels are based on the cross-validated predicted curves only.

Figure S7-11. FoF predicted-only reconstruction by tone

FoF predicted-only F1 curves for the SM /ai/ condition in the Tianjin-speaker dataset.

Figure S7-11. FoF predicted-only reconstruction by tone
Figure S7-11. FoF predicted-only reconstruction by tone

Figure S7-12. FoF predicted-only reconstruction by tone

FoF predicted-only F1 curves for the TJ /ai/ condition in the Tianjin-speaker dataset.

Figure S7-12. FoF predicted-only reconstruction by tone
Figure S7-12. FoF predicted-only reconstruction by tone

Figure S7-13. FoF predicted-only reconstruction by tone

FoF predicted-only F1 curves for the SM /au/ condition in the Tianjin-speaker dataset.

Figure S7-13. FoF predicted-only reconstruction by tone
Figure S7-13. FoF predicted-only reconstruction by tone

Figure S7-14. FoF predicted-only reconstruction by tone

FoF predicted-only F1 curves for the TJ /au/ condition in the Tianjin-speaker dataset.

Figure S7-14. FoF predicted-only reconstruction by tone
Figure S7-14. FoF predicted-only reconstruction by tone

Figure S7-15. FoF predicted-only reconstruction by tone

FoF predicted-only F1 curves for the SM /ai/ condition in the Xi'an-speaker dataset.

Figure S7-15. FoF predicted-only reconstruction by tone
Figure S7-15. FoF predicted-only reconstruction by tone

Figure S7-16. FoF predicted-only reconstruction by tone

FoF predicted-only F1 curves for the XA /ai/ condition in the Xi'an-speaker dataset.

Figure S7-16. FoF predicted-only reconstruction by tone
Figure S7-16. FoF predicted-only reconstruction by tone

Figure S7-17. FoF predicted-only reconstruction by tone

FoF predicted-only F1 curves for the SM /au/ condition in the Xi'an-speaker dataset.

Figure S7-17. FoF predicted-only reconstruction by tone
Figure S7-17. FoF predicted-only reconstruction by tone

Figure S7-18. FoF predicted-only reconstruction by tone

FoF predicted-only F1 curves for the XA /au/ condition in the Xi'an-speaker dataset.

Figure S7-18. FoF predicted-only reconstruction by tone
Figure S7-18. FoF predicted-only reconstruction by tone