BSETS First-Line Philosopher’s Stone Report

Cohort And Sleep Data

  • Unique scored recordings: 1753
  • Unique subjects: 258
  • Successfully scored by Philosopher’s Stone: 1753 recordings
  • ID prefix counts: B=14, F=39, M=199, R=6
  • Hypnogram TXT available: 1753 recordings
  • Device summary CSV available: 1747 recordings; missing: M001.6, M142.4, M191.1, R002.7, R003.6, R003.7
  • Artifact probability file available: 1752 recordings; missing: M136.3.1
  • Sleep-period clean fraction available: 1751 recordings
  • Sleep-period clean fraction >=20%: 1225 recordings; below threshold or unavailable: 528
  • Sleep-period clean fraction >=35%: 822 recordings
  • Sleep-period clean fraction >=50%: 538 recordings

Method Notes

  • Device summary CSV sidecars contain sleep metrics such as sleep duration, WASO, sleep efficiency, stage percentages, and record_quality_index; this report parses those values but does not compute record_quality_index.
  • TXT hypnograms contain 30-second sleep stages. Cohort sleep plots use device-summary metrics when present and hypnogram-derived fallbacks for duration, sleep efficiency, and stage percentages.
  • Artifact files contain 2-second channel-level artifact probabilities. Clean fraction is the selected channel’s artifact-free fraction within the hypnogram sleep period only, from first sleep epoch through the end of the last sleep epoch.
  • 20/75 QC: a recording passes the preprocessing screen when at least 20% of sleep-period artifact segments have artifact probability <=25%. Primary Phi summaries keep all scored sessions; QC-sensitive summaries can use threshold subsets.
  • Age-trend sensitivity fits score = intercept + slope x age, with age centered at the cohort median. The intercept is therefore the fitted score at median age, not age zero. Slopes are score units per year.
  • Threshold differences compare the fitted score-age line after keeping recordings with clean fraction >=50% versus >=20%; difference uncertainty is estimated by subject-level bootstrap. The QC<20% row is exploratory and shows recordings failing the 20/75 screen.
  • Reliability is repeatability from repeated nights: reliability = between-subject variance / (between-subject variance + within-subject variance). No external outcome is used.
  • Reliability for averaging k nights is reliability(k) = between-subject variance / (between-subject variance + within-subject variance / k). SEM = sqrt(within-subject variance), and MDC95 = 1.96 x sqrt(2) x SEM.
  • Adjacent Bland-Altman plots use consecutive recordings within each subject; delta is later night minus earlier night.
  • Bland-Altman delta magnitude is summarized against the cohort score IQR; this helps judge whether adjacent-night deltas are small or large relative to the available score spread.
  • Lag-dependent retest stability uses all within-subject pairs grouped by recording-order lag, not only adjacent recordings.
  • Subject-centered order drift subtracts each subject’s mean score first; the plotted value is therefore the average within-person deviation at each recording order.
  • PCA color association uses OLS R2 from colored variable ~ PC1 + PC2, with a permutation p-value from shuffled variable labels.
  • The four Phi scores are on different model scales, so their raw means should not be interpreted as directly comparable units.

Brain Health Score Distributions

  • brain_health_score: mean 0.000, IQR -0.118 to 0.134, SD 0.199, range -0.769 to 0.604
  • total_cognition_score: mean 0.478, IQR 0.349 to 0.646, SD 0.220, range -0.246 to 0.906
  • fluid_cognition_score: mean 0.499, IQR 0.346 to 0.696, SD 0.260, range -0.424 to 1.017
  • crystallized_cognition_score: mean -0.018, IQR -0.100 to 0.076, SD 0.099, range -0.180 to 0.270

Score Age Scatter Sensitivity

Comparison is clean fraction >=50% minus clean fraction >=20%. Slope/intercept point estimates use the fitted line shown in the plots; difference CIs and p-values use subject-level bootstrap.

  • brain_health_score: slope >=20% -0.0077 (-0.0085 to -0.0068), slope >=50% -0.0084 (-0.0099 to -0.0069), difference -0.0007 per year, 95% bootstrap CI -0.0019 to 0.0005, p=0.213; intercept >=20% 0.047 (0.036 to 0.059), intercept >=50% 0.025 (0.005 to 0.046), difference -0.022, 95% CI -0.036 to -0.006, p=0.007; age slope stable; score level shifts with threshold.
  • total_cognition_score: slope >=20% -0.0142 (-0.0148 to -0.0137), slope >=50% -0.0141 (-0.0151 to -0.0130), difference 0.0002 per year, 95% bootstrap CI -0.0007 to 0.0011, p=0.678; intercept >=20% 0.562 (0.554 to 0.570), intercept >=50% 0.533 (0.519 to 0.547), difference -0.029, 95% CI -0.041 to -0.016, p=0.007; age slope stable; score level shifts with threshold.
  • fluid_cognition_score: slope >=20% -0.0172 (-0.0178 to -0.0165), slope >=50% -0.0170 (-0.0182 to -0.0159), difference 0.0001 per year, 95% bootstrap CI -0.0008 to 0.0012, p=0.738; intercept >=20% 0.606 (0.596 to 0.615), intercept >=50% 0.573 (0.557 to 0.589), difference -0.033, 95% CI -0.045 to -0.019, p=0.007; age slope stable; score level shifts with threshold.
  • crystallized_cognition_score: slope >=20% 0.0031 (0.0027 to 0.0035), slope >=50% 0.0030 (0.0024 to 0.0036), difference -0.0000 per year, 95% bootstrap CI -0.0006 to 0.0005, p=0.837; intercept >=20% -0.041 (-0.046 to -0.035), intercept >=50% -0.028 (-0.036 to -0.020), difference 0.012, 95% CI 0.006 to 0.021, p=0.007; age slope stable; score level shifts with threshold.

Exploratory QC<20% fitted lines:

  • brain_health_score QC<20%: n=526 recordings / 167 subjects, slope -0.0061 (-0.0070 to -0.0051), intercept 0.054 (0.039 to 0.068), r=-0.49.
  • total_cognition_score QC<20%: n=526 recordings / 167 subjects, slope -0.0140 (-0.0147 to -0.0133), intercept 0.609 (0.598 to 0.619), r=-0.86.
  • fluid_cognition_score QC<20%: n=526 recordings / 167 subjects, slope -0.0168 (-0.0176 to -0.0159), intercept 0.641 (0.629 to 0.654), r=-0.87.
  • crystallized_cognition_score QC<20%: n=526 recordings / 167 subjects, slope 0.0034 (0.0028 to 0.0040), intercept -0.038 (-0.047 to -0.029), r=0.44.

The QC20-to-QC50 score-level shift is small and the age slopes are stable; for analyses that require artifact-QC filtering, QC20 remains a reasonable primary filter with QC50 retained as a sensitivity check for absolute score level.

Within-Subject Stability

Lag-dependent retest stability:

  • brain_health_score: lag 1 median delta 0.140 versus lag 7 0.144; similar median delta across displayed lags.
  • total_cognition_score: lag 1 median delta 0.092 versus lag 7 0.128; larger median delta at longer lag.
  • fluid_cognition_score: lag 1 median delta 0.103 versus lag 7 0.143; larger median delta at longer lag.
  • crystallized_cognition_score: lag 1 median delta 0.032 versus lag 7 0.030; similar median delta across displayed lags.

Subject-centered order drift:

  • brain_health_score: mean subject-centered slope -0.0018 (-0.0060 to 0.0025) per recording; ci overlaps zero.
  • crystallized_cognition_score: mean subject-centered slope 0.0039 (0.0009 to 0.0069) per recording; positive order drift.
  • fluid_cognition_score: mean subject-centered slope -0.0002 (-0.0038 to 0.0034) per recording; ci overlaps zero.
  • total_cognition_score: mean subject-centered slope 0.0001 (-0.0029 to 0.0030) per recording; ci overlaps zero.

Reliability

  • brain_health_score: ICC 0.513, MDC95 0.388, reliability k=7 0.881
  • total_cognition_score: ICC 0.808, MDC95 0.268, reliability k=7 0.967
  • fluid_cognition_score: ICC 0.817, MDC95 0.308, reliability k=7 0.969
  • crystallized_cognition_score: ICC 0.439, MDC95 0.204, reliability k=7 0.846

Inspection Flags

Routine data availability gaps and 20/75 preprocessing attrition are tracked in QC tables, not listed here as red flags.

  • extreme_low_score_p01: 72
  • extreme_high_score_p99: 72
  • extreme_adjacent_delta_p99: 60
  • duplicate_phi_rows: 12

Figures

data_flow cohort_overview sleep_overview sleep_age_sanity score_distributions score_age_scatter_sensitivity score_trajectories score_trajectory_extremes swimmer_plots lag_retest_stability subject_centered_order_drift bsets_first_line_adjacent_delta_distributions bsets_first_line_adjacent_bland_altman reliability sleep_association qc_association primary_score_correlation top_head_correlations latent_pca

Output Tables

  • scores_long: /home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_scores_long.csv
  • subject_summary: /home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_subject_summary.csv
  • variance_components: /home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_variance_components.csv
  • pairwise_retest: /home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_pairwise_retest.csv
  • lag_retest_stability: /home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_lag_retest_stability.csv
  • subject_centered_order_drift: /home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_subject_centered_order_drift.csv
  • subject_order_slope_summary: /home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_subject_order_slope_summary.csv
  • qc_audit: /home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_qc_audit.csv
  • data_flow_counts: /home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_data_flow_counts.csv
  • score_distribution_summary: /home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_score_distribution_summary.csv
  • score_age_trend_sensitivity: /home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_score_age_trend_sensitivity.csv
  • sleep_score_associations: /home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_sleep_score_associations.csv
  • qc_score_associations: /home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_qc_score_associations.csv
  • head_score_correlations: /home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_head_score_correlations.csv
  • latent_pca_scores: /home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_latent_pca_scores.csv
  • latent_pca_variance: /home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_latent_pca_variance.csv
  • latent_pca_color_associations: /home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_latent_pca_color_associations.csv
  • red_flags: /home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_red_flags.csv
  • overview_targets: /home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_prioritized_overview_targets.csv
  • duplicate_phi_rows: /home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_duplicate_phi_rows.csv
  • duplicate_manifest_rows: /home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_duplicate_manifest_rows.csv
  • duplicate_artifact_qc_rows: /home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_duplicate_artifact_qc_rows.csv

Limitations

  • F12 age remains unrecoverable. Age 22 is the median-age imputation used for model covariates.
  • Six recordings lack the adjacent device summary CSV, so sidecar-only metrics such as WASO and record_quality_index are unavailable for those sessions; hypnogram-derived sleep duration, efficiency, and stage percentages remain available.
  • M136.3.1 lacks the 2-second artifact-probability file for the scored segment, so clean-fraction and 20/75 QC status cannot be evaluated from the artifact source for that recording.
  • M24.7 has an artifact-probability file, but its TXT hypnogram contains no sleep epochs, so sleep-period clean fraction cannot be computed for that recording.
  • The primary BSETS Phi outputs use a wearable frontal-occipital derivation; device/domain differences from the PSG-like training setting remain a validation limitation.