BSETS First-Line Philosopher’s Stone Report
BSETS First-Line Philosopher’s Stone Report
Cohort And Sleep Data
- Unique scored recordings:
1753 - Unique subjects:
258 - Successfully scored by Philosopher’s Stone:
1753recordings - ID prefix counts: B=14, F=39, M=199, R=6
- Hypnogram TXT available:
1753recordings - Device summary CSV available:
1747recordings; missing:M001.6, M142.4, M191.1, R002.7, R003.6, R003.7 - Artifact probability file available:
1752recordings; missing:M136.3.1 - Sleep-period clean fraction available:
1751recordings - Sleep-period clean fraction >=20%:
1225recordings; below threshold or unavailable:528 - Sleep-period clean fraction >=35%:
822recordings - Sleep-period clean fraction >=50%:
538recordings
Method Notes
- Device summary CSV sidecars contain sleep metrics such as sleep duration, WASO, sleep efficiency, stage percentages, and record_quality_index; this report parses those values but does not compute record_quality_index.
- TXT hypnograms contain 30-second sleep stages. Cohort sleep plots use device-summary metrics when present and hypnogram-derived fallbacks for duration, sleep efficiency, and stage percentages.
- Artifact files contain 2-second channel-level artifact probabilities. Clean fraction is the selected channel’s artifact-free fraction within the hypnogram sleep period only, from first sleep epoch through the end of the last sleep epoch.
- 20/75 QC: a recording passes the preprocessing screen when at least 20% of sleep-period artifact segments have artifact probability <=25%. Primary Phi summaries keep all scored sessions; QC-sensitive summaries can use threshold subsets.
- Age-trend sensitivity fits score = intercept + slope x age, with age centered at the cohort median. The intercept is therefore the fitted score at median age, not age zero. Slopes are score units per year.
- Threshold differences compare the fitted score-age line after keeping recordings with clean fraction >=50% versus >=20%; difference uncertainty is estimated by subject-level bootstrap. The QC<20% row is exploratory and shows recordings failing the 20/75 screen.
- Reliability is repeatability from repeated nights: reliability = between-subject variance / (between-subject variance + within-subject variance). No external outcome is used.
- Reliability for averaging k nights is reliability(k) = between-subject variance / (between-subject variance + within-subject variance / k). SEM = sqrt(within-subject variance), and MDC95 = 1.96 x sqrt(2) x SEM.
- Adjacent Bland-Altman plots use consecutive recordings within each subject; delta is later night minus earlier night.
- Bland-Altman delta magnitude is summarized against the cohort score IQR; this helps judge whether adjacent-night deltas are small or large relative to the available score spread.
- Lag-dependent retest stability uses all within-subject pairs grouped by recording-order lag, not only adjacent recordings.
- Subject-centered order drift subtracts each subject’s mean score first; the plotted value is therefore the average within-person deviation at each recording order.
- PCA color association uses OLS R2 from colored variable ~ PC1 + PC2, with a permutation p-value from shuffled variable labels.
- The four Phi scores are on different model scales, so their raw means should not be interpreted as directly comparable units.
Brain Health Score Distributions
brain_health_score: mean 0.000, IQR -0.118 to 0.134, SD 0.199, range -0.769 to 0.604total_cognition_score: mean 0.478, IQR 0.349 to 0.646, SD 0.220, range -0.246 to 0.906fluid_cognition_score: mean 0.499, IQR 0.346 to 0.696, SD 0.260, range -0.424 to 1.017crystallized_cognition_score: mean -0.018, IQR -0.100 to 0.076, SD 0.099, range -0.180 to 0.270
Score Age Scatter Sensitivity
Comparison is clean fraction >=50% minus clean fraction >=20%. Slope/intercept point estimates use the fitted line shown in the plots; difference CIs and p-values use subject-level bootstrap.
brain_health_score: slope >=20% -0.0077 (-0.0085 to -0.0068), slope >=50% -0.0084 (-0.0099 to -0.0069), difference -0.0007 per year, 95% bootstrap CI -0.0019 to 0.0005, p=0.213; intercept >=20% 0.047 (0.036 to 0.059), intercept >=50% 0.025 (0.005 to 0.046), difference -0.022, 95% CI -0.036 to -0.006, p=0.007; age slope stable; score level shifts with threshold.total_cognition_score: slope >=20% -0.0142 (-0.0148 to -0.0137), slope >=50% -0.0141 (-0.0151 to -0.0130), difference 0.0002 per year, 95% bootstrap CI -0.0007 to 0.0011, p=0.678; intercept >=20% 0.562 (0.554 to 0.570), intercept >=50% 0.533 (0.519 to 0.547), difference -0.029, 95% CI -0.041 to -0.016, p=0.007; age slope stable; score level shifts with threshold.fluid_cognition_score: slope >=20% -0.0172 (-0.0178 to -0.0165), slope >=50% -0.0170 (-0.0182 to -0.0159), difference 0.0001 per year, 95% bootstrap CI -0.0008 to 0.0012, p=0.738; intercept >=20% 0.606 (0.596 to 0.615), intercept >=50% 0.573 (0.557 to 0.589), difference -0.033, 95% CI -0.045 to -0.019, p=0.007; age slope stable; score level shifts with threshold.crystallized_cognition_score: slope >=20% 0.0031 (0.0027 to 0.0035), slope >=50% 0.0030 (0.0024 to 0.0036), difference -0.0000 per year, 95% bootstrap CI -0.0006 to 0.0005, p=0.837; intercept >=20% -0.041 (-0.046 to -0.035), intercept >=50% -0.028 (-0.036 to -0.020), difference 0.012, 95% CI 0.006 to 0.021, p=0.007; age slope stable; score level shifts with threshold.
Exploratory QC<20% fitted lines:
brain_health_scoreQC<20%: n=526 recordings / 167 subjects, slope -0.0061 (-0.0070 to -0.0051), intercept 0.054 (0.039 to 0.068), r=-0.49.total_cognition_scoreQC<20%: n=526 recordings / 167 subjects, slope -0.0140 (-0.0147 to -0.0133), intercept 0.609 (0.598 to 0.619), r=-0.86.fluid_cognition_scoreQC<20%: n=526 recordings / 167 subjects, slope -0.0168 (-0.0176 to -0.0159), intercept 0.641 (0.629 to 0.654), r=-0.87.crystallized_cognition_scoreQC<20%: n=526 recordings / 167 subjects, slope 0.0034 (0.0028 to 0.0040), intercept -0.038 (-0.047 to -0.029), r=0.44.
The QC20-to-QC50 score-level shift is small and the age slopes are stable; for analyses that require artifact-QC filtering, QC20 remains a reasonable primary filter with QC50 retained as a sensitivity check for absolute score level.
Within-Subject Stability
Lag-dependent retest stability:
-
brain_health_score: lag 1 mediandelta 0.140 versus lag 7 0.144; similar median delta across displayed lags. -
total_cognition_score: lag 1 mediandelta 0.092 versus lag 7 0.128; larger median delta at longer lag. -
fluid_cognition_score: lag 1 mediandelta 0.103 versus lag 7 0.143; larger median delta at longer lag. -
crystallized_cognition_score: lag 1 mediandelta 0.032 versus lag 7 0.030; similar median delta across displayed lags.
Subject-centered order drift:
brain_health_score: mean subject-centered slope -0.0018 (-0.0060 to 0.0025) per recording; ci overlaps zero.crystallized_cognition_score: mean subject-centered slope 0.0039 (0.0009 to 0.0069) per recording; positive order drift.fluid_cognition_score: mean subject-centered slope -0.0002 (-0.0038 to 0.0034) per recording; ci overlaps zero.total_cognition_score: mean subject-centered slope 0.0001 (-0.0029 to 0.0030) per recording; ci overlaps zero.
Reliability
brain_health_score: ICC 0.513, MDC95 0.388, reliability k=7 0.881total_cognition_score: ICC 0.808, MDC95 0.268, reliability k=7 0.967fluid_cognition_score: ICC 0.817, MDC95 0.308, reliability k=7 0.969crystallized_cognition_score: ICC 0.439, MDC95 0.204, reliability k=7 0.846
Inspection Flags
Routine data availability gaps and 20/75 preprocessing attrition are tracked in QC tables, not listed here as red flags.
extreme_low_score_p01: 72extreme_high_score_p99: 72extreme_adjacent_delta_p99: 60duplicate_phi_rows: 12
Figures

Output Tables
scores_long:/home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_scores_long.csvsubject_summary:/home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_subject_summary.csvvariance_components:/home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_variance_components.csvpairwise_retest:/home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_pairwise_retest.csvlag_retest_stability:/home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_lag_retest_stability.csvsubject_centered_order_drift:/home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_subject_centered_order_drift.csvsubject_order_slope_summary:/home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_subject_order_slope_summary.csvqc_audit:/home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_qc_audit.csvdata_flow_counts:/home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_data_flow_counts.csvscore_distribution_summary:/home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_score_distribution_summary.csvscore_age_trend_sensitivity:/home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_score_age_trend_sensitivity.csvsleep_score_associations:/home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_sleep_score_associations.csvqc_score_associations:/home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_qc_score_associations.csvhead_score_correlations:/home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_head_score_correlations.csvlatent_pca_scores:/home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_latent_pca_scores.csvlatent_pca_variance:/home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_latent_pca_variance.csvlatent_pca_color_associations:/home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_latent_pca_color_associations.csvred_flags:/home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_red_flags.csvoverview_targets:/home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_prioritized_overview_targets.csvduplicate_phi_rows:/home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_duplicate_phi_rows.csvduplicate_manifest_rows:/home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_duplicate_manifest_rows.csvduplicate_artifact_qc_rows:/home/wolfgang/repos/philosophers-stone-internal/BSETS/results/first_line_analysis/bsets_first_line_duplicate_artifact_qc_rows.csv
Limitations
- F12 age remains unrecoverable. Age
22is the median-age imputation used for model covariates. - Six recordings lack the adjacent device summary CSV, so sidecar-only metrics such as WASO and record_quality_index are unavailable for those sessions; hypnogram-derived sleep duration, efficiency, and stage percentages remain available.
- M136.3.1 lacks the 2-second artifact-probability file for the scored segment, so clean-fraction and 20/75 QC status cannot be evaluated from the artifact source for that recording.
- M24.7 has an artifact-probability file, but its TXT hypnogram contains no sleep epochs, so sleep-period clean fraction cannot be computed for that recording.
- The primary BSETS Phi outputs use a wearable frontal-occipital derivation; device/domain differences from the PSG-like training setting remain a validation limitation.