Not a pitch-deck number — a tested result
Before a single proprietary iPSC sample is run, the core analytical pipeline has already been built and validated end-to-end on public data.
Co-expression clustering revealed these two anti-correlated biological blocks — the basis of the 10-gene composite score above.
All five hypotheses behind this analysis were pre-registered before the data was touched; two were honestly reported as rejected. That same standard of transparent, reproducible methodology carries forward to every proprietary dataset Engine 1 generates. View the code and methodology on GitHub →
Signal is distributed, not compressible
Contrary to the initial hypothesis, predictive performance did not concentrate in a small gene panel — it improved as more features were included, reaching an AUC of 0.704 at 5,000 genes. This demonstrates that neurodevelopmental- and psychiatric-relevant molecular signals are distributed across the transcriptome rather than compressible into a simple panel — directly validating the rationale for Engine 1's deep, multi-omics profiling approach over simpler single-marker assays.
How this compares to published benchmarks
Published academic work combining polygenic scores and family genetic risk scores for major depressive disorder reports prediction accuracy in the AUC 0.57–0.63 range, using the iPSYCH cohort. That range is directly comparable to — and in some configurations lower than — the AUC 0.654–0.772 range demonstrated in this proof-of-concept. In other words, these results are competitive with published state-of-the-art benchmarks, not merely a promising early signal.
Why genetics, genetic modifiers, and proteomics are valid risk signals
This isn't a novel claim — it's the direction the published literature already points. Click through the evidence base behind each layer of the platform.
Genome-wide association studies have linked thousands of variants to psychiatric disorders. Individually, each carries only a small effect — but aggregated into a polygenic risk score, they form a reproducible signal already used for research-grade risk stratification. Recent appraisals frame near-term clinical translation as chiefly a validation and communication challenge, not a scientific one.
Carrying a pathogenic variant is not the same as developing disease. A growing body of literature shows that genetic modifiers — secondary variants that enhance or suppress a primary mutation's effect — along with epigenetic and environmental factors, are a major reason penetrance is incomplete in neurogenetic and neurodevelopmental disorders. That gap is precisely what NeuroPenetrance is built to quantify.
No clinical guideline currently includes a blood-based test for depression — diagnosis still relies entirely on symptom interviews. Proteomic profiling of blood is the leading candidate to close that gap: a decade of studies has identified reproducible candidate protein signatures, and researchers increasingly frame large-scale validation, not discovery, as the remaining barrier to clinical use.