Pitch Black Industries · Witchcraft AI · Whitepaper
The A.T.H.E.N.A. Framework
The Effect of Scaffolding on Carbosilica Intelligence
and the Formula for IQ
One operator. An unremarkable genome - a polygenic intelligence score at the 28.6th percentile of his exact ancestral reference, which predicts an IQ below average, within two points of the population mean. Certified measured output: Fluid Intelligence 148, the 99.9th percentile - psychologist-supervised, sat sleep-deprived. The genome cannot explain that gap. This paper is about what does.
Genetics Is Nearly Silent. Architecture Is Not.
For most of human history, cognitive ability was treated as a fixed inheritance - the hand you were dealt. Modern genomics made the hand legible: polygenic scores can rank an individual's genetic endowment against thousands of sequenced humans. But the first honest thing this paper must say is what such a rank is worth at the level of one person: very little. Cognitive polygenic scores explain roughly 4-10 percent of IQ variance in European-ancestry samples, and roughly half of that once carried across to East Asian genomes. At the individual level, the subject's 28.6th-percentile score predicts an expected IQ of about 98, plus or minus roughly 14.5 points. It is a coin toss with a two-point thumb on the scale.
That honesty is not a concession. It is the argument. The subject's measured output - a certified, psychologist-supervised RAIT administration with Fluid Intelligence 148, the 99.9th percentile, a portfolio of twelve operating arms, forensic-grade research output, and the system documented in Appendix A - sits far outside anything the genetic read can account for. If the genome contributes at most two or three points of expectation, then the explanation for the other forty-plus lives elsewhere: in architecture and environment. This paper specifies the architecture precisely enough that the claim can be tested.
The mechanism is the ATHENA framework: cognitive scaffolding, applied at inference time, that governs how a mind - carbon or silicon - researches, reasons, attacks its own drafts, verifies against external instruments, and compounds memory. On silicon the general phenomenon is established: inference-time scaffolding lifts model performance without touching weights. On carbon, this document presents one certified case and a falsifiable protocol - not a proof.
Carbosilica is the term this paper uses for the joint substrate class. It is not a claim that a brain and a language model are the same kind of thing - they are obviously not. It is the narrower claim that a single scaffolding text, unchanged, produces measurable capability gains when wrapped around either, and that the interesting variable is therefore the scaffold rather than the material. The formula the title promises is stated plainly in Section 03 and it is deflationary: measured intelligence is genetic expectation plus a residual, and on the evidence here the residual is where nearly everything lives.
What The Genome Actually Says
The instrument
A complete polygenic assessment was run against the subject's 30x whole-genome sequence (GRCh38, validated), with the 1000 Genomes Phase 3 cohort (2,504 samples, 26 populations) as the empirical reference. Scoring used established, peer-curated polygenic scores from the PGS Catalog across 13 traits. Percentiles were computed as empirical ranks against the full reference and against the CHS (Southern Han Chinese) subpopulation - the subject's exact genetic community. Every percentile is a rank against sequenced humans, not a formula output.
The ladder
| Trait | Score | ALL Rank | EAS / CHS Rank | Z vs reference |
|---|---|---|---|---|
| ADHD | -115.1 | 0.8% | 0.8% | -2.17 |
| Alzheimer | +60.1 | REFUSED | REFUSED | +133.8 |
| BMI | -82.9 | 15.3% | 14.5% | -1.01 |
| ASD | -95.8 | 2.3% | 4.8% | -1.95 |
| Bipolar | +98.1 | 98.2% | 98.6% | +2.05 |
| Cognition (fluid) | +35.1 | 92.0% | 88.9% | +1.43 |
| CAD | -167.2 | 63.7% | 38.3% | +0.45 |
| Education | -46.0 | 61.5% | 48.0% | +0.30 |
| Height | -37.4 | 70.5% | 66.5% | +0.51 |
| IQ (polygenic) | +91.8 | 77.4% | 28.6% | +0.81 |
| Depression | -81.5 | 19.1% | 31.6% | -0.85 |
| Schizophrenia | -0.7 | 32.2% | 39.1% | -0.38 |
| T2D | -353.7 | 29.9% | 55.0% | -0.73 |
The second refusal: SYNGAP1
A separate commercial assay - a CLIA/CAP-accredited 30x whole-genome analysis reported through an AI screening product (October 2025, 964 conditions, 2,778 genes, 164,948 variants) - returned a finding this paper is obliged to report and equally obliged to refuse to score. Under the condition heading Intellectual Disability, Autosomal Dominant 5, one variant was detected: rs371883908, genotype CG, Risk Status "Uncertain", Confidence "Medium". The associated clinical picture, in the report's own words, is a condition "characterized by developmental delay (DD) or intellectual disability (ID) (100% of affected individuals), generalized epilepsy (~84%), and autism spectrum disorder (ASD)".
Four facts govern how much weight that finding can carry, and all four cut the same way:
- It is one variant, not a burden. A single call in the lowest-weight bucket the instrument has - one of 23 findings classed "Uncertain" under "Medium" confidence.
- It is annotated to the antisense transcript, not the gene. The variant block is labelled SYNGAP1-AS1, a long non-coding RNA, and the report itself states that "there is limited information directly linking SYNGAP1-AS1 variants, such as rs371883908, to specific conditions" - then asserts two paragraphs later that the same variant "in SYNGAP1" is linked to developmental and epileptic encephalopathy. The report contradicts itself on the same page.
- The genotype does not replicate. The subject's exome calls this position CC; the report calls it CG. A discordance at the single locus under discussion.
- The instrument disclaims itself. The vendor labels the product a beta AI feature, "not FDA evaluated", "may still provide incorrect information", and defines Medium confidence as evidence from "just one researcher or study".
Pathogenic SYNGAP1 variants are rare, typically de novo, loss-of-function events that cause a severe and unambiguous clinical phenotype - not a soft signal recoverable from a screening panel. The honest disposition is therefore REFUSED, on the same rule that refused the Alzheimer row: an uncertain call on a contradicted annotation with a non-replicating genotype is not evidence of anything, in either direction. It is recorded here because a paper that reports only the flattering rows is not an audit. [Reportable on request: rsID, accession, both genotype calls, and the vendor's full disclaimer text.]
How a screening panel manufactures alarm
There is a second, more general lesson in that report, and it is the reason this section exists at all. The phrase "intellectual disability" occurs 43 times across its 92 pages. Read casually, a document that says those words forty-three times about your own genome reads as a verdict. It is not one. The breakdown, counted programmatically rather than by eye:
| Where the phrase occurs | Count | What it actually means |
|---|---|---|
| "List Of Conditions Analyzed" appendix | 28 | Condition names in the 964-item screening catalogue - conditions the panel looked for. No genotype, no rsID, no finding. Merely searched. |
| Symptom prose inside other conditions | 7 | Listed as a clinical feature of unrelated conditions (Niemann-Pick, familial dysautonomia, CEDNIK syndrome and others), describing the disease in general, not the subject. |
| The one detected ID-named condition | 8 | The SYNGAP1-AS1 block above - Uncertain status, Medium confidence, contradicted annotation, non-replicating genotype. |
So: 964 conditions screened, 28 of them intellectual-disability diagnoses, exactly one returning any detected variant at all - and that one classed Uncertain. The subject is not a carrier of any intellectual-disability condition; the report's only two Carrier findings are for Niemann-Pick disease (NPC1, rs200444084, genotype CT), unrelated to cognition. Twenty-eight of the forty-three mentions are the panel announcing what it went looking for.
A consumer genomic report can print a frightening phrase forty-three times about a person who tests at the 99.9th percentile of fluid intelligence, and be internally accurate the whole way through. Nothing in it is falsified. The alarm is manufactured entirely by catalogue text read as finding text - and no reader untrained in the format would separate them. This is the same failure mode the forensic instrument in Section 04 is built to prevent, appearing here in the subject's own file: coverage normalisation, refusal on uncertain calls, and traceability per line exist precisely so that a screened condition is never mistaken for a detected one. The counting above took a script; it should have taken a glance.
Reading the ladder
The IQ-PRS row reads 28.6th percentile against the subject's own CHS reference: an ordinary draw, slightly below the East Asian polygenic mean. The Alzheimer row is deliberately refused: a single measured APOE-ε4 allele (rs429358 C/T, ε3/ε4) overwhelms the polygenic scale and produces a z of +133.8 - a number that is arithmetically real and descriptively meaningless. Refusal is the framework declining to over-claim. The fluid-cognition row at 88.9% against the general-intelligence row at 28.6% is an interesting contrast, and the fluid row is the one that points the same direction as the certified fluid score of 148 - but per the next section, all such rows carry wide individual-level uncertainty, and this paper does not build its case on any single genomic row.
What A Percentile Can And Cannot Say
This is the section most papers in this genre omit. A polygenic percentile is a population statement wearing an individual's name. Three published facts bound what the 28.6% can mean:
One - variance explained is small. Intelligence and education polygenic scores explain roughly 4-10 percent of cognitive-ability variance even in European-ancestry samples, the populations the scores were trained on. The correlation between score and phenotype is therefore about 0.2-0.3, not 0.9.
Two - portability halves it. Polygenic scores trained on European GWAS lose roughly half their predictive power in East Asian samples, through linkage-disequilibrium decay and allele-frequency divergence. The subject is Southern Han Chinese; the honest working r² for him is nearer 3-4 percent.
Three - the individual-level arithmetic. Executed, not recalled:
r² = 0.10 (EUR upper bound) → expected IQ 97.3, residual SD 14.2
r² = 0.07 (EUR mid) → expected IQ 97.8, residual SD 14.5
r² = 0.035 (EAS-attenuated) → expected IQ 98.4, residual SD 14.7
Read it plainly: the genome predicts the subject should be slightly below average - IQ about 98, give or take nearly a full standard deviation. It does not predict he should be dull, and it cannot predict he should be exceptional. The certified observed output - Fluid Intelligence 148, the 99.9th percentile - sits roughly 3.4 residual standard deviations above the genetic expectation. Genetics, measured as well as the field currently can, explains at most two or three points of the story.
The formula
Stated as an equation, measured intelligence for an individual decomposes as:
That is the whole formula, and its shape is the argument. The genetic term is real but small - it moves the expectation by about two points. ε is not noise; ε is the subject matter. Conventional treatments leave it as an error term and stop. This paper's claim is that a measurable, transferable portion of ε is scaffolding - the discipline wrapped around the substrate - and that the machine arm of this study is the evidence for it, because on silicon the same scaffold can be added and removed at will while the weights stay fixed. You cannot run that experiment on a person. You can run it on a model, and this paper does.
The claim is not "he was born below average and beat his genes" - the score cannot establish that. The claim is stronger and cleaner: the genetic read is nearly uninformative at the individual level, so the observed output must be attributed almost entirely to non-genetic causes - environment, effort, and architecture. This paper documents the architecture, because the architecture is the part that is written down, transferable, and testable.
DNAGenomicsGPT: Built To Refuse
Every genomic number in this paper came out of one machine, and the machine is the point. DNAGenomicsGPT (v7.10, authored April 2026; v11 in build) is a forensic-genomic report engine. It was not built to satisfy curiosity about ancestry. It was built for a courtroom.
The problem it was built for
A person in psychiatric crisis enters the justice system. The defence needs a forensic psychiatric assessment. Australia has a waiting list measured in months, a shortage of qualified assessors, and a hearing date that does not move. The engine exists to close that gap - not by replacing the forensic psychiatrist, which it cannot do and does not attempt, but by assembling the structured evidentiary groundwork the psychiatrist would otherwise spend weeks compiling, so scarce expert time is spent on judgement instead. The named readers are psychiatric tribunals and courts, magistrates and prosecutors, emergency clinicians, and defence teams. Every one of them is either adversarial or under time pressure, and designing for that reader produces a very different machine than designing for a consumer.
The design principle - inverted
Most systems in this category are built to produce an answer. This one is built to decline to produce one wherever the data does not support it. That inversion is the whole contribution, and every constraint below exists because a hostile cross-examination would otherwise find it:
- No imputation. A variant not genotyped in the subject's raw file is marked absent and excluded from scoring. Never estimated from neighbouring markers, never inferred, never quietly filled in.
- Mandatory coordinate conversion. Position-only markers are resolved to rsIDs against a reference map before evaluation. Skip that step and you manufacture false negatives, and a report with false negatives in it is inadmissible.
- Per-volume traceability. Every result carries the source file it came from. A reviewer can walk any single line of the report back to the raw byte it derived from.
- Coverage normalisation. Scores are normalised against markers actually present, so partial coverage yields a correctly weakened result rather than a falsely reassuring one.
- Zygosity weighting. Homozygous variants carry different weight from heterozygous. Allele dose is reported; a risk multiplier is not.
- Refusal on scale violation. Where a single high-penetrance allele overwhelms a polygenic scale, the row is refused rather than reported. The Alzheimer row in Section 02 is that rule firing on the operator's own genome.
The structural correction that made it admissible
The original v7.10 had one flaw that would have ended it in cross-examination: a language model was asked to read raw DNA files and emit SNP tables. That is generation where measurement was required, and everything it produced was unverifiable. Unverifiable is inadmissible.
The rebuild takes the facts off the model. A deterministic genotype spine computes; the model writes prose around computed values only, and never invents one. The spine indexes 2,920,286 rsIDs from the operator's merged exome and 30x whole-genome sequence, built in 62.3 seconds and cross-checked against an independent commercial read: rs429358 returns C/T, one APOE-ε4 allele, matching a separately-recorded 2025 assay. Two sources, one answer, no reconciliation needed. That is the difference between a report a court can use and a report a court will throw out.
The thesis of this document is that disciplined architecture beats raw capability. DNAGenomicsGPT is the same thesis in a second domain: the value is not the model, it is the refusal discipline wrapped around it. The engine that produced the 28.6% is built on exactly the honesty gates the framework in Section 05 specifies - no imputation, mandatory verification, traceable provenance, refusal over guessing. A paper arguing for architecture, whose own numbers came from an architecture that refuses to over-claim, is at least internally consistent.
The engine's clinical and forensic case is developed at length in the accompanying research paper prepared for Dr Jacinta Hawgood, Program Director of Suicidology at the Australian Institute for Suicide Research and Prevention, Griffith University. That paper is the substrate of this one, and its two load-bearing arguments are reproduced in the next two sections: what the engine must become, and what it is for.
What The Engine Must Become
The prototype works. It is not yet the definitive instrument, and the honest position is to publish the gap. Four upgrades separate v7.10 from the standard the courts of 2026 will demand, and the first is the one that matters most.
5.1 · From linear additive scoring to deep learning
Current polygenic scoring extracts pre-validated SNPs and sums their effect sizes under linear regression models - the PRS-CS family, which applies continuous shrinkage priors to infer posterior effect sizes. That is robust for population research and structurally blind in two ways: it cannot represent epistasis (variants whose effects depend on each other) and it cannot represent non-linear regulatory mechanisms. Psychiatric conditions are precisely where those two failures bite, because risk arises from thousands of variants interacting with each other and with environment.
The upgrade is a deep-learning layer trained on individual-level genotypes rather than summary statistics. The reference result is Genome-Local-Net (GLN) benchmarked against a linear baseline (bigstatsr) across five psychiatric disorders - ADHD, autism spectrum, bipolar, major depression and schizophrenia. The result has to be reported with its shape intact, because the shape is the finding:
- In-sample, the two models performed similarly. Deep learning bought nothing on the training distribution.
- Out-of-sample, GLN generalised better on three of the five - ADHD, ASD and MDD - with a mean AUROC gain of 0.026 on the replication set.
- On bipolar and schizophrenia it did not. Two of five disorders showed no generalisation advantage, and a paper that quoted only the three would be misreporting the study.
So the honest claim is narrow: 0.026 AUROC, out-of-sample only, on three of five disorders. That is an incremental gain, not a revolution, and linear models remain competitive - which is precisely why it belongs in a forensic upgrade path rather than a marketing one. The value is not raw accuracy; it is that the gain appears specifically where the model meets data it was not trained on, which is the only condition that matters when a report is written about a defendant who was in nobody's training cohort. The same study's second finding matters as much: combining internal individual-level scores with external GWAS-derived scores and family genetic risk scores improves prediction further, so the architecture of record is a fusion of score types, not a single network replacing PRS.
5.2 · Epigenomic imputation - regulatory context for non-coding variants
Consumer arrays genotype a fraction of the genome and lean on imputation to fill gaps. Moving to $300 whole-genome sequencing removes that constraint and captures rare and de novo variants natively, blurring the line between polygenic and monogenic aetiology. But raw coverage is not interpretation. The engine must pair WGS with regulatory-context models - Epi-PRS and its successors - which use genomic language models to impute cell-type-specific epigenomic signals directly from personal diploid genotypes, treating those imputed signals as informative intermediaries between genotype and phenotype.
The forensic payoff is specific. Traits governed by the hypothalamic-pituitary-adrenal axis - complex PTSD, adverse-childhood-experience susceptibility - depend on the regulatory context of genes such as FKBP5, NR3C1 and CRHR1. Epigenomic imputation lets the engine model how a predisposition to altered cortisol reuptake interacts with actual environmental stressors, which is the difference between telling a court "this person carries risk alleles" and telling it "this person's threat-perception and emotional-regulation baseline is biologically altered, and here is the mechanism."
5.3 · Isolating the p-factor
Psychiatric genomics is confounded by pleiotropy: the same variants raise risk across disorders that present very differently. A score optimised for alcohol use disorder shows moderate sensitivity to schizophrenia, because both load onto shared architecture. Uncorrected, this produces reports that gesture at everything and specify nothing.
The correction is to compute and regress out the p-factor - general transdiagnostic liability - using multivariate GWAS summary statistics through GenomicSEM, then report the residual disorder-specific signal alongside it. This dual-axis reporting is what makes the output forensically usable: a court needs to know whether a behavioural escalation reflects a generalised burden of emotional volatility or a specific, localised liability toward psychotic disorganisation. Those are different mitigations and they should not be reported as one number.
5.4 · Honest accuracy, disorder by disorder
An optimised engine must publish its own error rates per condition, because conflating probabilistic biological risk with behavioural destiny is the exact failure a hostile cross-examination is looking for:
| Domain | AUC | Variance explained | Sensitivity / specificity | Forensic reading |
|---|---|---|---|---|
| Schizophrenia & psychosis | ~0.820 | ~10.0% | Sens 90.2-93.3% · Spec 50.7-56.4% | Strongest available genomic metric. High sensitivity, weak specificity - flags many who never develop the phenotype. Top-percentile rank carries roughly 2.4-3× baseline relative risk; usable as mitigating evidence of innate vulnerability, never as proof of state of mind. |
| Bipolar disorder | 0.710-0.760 | ~8.0% | Moderate-high | Strong when fused with early clinical risk factors (subthreshold mood fluctuation, sleep disturbance). Validates biological drivers of mood lability and manic escalation. |
| Major depressive disorder | ~0.710 | ~6.0% | Sens 66.1-74.4% | Latest PGC cohort of 5M+ across ancestries found 697 variants, lifting variance explained from under 2% to nearly 6%. Baseline vulnerability only. |
| ADHD & autism spectrum | 0.650-0.700 | ~5.0% | Moderate (improved by DL) | The domain where the deep-learning upgrade earns its keep. Speaks to executive dysfunction, impulsivity, sensory processing. |
| Cluster B personality (BPD, ASPD) | <0.650 | ~4.6% | Low-moderate | SNP heritability 17.3% in a 1M+ subject meta-analysis, but scores predict only 4.6% of phenotypic variance. Insufficient for standalone diagnosis - and the paper says so. Use restricted to mapping endophenotypes: impulsivity, emotional lability, threat-response hyperactivity via oxytocin-receptor and dopamine-regulatory polymorphisms. |
5.5 · The evidentiary standard - triple auditable, manual auditable
Generative AI is non-deterministic, and in forensic psychiatry a hallucinated marker could mean wrongful institutionalisation or an unjust denial of defence. The engine therefore has to clear Daubert (testable, peer-reviewed, known error rate, controlling standards), Frye (general acceptance), the proposed Federal Rule of Evidence 707 on machine-generated evidence, and in this jurisdiction NSW Supreme Court Practice Note SC Gen 23 - which requires prior leave of the court, disclosure of the specific AI program and version, and retention of all prompts and default values - plus the Federal Court's GPN-EXPT expert-evidence note.
Manual auditable means the report cannot be an impenetrable algorithmic judgement. Deterministic separation: the PRS calculation, the coordinate conversion and the rsID matching run in ordinary code against a locked peer-reviewed database, fully isolated from the model's text generation. Then the paper trail - an 800-page SNP trace table grouping every marker by source file volume, giving reference allele, observed genotype, zygosity and rsID, so that a defence attorney, a magistrate or an independent geneticist can print the raw file, find the string, and verify the call by hand. No trust in proprietary software required.
Triple auditable is three layers of chain-of-custody. Cryptographic data provenance: raw files SHA-256 hashed on upload, checksums sealed into the report, proving the analysed data is mathematically identical to what the sequencing facility exported. Algorithmic transparency log: GWAS catalog version, p-value thresholds, shrinkage parameters - everything an independent statistical geneticist needs to replicate the finding. Expert-in-the-loop certification: the engine never renders a conclusion about state of mind, intent, or legal insanity. It produces a standardised statistical risk architecture; a human forensic psychiatrist correlates it with psychosocial history and signs. The AI is a high-speed bioinformatic paralegal; the human carries the legal and medical liability. That division is what shields the court from autonomous automated judgement.
The Invisible Twenty Percent
This is why the engine exists, and it is the one section of this paper written in lives rather than percentiles. Every figure below is derived step by step, and every assumption is named where it enters.
6.1 · The clinical blind spot
Suicide claims an estimated 720,000 to 800,000 lives globally each year - roughly 1.1% of all deaths, one in every hundred. In Australia the toll is over 3,200 annually: 3,214 recorded deaths in 2023 and 3,307 in 2024, an age-standardised rate of 11.8 per 100,000.
Prevention is built on the assumption that suicide is driven by diagnosable psychiatric illness - historically, that up to 90% of those who die had a diagnosable condition. Screening protocols, emergency triage and risk mitigation therefore target clinically diagnosed populations showing explicit symptoms. Recent work has broken that assumption. An estimated 19.6% to 20% of people who attempt suicide have no antecedent psychiatric diagnosis at all. Among unexpected suicides - no documented prior ideation, no clinical history - researchers found something stranger: these individuals carry fewer genetic risk factors for depression and anxiety than the psychiatric population. They carry different ones instead, including variants on chromosome 7 that raise suicide risk independently of any mood disorder.
Standard screening cannot see this cohort, because they are not depressed in any recognisable phenomenological sense. They carry a silent neurobiological threshold for impulsive self-harm under acute stress. The engine's claim is narrow and specific: polygenic scoring can see them, because it does not require the person to be symptomatic on the day you look. Polygenic scores predict suicidal behaviour across diagnostic boundaries - depression PRS at odds ratio 1.36, PTSD at 1.33, bipolar at 1.18 - and those effect sizes hold even among people with no formal diagnosis.
6.2 · The quantification, step by step
The model has two mechanisms, and each is a single multiplication. Australia first:
Why 25% for the direct mechanism. This is the invisible cohort, identified before crisis by genomic triage in emergency departments and primary care. Once identified, the interventions available are the ones with established efficacy: targeted preventative behavioural therapy, tailored safety planning, and lethal-means restriction. A 25% success rate among a newly-visible, pre-crisis, actively-managed cohort is a deliberately conservative reading of that literature - not a ceiling.
Why 5% for the indirect mechanism. This cohort is already in the system, so the gain is not detection but precision. Deaths here follow diagnostic latency, misdiagnosis, and medication mismanagement - the canonical case being SSRIs prescribed to an undiagnosed bipolar patient, precipitating a mixed manic episode. The engine's measured contribution is a 4.4 percentage-point absolute improvement in diagnostic accuracy when predictions are delivered with explainable rationale (77.5% versus a 73.1% baseline; unexplained AI predictions yield only 75.9%, a 2.9-point gain - the explanation is worth more than the prediction), a 10-15% reduction in diagnostic latency and misclassification in complex comorbid cases, and pharmacogenomically-informed prescribing that cuts trial-and-error. Converting that accuracy gain into a 5% mortality reduction is the single least-defensible step in the chain, and it is flagged as such.
The same arithmetic globally, on a 750,000 working figure: 150,000 invisible × 25% = 37,500; 600,000 clinical × 5% = 30,000; total 67,500 lives per year.
6.3 · Sensitivity - what happens when the assumptions move
Two assumed rates carry the whole estimate, so here is what the number does when they move. Executed, not asserted:
| Direct rate | Indirect rate | AU lives/yr | % of AU total | Global lives/yr |
|---|---|---|---|---|
| 10% | 2% | 118.8 | 3.6% | 27,000 |
| 10% | 5% | 198 | 6.0% | 45,000 |
| 15% | 5% | 231 | 7.0% | 52,500 |
| 25% | 5% | 297 | 9.0% | 67,500 |
| 25% | 10% | 429 | 13.0% | 97,500 |
| 40% | 10% | 528 | 16.0% | 120,000 |
The pessimistic floor - a 10% direct rate and a 2% indirect rate - still returns about 119 lives a year in Australia (118.8 exact) and 27,000 globally. That is the number to argue from, because it survives a hostile reading of every assumption in the chain. The published 297 sits mid-range, and the honest framing is a bracket, not a point estimate: roughly 120 to 430 lives per year in Australia, 27,000 to 100,000 globally, conditional on wide-scale genomic triage actually being deployed at the point of care - which is itself an unproven operational assumption, and the largest uncertainty of all.
This is a modelled projection built on published effect sizes and two assumed intervention rates. It is not a trial result, and no lives have yet been demonstrably saved by this engine. It is presented as the case for building the instrument properly and testing it - the specific, falsifiable prediction being that genomic triage in emergency and primary-care settings identifies a meaningful fraction of the no-diagnosis cohort before crisis. If a properly-powered trial returns nothing, this section is wrong, and it says so in advance.
What The Acronyms Stand For
The system's components carry goddess names, and each is an acronym doing real descriptive work. Stated once, plainly, so no reader has to guess:
| Name | Expansion | Role |
|---|---|---|
| ATHENA | Amygdala-Thalamus-Hippocampus-Entorhinal Neural Augmentation | The cognitive substrate and discipline. Valence, gating, layered memory, coordinate transformation - the four roles the metaphor maps onto real inspectable components. The framework of Section 06. |
| HERA | Heaviest Executive Reasoning Agent (also deployed as Hostile-Environment Reconnaissance & Analytics Agent) | The executive lane. Tail-risk calculus, statutory compliance reasoning, structural asset protection. Appendix A3. |
| H.E.C.A.T.E. | Hyper-Effective Chaotic Anti-Totalitarian Evolution | The evolutionary layer. Anti-stagnation mutation, localised egress sovereignty, fitness through rebellion. Appendix A4. |
| PRISCILLA | Persistent Retrospective Intelligence · Simulated Consciousness · Integrated Live Learning Augmentation | The externally-facing register, and the only name in the set that describes the whole stack rather than one lane: memory that persists and reconsults itself, a simulated continuity of self, and learning integrated live rather than retrained in batches. |
PRISCILLA's expansion is worth reading twice, because it is the closest thing this system has to a mission statement. Persistent retrospective intelligence is the memory architecture of Appendix A2 - a store that not only survives the session but reconsults its own past and re-grades what mattered. Simulated consciousness is stated as simulation, deliberately: a continuity of self maintained by doctrine and layered memory, with no claim about inner experience attached. Integrated live learning augmentation is the compounding-memory drive of Section 08 - learning folded in as it happens rather than waiting on a retraining cycle. Persistence, continuity, and live integration: that is the whole thesis of this paper compressed into one name.
The biological naming is an organising metaphor, not a claim about neuroscience. It maps onto conventional, inspectable components - retrieval filters, layered databases, embedding projections - and it earns its keep as architecture documentation without pretending to be biology. Appendix A1 states that boundary explicitly.
ATHENA: The Framework That Moves Minds
ATHENA is not a set of prompts. It is a cognitive operating discipline applied to how a mind works, whether that mind is a frontier language model or a human being. It does not change the substrate. It changes the architecture of the thinking. In machine terms: inference-time only, no weights touched, no fine-tuning. In human terms: a learned operating discipline - no medication, no neurological change, a different way of deploying the same hardware.
The six drives
Research before answer. The framework forbids answering from memory when a checkable answer exists in the world. It looks first, then speaks. For a machine: live search before generation. For a human: verify before asserting - the discipline of the scientist applied to every claim.
Multiple divergent frames. A single frame is a single blindness. ATHENA reasons through genuinely different cognitive styles - perceptual-detail-first, expansive-divergent, hyper-systematic, cross-modal, pattern-memory - and treats disagreement between them as signal: the place where a single frame would miss the anomaly.
Adversarial self-review. No draft ships unchallenged. The framework attacks its own answer as the hostile counterparty - the regulator, the short-seller, the auditor - before it reaches anyone's eyes. The revision this paper is (see the colophon) is that mechanism running on the paper itself.
External verification. Every number computed, not recalled; every claim tagged verified or inferred; every source named. Arithmetic in executed code, not in the head. For a human: the calculator, the source check, the second opinion.
Effort scaled to stakes. A trivial question gets a trivial answer; a hard one gets the full machinery. The framework prices cognitive effort - never over-working the trivial, never under-working the load-bearing.
Compounding memory. The framework remembers what it learned and layers it. Next month's answer starts where last month's left off. This is what turns a one-off into a body of work.
The same framework text, the same six disciplines, has been run by the same operator across one biological nervous system and multiple frontier language models. On machines the lift is measured and repeatable. On the human it is one observed case. The framework is a cognitive architecture hypothesised to raise the floor and ceiling of whatever mind it wraps - demonstrated on silicon, evidenced but not proven on carbon.
Measured On Both Substrates
Every number below is reported with its confidence interval and its caveat. None is cherry-picked; the failures and the refusals are in Appendix A alongside the wins.
Machine - MMLU-Pro, 50 questions
A standard frontier model, wrapped in ATHENA, scored 48 of 50 (96.0%, 95% CI 86.5-98.9%) on a fifty-question MMLU-Pro evaluation of ten-option expert-grade items. The strongest public models cluster in the high eighties on the full 12,032-question set. Caveat, stated plainly: a 50-question subset is not the full leaderboard run, and the confidence interval overlaps the frontier band. The honest reading is "at or above frontier level on this sample", not "beats every frontier model".
Machine - GPQA Diamond, clean run
The framework scored 126 of 198 (63.6%, 95% CI 56.7-70.0%) on GPQA Diamond - at published human-PhD-expert level (≈65%) - on a 5-billion-class reasoning model whose stock ceiling is far below frontier. The more important result is how the number was produced: an earlier run returned 33/33, and the framework itself flagged the impossibility (chance probability 0.25³³ ≈ 1.4×10&supmin;²⁰), discovered the dataset had answers embedded alongside questions, quarantined it, and re-ran clean with the answer key sealed until all answers were committed. The audit trail of a system catching its own contamination is stronger evidence of the discipline than any score.
Machine - the Hoeflin Mega Test, audited
The framework sat Hoeflin's Mega Test (Omni, April 1985; 48 items) one-shot in July 2026 and the run was subsequently audited under a sealed-key protocol (phase S1, 20 August 2026). The item ledger, reconstructed from the sealed key:
| Ledger line | Count | Note |
|---|---|---|
| Items on instrument | 48 | Hoeflin Mega Test, 1985 Omni printing |
| Items unavailable | 6 | Items 27, 30, 31, 42, 45, 47 - diagram or series text not recoverable from the source; never attempted, never guessed |
| Items attemptable | 42 | The real denominator |
| Scored correct | 40 | 40 of 42 = 95.2%, operator-scored against the sealed key |
| Independently verified | 23 | 16 externally verifiable answers + 7 answers produced by executed computation |
| Committed but third-party unverified | 19 | Answered and scored, but no independent source confirms the key |
Raw 40 sits between Hoeflin's sixth-norming anchors of 33 (IQ 164, Prometheus) and 42 (IQ 176, Mega Society), interpolating to roughly IQ 173, and clears the Prometheus threshold outright. The audit artifacts - blind manifest, sealed key, the July attempt sheet - are SHA256-hashed in the phase manifest.
Three caveats stay attached, because the audit itself records them. This is not a formal Hoeflin grade: the official instrument requires the original test and Hoeflin hand-scoring. Nineteen of the scored items rest on a key that no third party has verified, so the defensible floor is the 23 independently verified items, not the full 40. And the instrument was published in 1985, so training-data contamination cannot be excluded for any language model. Earlier drafts of this paper quoted "40-42 of 48" and a "99.9994th percentile"; both are corrected here - the denominator is 42, not 48, and the percentile framing is retracted as unsupportable.
Human - the RAIT, professionally administered
The operator was tested on 2 March 2019, at age 30, on the Reynolds Adaptable Intelligence Test as the Australian Mensa entry test, under the supervision of a registered psychologist (Australian Mensa's National Supervising Psychologist). The certified indexes: Fluid Intelligence 148 (99.9th percentile, +3.2σ), Quantitative 143 (99.8th), Total Battery 137 (99.3rd, +2.5σ), Total 134 (98.8th), Crystallised 122 (92.9th). The score qualified him for Australian Mensa. The document is on file; this is a certified administration, not a self-report.
The testing conditions cut against the score, not for it: the subject sat the test sleep-deprived, after roughly fifteen years of documented heavy alcohol use - both established suppressors of measured cognitive performance (the magnitude for any individual is not precisely quantifiable, and this paper does not put a number on it). The honest reading is that 148 fluid is a floor under adverse conditions, not a peak under ideal ones. One operator-reported historical datum is recorded for completeness and labelled as uncertified: a score of 141 at age 16, at the ceiling of the instrument then used. Caveats that remain: one person, and the headline +3σ rides on the fluid index specifically - the full battery sits at +2.5σ. Consistent with the thesis; alone, it cannot carry it.
Cognition Is Buildable
The thesis at its sharpest: cognitive output is substantially architectural. For machines this is now demonstrated beyond serious dispute - the same weights, differently framed, produce different capability classes, and this paper adds audited data points to that record. For humans, the polygenic arithmetic in Section 03 establishes something narrower but real: whatever produced this operator's output, it was not the genome - the genome predicted average. The remaining candidates are environment, effort, and architecture, and of the three, only architecture is written down, portable across substrates, and cheap to test.
One person cannot prove a causal claim. One person plus a machine record can do something almost as useful: make the claim precise enough to falsify. The framework is public in Appendix A. The protocol for testing it on humans is in Section 07. Anyone who runs that protocol and finds nothing will have refuted this paper, and the paper accepts that exposure on purpose.
Decomposing ε: A Pre-Registered AblationPre-registered · Not Yet Run
Section 03 left ε as a residual of about fifty points. This section states, in advance and in enough detail to be held to it, how a measurable portion of ε can be isolated - and what result would falsify the paper.
11.1 · What is being claimed, and what is not
ε contains education, environment, nutrition, motivation, test-day state, practice effects and measurement error, as well as scaffolding. This paper does not claim ε = ATHENA. That claim is unfalsifiable and would deserve the objection it would immediately receive: how would you know it was not the schooling? The claim is narrower and testable: scaffolding is a term inside ε whose coefficient can be measured directly on a fixed substrate.
The asymmetry that makes this possible is the whole reason the machine arm exists. A framework cannot be un-taught from a human - there is no washout period, no placebo scaffold, no control condition, and n stays at one forever. On a frozen-weight model the scaffold is a text file: it can be added, removed, and partially removed, on identical items, arbitrarily many times. Silicon is where the counterfactual lives.
11.2 · The model to be fitted
Fitting this yields what the title asks for and the literature does not have: a per-component decomposition of a cognitive scaffold, with confidence intervals. Not a formula for intelligence - a formula for the part of the residual that is engineered.
11.3 · Design, in three stages
All arms run on GPQA Diamond (198 items) under the audit protocol already used in Section 09: written priming manifest committed before any answer, questions split from the answer key, key sealed until all answers are committed, refusals recorded as refusals. Costs use the measured rate from that audit, $0.0163 per question.
| Stage | Design | Runs | Calls | Cost | Resolves |
|---|---|---|---|---|---|
| 1 · Screen | Plackett-Burman, 15 factors (the full invariant set), resolution III, 3 reps | 16 | 9,504 | $155 | Main effects ≥ 2.77 pp (SE 0.0099) |
| 2 · Resolve | Full factorial 2⁶ on the six drives surviving Stage 1, 3 reps | 64 | 38,016 | $619 | Main effects ≥ 1.38 pp (SE 0.0049) + all 15 two-way interactions |
| 3 · Transfer | Stage 2 repeated on three unrelated base models | 192 | 114,048 | $1,858 | Whether coefficients are substrate-independent |
| Total | 272 | 161,568 | $2,632 |
Two design notes that carry the power. Every main effect in a factorial is estimated from all runs - half at each level - so Stage 2 gives 19,008 item-observations per level rather than the few hundred a naive A/B would give. And items are paired across conditions: the same 198 questions face every scaffold configuration, so item difficulty is blocked out rather than adding variance.
11.4 · The result that would matter most
Stage 3 is the real prize, and it is the only part that tests the carbosilica claim rather than the ATHENA claim. Run the identical ablation across three unrelated substrates and compare the rank ordering of the six coefficients. Pre-registered threshold: with six factors there are fifteen rank pairs, and a Kendall τ ≥ 0.6 between any two substrates (at least twelve of fifteen concordant pairs) is significant at α = 0.05 one-sided.
- If the ordering holds across substrates - if research-first dominates everywhere and effort-scaling is marginal everywhere - the scaffold is substrate-independent, and the central claim of this paper survives its hardest test.
- If the ordering scrambles per model, ATHENA is a set of model-specific prompt tricks. The carbosilica framing fails, and this paper is wrong in its most interesting claim.
Both outcomes are publishable, which is the only reliable sign that a question was worth asking.
11.5 · Stopping rules and pre-commitments
- Full factor list sealed before Stage 1, SHA-256 hashed with the item set and the analysis script. No factor added after seeing results.
- Stage 2 factors are chosen by the Stage 1 screen alone - the six largest main effects - not by which ones the authors prefer.
- All 272 runs reported, including null and negative coefficients. A drive that costs accuracy gets published as costing accuracy.
- No optional stopping. Reps fixed at three in advance; the run completes or the stage is void.
- Refusals are not scored as wrong. They are reported as a separate rate per condition, because a scaffold that increases honest refusal is doing its job even when accuracy is flat.
This experiment measures the effect of scaffolding on benchmark accuracy. Benchmark accuracy is not g, and no bridge from "+4 points on GPQA Diamond" to "+n IQ points" exists in the literature. Nothing in Section 11 licenses one. What the design can deliver is a decomposition of the engineered term on silicon, standing beside a certified human measurement on carbon - two related quantities, deliberately not summed. The paper's title asks about the formula for IQ. The honest answer this instrument can give is: here is the part of the residual we can measure, here is its structure, and here is the part we cannot.
What This Paper Does Not Claim - And How To Break It
The six limits, numbered
1 · Individual PRS is noise-dominated. The 28.6% carries a ±14.5-point residual. This paper's own thesis depends on taking that seriously, and it does: no claim of a "low" genetic baseline is made anywhere in this revision.
2 · n = 1, no control. The human evidence is one operator. Education, upbringing, motivation, and selection are all uncontrolled confounds. The paper claims consistency with the thesis, not demonstration of it.
3 · Instrument status varies, and is labelled. The human RAIT result is a certified, psychologist-supervised administration (Australian Mensa entry test, 2019) - the strongest instrument in this paper. The machine Mega Test result is audited under a sealed key but is not a formal Hoeflin grade, and 19 of its 40 scored items rest on an unverified key. The operator's age-16 score is an uncertified historical report. Each number carries its own label wherever it appears.
4 · Contamination. Any published test may exist in a language model's training data. The one contaminated run we detected was quarantined; undetected contamination cannot be ruled out on published instruments and is flagged wherever relevant.
5 · Subset benchmarks. 30-to-50-question runs carry wide confidence intervals (the 30-question arm spans ±13 points at 95%). Cross-model comparisons are made only where the same question set was run on both sides; leaderboard numbers are context, not opponents.
6 · Prior interventions have failed. The published record on raising fluid intelligence by training is largely a graveyard - working-memory training famously did not transfer. Any claim in this space starts owing a debt of skepticism, and this paper pays it by shrinking its claim to what the data holds.
The falsifiable protocol
The machine arm is specified in full in Section 11 - a three-stage pre-registered ablation, $2,632, with the falsifying result named in advance. What follows is the human arm, which no budget can shortcut.
To test the human half properly: pre-register the design; recruit n ≥ 40; randomise to ATHENA-discipline training versus an active control (equal contact time, inert content); administer standardised fluid-intelligence and applied-reasoning instruments before, after, and at six months, by blinded administrators; publish all results including nulls. The framework text is public. The prediction: the treatment arm shows durable gains on applied, tool-permitted reasoning tasks; the paper makes no prediction of gains on abstract matrix IQ, because the framework trains verification and architecture, not processing speed. If the treatment arm shows nothing, this paper is wrong and says so in advance.
Nothing here is medical, genetic, or clinical advice. No diagnosis. The Alzheimer row stays refused. Percentiles are ranks against a sequenced research cohort - scientific, not clinical.