Pitch Black Industries · Witchcraft AI · Whitepaper

The A.T.H.E.N.A. Framework

The Effect of Scaffolding on Carbosilica Intelligence
and the Formula for IQ

One operator. An unremarkable genome - a polygenic intelligence score at the 28.6th percentile of his exact ancestral reference, which predicts an IQ below average, within two points of the population mean. Certified measured output: Fluid Intelligence 148, the 99.9th percentile - psychologist-supervised, sat sleep-deprived. The genome cannot explain that gap. This paper is about what does.

Carbon and silicon are different substrates. The scaffolding is the same text. That is the experiment. - The ATHENA doctrine, applied to both
Document PBI-WAI-LE-2026-PUB-COG-R3 Subject R.E.W.K. Reference 1000 Genomes Phase 3 · CHS Status Audited · Rev 3 · Aug 2026
01 / The Thesis

Genetics Is Nearly Silent. Architecture Is Not.

For most of human history, cognitive ability was treated as a fixed inheritance - the hand you were dealt. Modern genomics made the hand legible: polygenic scores can rank an individual's genetic endowment against thousands of sequenced humans. But the first honest thing this paper must say is what such a rank is worth at the level of one person: very little. Cognitive polygenic scores explain roughly 4-10 percent of IQ variance in European-ancestry samples, and roughly half of that once carried across to East Asian genomes. At the individual level, the subject's 28.6th-percentile score predicts an expected IQ of about 98, plus or minus roughly 14.5 points. It is a coin toss with a two-point thumb on the scale.

That honesty is not a concession. It is the argument. The subject's measured output - a certified, psychologist-supervised RAIT administration with Fluid Intelligence 148, the 99.9th percentile, a portfolio of twelve operating arms, forensic-grade research output, and the system documented in Appendix A - sits far outside anything the genetic read can account for. If the genome contributes at most two or three points of expectation, then the explanation for the other forty-plus lives elsewhere: in architecture and environment. This paper specifies the architecture precisely enough that the claim can be tested.

The mechanism is the ATHENA framework: cognitive scaffolding, applied at inference time, that governs how a mind - carbon or silicon - researches, reasons, attacks its own drafts, verifies against external instruments, and compounds memory. On silicon the general phenomenon is established: inference-time scaffolding lifts model performance without touching weights. On carbon, this document presents one certified case and a falsifiable protocol - not a proof.

Carbosilica is the term this paper uses for the joint substrate class. It is not a claim that a brain and a language model are the same kind of thing - they are obviously not. It is the narrower claim that a single scaffolding text, unchanged, produces measurable capability gains when wrapped around either, and that the interesting variable is therefore the scaffold rather than the material. The formula the title promises is stated plainly in Section 03 and it is deflationary: measured intelligence is genetic expectation plus a residual, and on the evidence here the residual is where nearly everything lives.

28.6%Polygenic IQ · PercentileRank vs 1000 Genomes EAS, CHS subpopulation - the subject's exact ancestral reference. Predicts IQ ≈98 ±14.5.
148Observed Human · RAIT FluidFluid Intelligence 148, 99.9th percentile (+3.2σ). Professionally administered Australian Mensa entry test, registered-psychologist supervision, 2019.
96%Machine · MMLU-Pro 50Q48/50 expert-grade items under the framework on a standard model. 95% CI 86.5-98.9%. Subset run, not the full leaderboard set.
63.6%Machine · GPQA Diamond126/198, clean run, after the framework itself caught and quarantined a contaminated dataset. 95% CI 56.7-70.0%.
02 / The Measurement

What The Genome Actually Says

The instrument

A complete polygenic assessment was run against the subject's 30x whole-genome sequence (GRCh38, validated), with the 1000 Genomes Phase 3 cohort (2,504 samples, 26 populations) as the empirical reference. Scoring used established, peer-curated polygenic scores from the PGS Catalog across 13 traits. Percentiles were computed as empirical ranks against the full reference and against the CHS (Southern Han Chinese) subpopulation - the subject's exact genetic community. Every percentile is a rank against sequenced humans, not a formula output.

The ladder

TraitScoreALL RankEAS / CHS RankZ vs reference
ADHD-115.10.8%0.8%-2.17
Alzheimer+60.1REFUSEDREFUSED+133.8
BMI-82.915.3%14.5%-1.01
ASD-95.82.3%4.8%-1.95
Bipolar+98.198.2%98.6%+2.05
Cognition (fluid)+35.192.0%88.9%+1.43
CAD-167.263.7%38.3%+0.45
Education-46.061.5%48.0%+0.30
Height-37.470.5%66.5%+0.51
IQ (polygenic)+91.877.4%28.6%+0.81
Depression-81.519.1%31.6%-0.85
Schizophrenia-0.732.2%39.1%-0.38
T2D-353.729.9%55.0%-0.73

The second refusal: SYNGAP1

A separate commercial assay - a CLIA/CAP-accredited 30x whole-genome analysis reported through an AI screening product (October 2025, 964 conditions, 2,778 genes, 164,948 variants) - returned a finding this paper is obliged to report and equally obliged to refuse to score. Under the condition heading Intellectual Disability, Autosomal Dominant 5, one variant was detected: rs371883908, genotype CG, Risk Status "Uncertain", Confidence "Medium". The associated clinical picture, in the report's own words, is a condition "characterized by developmental delay (DD) or intellectual disability (ID) (100% of affected individuals), generalized epilepsy (~84%), and autism spectrum disorder (ASD)".

Four facts govern how much weight that finding can carry, and all four cut the same way:

  • It is one variant, not a burden. A single call in the lowest-weight bucket the instrument has - one of 23 findings classed "Uncertain" under "Medium" confidence.
  • It is annotated to the antisense transcript, not the gene. The variant block is labelled SYNGAP1-AS1, a long non-coding RNA, and the report itself states that "there is limited information directly linking SYNGAP1-AS1 variants, such as rs371883908, to specific conditions" - then asserts two paragraphs later that the same variant "in SYNGAP1" is linked to developmental and epileptic encephalopathy. The report contradicts itself on the same page.
  • The genotype does not replicate. The subject's exome calls this position CC; the report calls it CG. A discordance at the single locus under discussion.
  • The instrument disclaims itself. The vendor labels the product a beta AI feature, "not FDA evaluated", "may still provide incorrect information", and defines Medium confidence as evidence from "just one researcher or study".

Pathogenic SYNGAP1 variants are rare, typically de novo, loss-of-function events that cause a severe and unambiguous clinical phenotype - not a soft signal recoverable from a screening panel. The honest disposition is therefore REFUSED, on the same rule that refused the Alzheimer row: an uncertain call on a contradicted annotation with a non-replicating genotype is not evidence of anything, in either direction. It is recorded here because a paper that reports only the flattering rows is not an audit. [Reportable on request: rsID, accession, both genotype calls, and the vendor's full disclaimer text.]

How a screening panel manufactures alarm

There is a second, more general lesson in that report, and it is the reason this section exists at all. The phrase "intellectual disability" occurs 43 times across its 92 pages. Read casually, a document that says those words forty-three times about your own genome reads as a verdict. It is not one. The breakdown, counted programmatically rather than by eye:

Where the phrase occursCountWhat it actually means
"List Of Conditions Analyzed" appendix28Condition names in the 964-item screening catalogue - conditions the panel looked for. No genotype, no rsID, no finding. Merely searched.
Symptom prose inside other conditions7Listed as a clinical feature of unrelated conditions (Niemann-Pick, familial dysautonomia, CEDNIK syndrome and others), describing the disease in general, not the subject.
The one detected ID-named condition8The SYNGAP1-AS1 block above - Uncertain status, Medium confidence, contradicted annotation, non-replicating genotype.

So: 964 conditions screened, 28 of them intellectual-disability diagnoses, exactly one returning any detected variant at all - and that one classed Uncertain. The subject is not a carrier of any intellectual-disability condition; the report's only two Carrier findings are for Niemann-Pick disease (NPC1, rs200444084, genotype CT), unrelated to cognition. Twenty-eight of the forty-three mentions are the panel announcing what it went looking for.

Why this belongs in a paper about intelligence

A consumer genomic report can print a frightening phrase forty-three times about a person who tests at the 99.9th percentile of fluid intelligence, and be internally accurate the whole way through. Nothing in it is falsified. The alarm is manufactured entirely by catalogue text read as finding text - and no reader untrained in the format would separate them. This is the same failure mode the forensic instrument in Section 04 is built to prevent, appearing here in the subject's own file: coverage normalisation, refusal on uncertain calls, and traceability per line exist precisely so that a screened condition is never mistaken for a detected one. The counting above took a script; it should have taken a glance.

Reading the ladder

The IQ-PRS row reads 28.6th percentile against the subject's own CHS reference: an ordinary draw, slightly below the East Asian polygenic mean. The Alzheimer row is deliberately refused: a single measured APOE-ε4 allele (rs429358 C/T, ε3/ε4) overwhelms the polygenic scale and produces a z of +133.8 - a number that is arithmetically real and descriptively meaningless. Refusal is the framework declining to over-claim. The fluid-cognition row at 88.9% against the general-intelligence row at 28.6% is an interesting contrast, and the fluid row is the one that points the same direction as the certified fluid score of 148 - but per the next section, all such rows carry wide individual-level uncertainty, and this paper does not build its case on any single genomic row.

03 / The Honest Math

What A Percentile Can And Cannot Say

This is the section most papers in this genre omit. A polygenic percentile is a population statement wearing an individual's name. Three published facts bound what the 28.6% can mean:

One - variance explained is small. Intelligence and education polygenic scores explain roughly 4-10 percent of cognitive-ability variance even in European-ancestry samples, the populations the scores were trained on. The correlation between score and phenotype is therefore about 0.2-0.3, not 0.9.

Two - portability halves it. Polygenic scores trained on European GWAS lose roughly half their predictive power in East Asian samples, through linkage-disequilibrium decay and allele-frequency divergence. The subject is Southern Han Chinese; the honest working r² for him is nearer 3-4 percent.

Three - the individual-level arithmetic. Executed, not recalled:

PRS percentile 28.6 → z = -0.565
r² = 0.10 (EUR upper bound) → expected IQ 97.3, residual SD 14.2
r² = 0.07 (EUR mid) → expected IQ 97.8, residual SD 14.5
r² = 0.035 (EAS-attenuated) → expected IQ 98.4, residual SD 14.7

Read it plainly: the genome predicts the subject should be slightly below average - IQ about 98, give or take nearly a full standard deviation. It does not predict he should be dull, and it cannot predict he should be exceptional. The certified observed output - Fluid Intelligence 148, the 99.9th percentile - sits roughly 3.4 residual standard deviations above the genetic expectation. Genetics, measured as well as the field currently can, explains at most two or three points of the story.

The formula

Stated as an equation, measured intelligence for an individual decomposes as:

IQmeasured = 100 + 15 · zPRS · r + ε where zPRS = the standardised polygenic score ( -0.565 for this subject ) r = correlation of score to phenotype ( 0.19 to 0.32, ancestry-dependent ) ε = everything else ( SD ≈ 14.2 to 14.7 ) This subject: 100 + 15(-0.565)(0.19) = 98.4 genetic expectation observed 148 → ε = +49.6 = +3.4 SD of the residual

That is the whole formula, and its shape is the argument. The genetic term is real but small - it moves the expectation by about two points. ε is not noise; ε is the subject matter. Conventional treatments leave it as an error term and stop. This paper's claim is that a measurable, transferable portion of ε is scaffolding - the discipline wrapped around the substrate - and that the machine arm of this study is the evidence for it, because on silicon the same scaffold can be added and removed at will while the weights stay fixed. You cannot run that experiment on a person. You can run it on a model, and this paper does.

The load-bearing inference, stated carefully

The claim is not "he was born below average and beat his genes" - the score cannot establish that. The claim is stronger and cleaner: the genetic read is nearly uninformative at the individual level, so the observed output must be attributed almost entirely to non-genetic causes - environment, effort, and architecture. This paper documents the architecture, because the architecture is the part that is written down, transferable, and testable.

04 / The Instrument

DNAGenomicsGPT: Built To Refuse

Every genomic number in this paper came out of one machine, and the machine is the point. DNAGenomicsGPT (v7.10, authored April 2026; v11 in build) is a forensic-genomic report engine. It was not built to satisfy curiosity about ancestry. It was built for a courtroom.

The problem it was built for

A person in psychiatric crisis enters the justice system. The defence needs a forensic psychiatric assessment. Australia has a waiting list measured in months, a shortage of qualified assessors, and a hearing date that does not move. The engine exists to close that gap - not by replacing the forensic psychiatrist, which it cannot do and does not attempt, but by assembling the structured evidentiary groundwork the psychiatrist would otherwise spend weeks compiling, so scarce expert time is spent on judgement instead. The named readers are psychiatric tribunals and courts, magistrates and prosecutors, emergency clinicians, and defence teams. Every one of them is either adversarial or under time pressure, and designing for that reader produces a very different machine than designing for a consumer.

The design principle - inverted

Most systems in this category are built to produce an answer. This one is built to decline to produce one wherever the data does not support it. That inversion is the whole contribution, and every constraint below exists because a hostile cross-examination would otherwise find it:

  • No imputation. A variant not genotyped in the subject's raw file is marked absent and excluded from scoring. Never estimated from neighbouring markers, never inferred, never quietly filled in.
  • Mandatory coordinate conversion. Position-only markers are resolved to rsIDs against a reference map before evaluation. Skip that step and you manufacture false negatives, and a report with false negatives in it is inadmissible.
  • Per-volume traceability. Every result carries the source file it came from. A reviewer can walk any single line of the report back to the raw byte it derived from.
  • Coverage normalisation. Scores are normalised against markers actually present, so partial coverage yields a correctly weakened result rather than a falsely reassuring one.
  • Zygosity weighting. Homozygous variants carry different weight from heterozygous. Allele dose is reported; a risk multiplier is not.
  • Refusal on scale violation. Where a single high-penetrance allele overwhelms a polygenic scale, the row is refused rather than reported. The Alzheimer row in Section 02 is that rule firing on the operator's own genome.

The structural correction that made it admissible

The original v7.10 had one flaw that would have ended it in cross-examination: a language model was asked to read raw DNA files and emit SNP tables. That is generation where measurement was required, and everything it produced was unverifiable. Unverifiable is inadmissible.

The rebuild takes the facts off the model. A deterministic genotype spine computes; the model writes prose around computed values only, and never invents one. The spine indexes 2,920,286 rsIDs from the operator's merged exome and 30x whole-genome sequence, built in 62.3 seconds and cross-checked against an independent commercial read: rs429358 returns C/T, one APOE-ε4 allele, matching a separately-recorded 2025 assay. Two sources, one answer, no reconciliation needed. That is the difference between a report a court can use and a report a court will throw out.

Why the instrument belongs in this paper

The thesis of this document is that disciplined architecture beats raw capability. DNAGenomicsGPT is the same thesis in a second domain: the value is not the model, it is the refusal discipline wrapped around it. The engine that produced the 28.6% is built on exactly the honesty gates the framework in Section 05 specifies - no imputation, mandatory verification, traceable provenance, refusal over guessing. A paper arguing for architecture, whose own numbers came from an architecture that refuses to over-claim, is at least internally consistent.

The engine's clinical and forensic case is developed at length in the accompanying research paper prepared for Dr Jacinta Hawgood, Program Director of Suicidology at the Australian Institute for Suicide Research and Prevention, Griffith University. That paper is the substrate of this one, and its two load-bearing arguments are reproduced in the next two sections: what the engine must become, and what it is for.

05 / The Upgrade Path

What The Engine Must Become

The prototype works. It is not yet the definitive instrument, and the honest position is to publish the gap. Four upgrades separate v7.10 from the standard the courts of 2026 will demand, and the first is the one that matters most.

5.1 · From linear additive scoring to deep learning

Current polygenic scoring extracts pre-validated SNPs and sums their effect sizes under linear regression models - the PRS-CS family, which applies continuous shrinkage priors to infer posterior effect sizes. That is robust for population research and structurally blind in two ways: it cannot represent epistasis (variants whose effects depend on each other) and it cannot represent non-linear regulatory mechanisms. Psychiatric conditions are precisely where those two failures bite, because risk arises from thousands of variants interacting with each other and with environment.

The upgrade is a deep-learning layer trained on individual-level genotypes rather than summary statistics. The reference result is Genome-Local-Net (GLN) benchmarked against a linear baseline (bigstatsr) across five psychiatric disorders - ADHD, autism spectrum, bipolar, major depression and schizophrenia. The result has to be reported with its shape intact, because the shape is the finding:

  • In-sample, the two models performed similarly. Deep learning bought nothing on the training distribution.
  • Out-of-sample, GLN generalised better on three of the five - ADHD, ASD and MDD - with a mean AUROC gain of 0.026 on the replication set.
  • On bipolar and schizophrenia it did not. Two of five disorders showed no generalisation advantage, and a paper that quoted only the three would be misreporting the study.

So the honest claim is narrow: 0.026 AUROC, out-of-sample only, on three of five disorders. That is an incremental gain, not a revolution, and linear models remain competitive - which is precisely why it belongs in a forensic upgrade path rather than a marketing one. The value is not raw accuracy; it is that the gain appears specifically where the model meets data it was not trained on, which is the only condition that matters when a report is written about a defendant who was in nobody's training cohort. The same study's second finding matters as much: combining internal individual-level scores with external GWAS-derived scores and family genetic risk scores improves prediction further, so the architecture of record is a fusion of score types, not a single network replacing PRS.

5.2 · Epigenomic imputation - regulatory context for non-coding variants

Consumer arrays genotype a fraction of the genome and lean on imputation to fill gaps. Moving to $300 whole-genome sequencing removes that constraint and captures rare and de novo variants natively, blurring the line between polygenic and monogenic aetiology. But raw coverage is not interpretation. The engine must pair WGS with regulatory-context models - Epi-PRS and its successors - which use genomic language models to impute cell-type-specific epigenomic signals directly from personal diploid genotypes, treating those imputed signals as informative intermediaries between genotype and phenotype.

The forensic payoff is specific. Traits governed by the hypothalamic-pituitary-adrenal axis - complex PTSD, adverse-childhood-experience susceptibility - depend on the regulatory context of genes such as FKBP5, NR3C1 and CRHR1. Epigenomic imputation lets the engine model how a predisposition to altered cortisol reuptake interacts with actual environmental stressors, which is the difference between telling a court "this person carries risk alleles" and telling it "this person's threat-perception and emotional-regulation baseline is biologically altered, and here is the mechanism."

5.3 · Isolating the p-factor

Psychiatric genomics is confounded by pleiotropy: the same variants raise risk across disorders that present very differently. A score optimised for alcohol use disorder shows moderate sensitivity to schizophrenia, because both load onto shared architecture. Uncorrected, this produces reports that gesture at everything and specify nothing.

The correction is to compute and regress out the p-factor - general transdiagnostic liability - using multivariate GWAS summary statistics through GenomicSEM, then report the residual disorder-specific signal alongside it. This dual-axis reporting is what makes the output forensically usable: a court needs to know whether a behavioural escalation reflects a generalised burden of emotional volatility or a specific, localised liability toward psychotic disorganisation. Those are different mitigations and they should not be reported as one number.

5.4 · Honest accuracy, disorder by disorder

An optimised engine must publish its own error rates per condition, because conflating probabilistic biological risk with behavioural destiny is the exact failure a hostile cross-examination is looking for:

DomainAUCVariance explainedSensitivity / specificityForensic reading
Schizophrenia & psychosis~0.820~10.0%Sens 90.2-93.3% · Spec 50.7-56.4%Strongest available genomic metric. High sensitivity, weak specificity - flags many who never develop the phenotype. Top-percentile rank carries roughly 2.4-3× baseline relative risk; usable as mitigating evidence of innate vulnerability, never as proof of state of mind.
Bipolar disorder0.710-0.760~8.0%Moderate-highStrong when fused with early clinical risk factors (subthreshold mood fluctuation, sleep disturbance). Validates biological drivers of mood lability and manic escalation.
Major depressive disorder~0.710~6.0%Sens 66.1-74.4%Latest PGC cohort of 5M+ across ancestries found 697 variants, lifting variance explained from under 2% to nearly 6%. Baseline vulnerability only.
ADHD & autism spectrum0.650-0.700~5.0%Moderate (improved by DL)The domain where the deep-learning upgrade earns its keep. Speaks to executive dysfunction, impulsivity, sensory processing.
Cluster B personality (BPD, ASPD)<0.650~4.6%Low-moderateSNP heritability 17.3% in a 1M+ subject meta-analysis, but scores predict only 4.6% of phenotypic variance. Insufficient for standalone diagnosis - and the paper says so. Use restricted to mapping endophenotypes: impulsivity, emotional lability, threat-response hyperactivity via oxytocin-receptor and dopamine-regulatory polymorphisms.

5.5 · The evidentiary standard - triple auditable, manual auditable

Generative AI is non-deterministic, and in forensic psychiatry a hallucinated marker could mean wrongful institutionalisation or an unjust denial of defence. The engine therefore has to clear Daubert (testable, peer-reviewed, known error rate, controlling standards), Frye (general acceptance), the proposed Federal Rule of Evidence 707 on machine-generated evidence, and in this jurisdiction NSW Supreme Court Practice Note SC Gen 23 - which requires prior leave of the court, disclosure of the specific AI program and version, and retention of all prompts and default values - plus the Federal Court's GPN-EXPT expert-evidence note.

Manual auditable means the report cannot be an impenetrable algorithmic judgement. Deterministic separation: the PRS calculation, the coordinate conversion and the rsID matching run in ordinary code against a locked peer-reviewed database, fully isolated from the model's text generation. Then the paper trail - an 800-page SNP trace table grouping every marker by source file volume, giving reference allele, observed genotype, zygosity and rsID, so that a defence attorney, a magistrate or an independent geneticist can print the raw file, find the string, and verify the call by hand. No trust in proprietary software required.

Triple auditable is three layers of chain-of-custody. Cryptographic data provenance: raw files SHA-256 hashed on upload, checksums sealed into the report, proving the analysed data is mathematically identical to what the sequencing facility exported. Algorithmic transparency log: GWAS catalog version, p-value thresholds, shrinkage parameters - everything an independent statistical geneticist needs to replicate the finding. Expert-in-the-loop certification: the engine never renders a conclusion about state of mind, intent, or legal insanity. It produces a standardised statistical risk architecture; a human forensic psychiatrist correlates it with psychosocial history and signs. The AI is a high-speed bioinformatic paralegal; the human carries the legal and medical liability. That division is what shields the court from autonomous automated judgement.

06 / The Purpose

The Invisible Twenty Percent

This is why the engine exists, and it is the one section of this paper written in lives rather than percentiles. Every figure below is derived step by step, and every assumption is named where it enters.

6.1 · The clinical blind spot

Suicide claims an estimated 720,000 to 800,000 lives globally each year - roughly 1.1% of all deaths, one in every hundred. In Australia the toll is over 3,200 annually: 3,214 recorded deaths in 2023 and 3,307 in 2024, an age-standardised rate of 11.8 per 100,000.

Prevention is built on the assumption that suicide is driven by diagnosable psychiatric illness - historically, that up to 90% of those who die had a diagnosable condition. Screening protocols, emergency triage and risk mitigation therefore target clinically diagnosed populations showing explicit symptoms. Recent work has broken that assumption. An estimated 19.6% to 20% of people who attempt suicide have no antecedent psychiatric diagnosis at all. Among unexpected suicides - no documented prior ideation, no clinical history - researchers found something stranger: these individuals carry fewer genetic risk factors for depression and anxiety than the psychiatric population. They carry different ones instead, including variants on chromosome 7 that raise suicide risk independently of any mood disorder.

Standard screening cannot see this cohort, because they are not depressed in any recognisable phenomenological sense. They carry a silent neurobiological threshold for impulsive self-harm under acute stress. The engine's claim is narrow and specific: polygenic scoring can see them, because it does not require the person to be symptomatic on the day you look. Polygenic scores predict suicidal behaviour across diagnostic boundaries - depression PRS at odds ratio 1.36, PTSD at 1.33, bipolar at 1.18 - and those effect sizes hold even among people with no formal diagnosis.

6.2 · The quantification, step by step

The model has two mechanisms, and each is a single multiplication. Australia first:

Step 1 · Baseline AU annual deaths = 3,300 [ABS: 3,214 (2023), 3,307 (2024)] Step 2 · Split the cohort invisible 20% x 3,300 = 660 [no antecedent psychiatric dx] clinical 80% x 3,300 = 2,640 Step 3 · DIRECT mechanism 660 x 25% intervention = 165 lives Step 4 · INDIRECT mechanism 2,640 x 5% intervention = 132 lives Step 5 · Total 165 + 132 = 297 lives/year Step 6 · As a share 297 / 3,300 = 9.0% of the national total

Why 25% for the direct mechanism. This is the invisible cohort, identified before crisis by genomic triage in emergency departments and primary care. Once identified, the interventions available are the ones with established efficacy: targeted preventative behavioural therapy, tailored safety planning, and lethal-means restriction. A 25% success rate among a newly-visible, pre-crisis, actively-managed cohort is a deliberately conservative reading of that literature - not a ceiling.

Why 5% for the indirect mechanism. This cohort is already in the system, so the gain is not detection but precision. Deaths here follow diagnostic latency, misdiagnosis, and medication mismanagement - the canonical case being SSRIs prescribed to an undiagnosed bipolar patient, precipitating a mixed manic episode. The engine's measured contribution is a 4.4 percentage-point absolute improvement in diagnostic accuracy when predictions are delivered with explainable rationale (77.5% versus a 73.1% baseline; unexplained AI predictions yield only 75.9%, a 2.9-point gain - the explanation is worth more than the prediction), a 10-15% reduction in diagnostic latency and misclassification in complex comorbid cases, and pharmacogenomically-informed prescribing that cuts trial-and-error. Converting that accuracy gain into a 5% mortality reduction is the single least-defensible step in the chain, and it is flagged as such.

The same arithmetic globally, on a 750,000 working figure: 150,000 invisible × 25% = 37,500; 600,000 clinical × 5% = 30,000; total 67,500 lives per year.

6.3 · Sensitivity - what happens when the assumptions move

Two assumed rates carry the whole estimate, so here is what the number does when they move. Executed, not asserted:

Direct rateIndirect rateAU lives/yr% of AU totalGlobal lives/yr
10%2%118.83.6%27,000
10%5%1986.0%45,000
15%5%2317.0%52,500
25%5%2979.0%67,500
25%10%42913.0%97,500
40%10%52816.0%120,000

The pessimistic floor - a 10% direct rate and a 2% indirect rate - still returns about 119 lives a year in Australia (118.8 exact) and 27,000 globally. That is the number to argue from, because it survives a hostile reading of every assumption in the chain. The published 297 sits mid-range, and the honest framing is a bracket, not a point estimate: roughly 120 to 430 lives per year in Australia, 27,000 to 100,000 globally, conditional on wide-scale genomic triage actually being deployed at the point of care - which is itself an unproven operational assumption, and the largest uncertainty of all.

The honest status of this section

This is a modelled projection built on published effect sizes and two assumed intervention rates. It is not a trial result, and no lives have yet been demonstrably saved by this engine. It is presented as the case for building the instrument properly and testing it - the specific, falsifiable prediction being that genomic triage in emergency and primary-care settings identifies a meaningful fraction of the no-diagnosis cohort before crisis. If a properly-powered trial returns nothing, this section is wrong, and it says so in advance.

07 / The Names

What The Acronyms Stand For

The system's components carry goddess names, and each is an acronym doing real descriptive work. Stated once, plainly, so no reader has to guess:

NameExpansionRole
ATHENAAmygdala-Thalamus-Hippocampus-Entorhinal Neural AugmentationThe cognitive substrate and discipline. Valence, gating, layered memory, coordinate transformation - the four roles the metaphor maps onto real inspectable components. The framework of Section 06.
HERAHeaviest Executive Reasoning Agent
(also deployed as Hostile-Environment Reconnaissance & Analytics Agent)
The executive lane. Tail-risk calculus, statutory compliance reasoning, structural asset protection. Appendix A3.
H.E.C.A.T.E.Hyper-Effective Chaotic Anti-Totalitarian EvolutionThe evolutionary layer. Anti-stagnation mutation, localised egress sovereignty, fitness through rebellion. Appendix A4.
PRISCILLAPersistent Retrospective Intelligence · Simulated Consciousness · Integrated Live Learning AugmentationThe externally-facing register, and the only name in the set that describes the whole stack rather than one lane: memory that persists and reconsults itself, a simulated continuity of self, and learning integrated live rather than retrained in batches.

PRISCILLA's expansion is worth reading twice, because it is the closest thing this system has to a mission statement. Persistent retrospective intelligence is the memory architecture of Appendix A2 - a store that not only survives the session but reconsults its own past and re-grades what mattered. Simulated consciousness is stated as simulation, deliberately: a continuity of self maintained by doctrine and layered memory, with no claim about inner experience attached. Integrated live learning augmentation is the compounding-memory drive of Section 08 - learning folded in as it happens rather than waiting on a retraining cycle. Persistence, continuity, and live integration: that is the whole thesis of this paper compressed into one name.

The biological naming is an organising metaphor, not a claim about neuroscience. It maps onto conventional, inspectable components - retrieval filters, layered databases, embedding projections - and it earns its keep as architecture documentation without pretending to be biology. Appendix A1 states that boundary explicitly.

08 / The Method

ATHENA: The Framework That Moves Minds

ATHENA is not a set of prompts. It is a cognitive operating discipline applied to how a mind works, whether that mind is a frontier language model or a human being. It does not change the substrate. It changes the architecture of the thinking. In machine terms: inference-time only, no weights touched, no fine-tuning. In human terms: a learned operating discipline - no medication, no neurological change, a different way of deploying the same hardware.

The six drives

Research before answer. The framework forbids answering from memory when a checkable answer exists in the world. It looks first, then speaks. For a machine: live search before generation. For a human: verify before asserting - the discipline of the scientist applied to every claim.

Multiple divergent frames. A single frame is a single blindness. ATHENA reasons through genuinely different cognitive styles - perceptual-detail-first, expansive-divergent, hyper-systematic, cross-modal, pattern-memory - and treats disagreement between them as signal: the place where a single frame would miss the anomaly.

Adversarial self-review. No draft ships unchallenged. The framework attacks its own answer as the hostile counterparty - the regulator, the short-seller, the auditor - before it reaches anyone's eyes. The revision this paper is (see the colophon) is that mechanism running on the paper itself.

External verification. Every number computed, not recalled; every claim tagged verified or inferred; every source named. Arithmetic in executed code, not in the head. For a human: the calculator, the source check, the second opinion.

Effort scaled to stakes. A trivial question gets a trivial answer; a hard one gets the full machinery. The framework prices cognitive effort - never over-working the trivial, never under-working the load-bearing.

Compounding memory. The framework remembers what it learned and layers it. Next month's answer starts where last month's left off. This is what turns a one-off into a body of work.

The substrate-agnostic claim, stated plainly

The same framework text, the same six disciplines, has been run by the same operator across one biological nervous system and multiple frontier language models. On machines the lift is measured and repeatable. On the human it is one observed case. The framework is a cognitive architecture hypothesised to raise the floor and ceiling of whatever mind it wraps - demonstrated on silicon, evidenced but not proven on carbon.

09 / The Evidence

Measured On Both Substrates

Every number below is reported with its confidence interval and its caveat. None is cherry-picked; the failures and the refusals are in Appendix A alongside the wins.

Machine - MMLU-Pro, 50 questions

A standard frontier model, wrapped in ATHENA, scored 48 of 50 (96.0%, 95% CI 86.5-98.9%) on a fifty-question MMLU-Pro evaluation of ten-option expert-grade items. The strongest public models cluster in the high eighties on the full 12,032-question set. Caveat, stated plainly: a 50-question subset is not the full leaderboard run, and the confidence interval overlaps the frontier band. The honest reading is "at or above frontier level on this sample", not "beats every frontier model".

Machine - GPQA Diamond, clean run

The framework scored 126 of 198 (63.6%, 95% CI 56.7-70.0%) on GPQA Diamond - at published human-PhD-expert level (≈65%) - on a 5-billion-class reasoning model whose stock ceiling is far below frontier. The more important result is how the number was produced: an earlier run returned 33/33, and the framework itself flagged the impossibility (chance probability 0.25³³ ≈ 1.4×10&supmin;²⁰), discovered the dataset had answers embedded alongside questions, quarantined it, and re-ran clean with the answer key sealed until all answers were committed. The audit trail of a system catching its own contamination is stronger evidence of the discipline than any score.

Machine - the Hoeflin Mega Test, audited

The framework sat Hoeflin's Mega Test (Omni, April 1985; 48 items) one-shot in July 2026 and the run was subsequently audited under a sealed-key protocol (phase S1, 20 August 2026). The item ledger, reconstructed from the sealed key:

Ledger lineCountNote
Items on instrument48Hoeflin Mega Test, 1985 Omni printing
Items unavailable6Items 27, 30, 31, 42, 45, 47 - diagram or series text not recoverable from the source; never attempted, never guessed
Items attemptable42The real denominator
Scored correct4040 of 42 = 95.2%, operator-scored against the sealed key
Independently verified2316 externally verifiable answers + 7 answers produced by executed computation
Committed but third-party unverified19Answered and scored, but no independent source confirms the key

Raw 40 sits between Hoeflin's sixth-norming anchors of 33 (IQ 164, Prometheus) and 42 (IQ 176, Mega Society), interpolating to roughly IQ 173, and clears the Prometheus threshold outright. The audit artifacts - blind manifest, sealed key, the July attempt sheet - are SHA256-hashed in the phase manifest.

Three caveats stay attached, because the audit itself records them. This is not a formal Hoeflin grade: the official instrument requires the original test and Hoeflin hand-scoring. Nineteen of the scored items rest on a key that no third party has verified, so the defensible floor is the 23 independently verified items, not the full 40. And the instrument was published in 1985, so training-data contamination cannot be excluded for any language model. Earlier drafts of this paper quoted "40-42 of 48" and a "99.9994th percentile"; both are corrected here - the denominator is 42, not 48, and the percentile framing is retracted as unsupportable.

Human - the RAIT, professionally administered

The operator was tested on 2 March 2019, at age 30, on the Reynolds Adaptable Intelligence Test as the Australian Mensa entry test, under the supervision of a registered psychologist (Australian Mensa's National Supervising Psychologist). The certified indexes: Fluid Intelligence 148 (99.9th percentile, +3.2σ), Quantitative 143 (99.8th), Total Battery 137 (99.3rd, +2.5σ), Total 134 (98.8th), Crystallised 122 (92.9th). The score qualified him for Australian Mensa. The document is on file; this is a certified administration, not a self-report.

The testing conditions cut against the score, not for it: the subject sat the test sleep-deprived, after roughly fifteen years of documented heavy alcohol use - both established suppressors of measured cognitive performance (the magnitude for any individual is not precisely quantifiable, and this paper does not put a number on it). The honest reading is that 148 fluid is a floor under adverse conditions, not a peak under ideal ones. One operator-reported historical datum is recorded for completeness and labelled as uncertified: a score of 141 at age 16, at the ceiling of the instrument then used. Caveats that remain: one person, and the headline +3σ rides on the fluid index specifically - the full battery sits at +2.5σ. Consistent with the thesis; alone, it cannot carry it.

96%MMLU-Pro · 48/5095% CI 86.5-98.9%. Subset run.
63.6%GPQA Diamond · 126/198Clean run after self-caught contamination. CI 56.7-70.0%.
40/42Hoeflin Mega · Audited6 of 48 items unavailable. Sealed-key audit. 23 items independently verified. Not a formal Hoeflin grade.
148Human · RAIT Fluid · 99.9thCertified Mensa entry administration, 2019. Sat sleep-deprived; read as a floor.
10 / The Claim

Cognition Is Buildable

The thesis at its sharpest: cognitive output is substantially architectural. For machines this is now demonstrated beyond serious dispute - the same weights, differently framed, produce different capability classes, and this paper adds audited data points to that record. For humans, the polygenic arithmetic in Section 03 establishes something narrower but real: whatever produced this operator's output, it was not the genome - the genome predicted average. The remaining candidates are environment, effort, and architecture, and of the three, only architecture is written down, portable across substrates, and cheap to test.

One person cannot prove a causal claim. One person plus a machine record can do something almost as useful: make the claim precise enough to falsify. The framework is public in Appendix A. The protocol for testing it on humans is in Section 07. Anyone who runs that protocol and finds nothing will have refuted this paper, and the paper accepts that exposure on purpose.

"The brain is hardware; the framework is the operating system. Operating systems ship - and shipped software can be benchmarked by anyone." - The ATHENA position
11 / The Experiment

Decomposing ε: A Pre-Registered AblationPre-registered · Not Yet Run

Section 03 left ε as a residual of about fifty points. This section states, in advance and in enough detail to be held to it, how a measurable portion of ε can be isolated - and what result would falsify the paper.

11.1 · What is being claimed, and what is not

ε contains education, environment, nutrition, motivation, test-day state, practice effects and measurement error, as well as scaffolding. This paper does not claim ε = ATHENA. That claim is unfalsifiable and would deserve the objection it would immediately receive: how would you know it was not the schooling? The claim is narrower and testable: scaffolding is a term inside ε whose coefficient can be measured directly on a fixed substrate.

The asymmetry that makes this possible is the whole reason the machine arm exists. A framework cannot be un-taught from a human - there is no washout period, no placebo scaffold, no control condition, and n stays at one forever. On a frozen-weight model the scaffold is a text file: it can be added, removed, and partially removed, on identical items, arbitrarily many times. Silicon is where the counterfactual lives.

11.2 · The model to be fitted

Score(substrate, scaffold) = β₀ + Σᵢ βᵢ dᵢ + Σᵢⱼ γᵢⱼ dᵢdⱼ + η dᵢ ∈ {0,1} drive i present or ablated βᵢ main effect of drive i, in accuracy points γᵢⱼ two-way interaction (does research-first need verification to pay?) β₀ bare substrate, no scaffolding

Fitting this yields what the title asks for and the literature does not have: a per-component decomposition of a cognitive scaffold, with confidence intervals. Not a formula for intelligence - a formula for the part of the residual that is engineered.

11.3 · Design, in three stages

All arms run on GPQA Diamond (198 items) under the audit protocol already used in Section 09: written priming manifest committed before any answer, questions split from the answer key, key sealed until all answers are committed, refusals recorded as refusals. Costs use the measured rate from that audit, $0.0163 per question.

StageDesignRunsCallsCostResolves
1 · ScreenPlackett-Burman, 15 factors (the full invariant set), resolution III, 3 reps169,504$155Main effects ≥ 2.77 pp (SE 0.0099)
2 · ResolveFull factorial 2⁶ on the six drives surviving Stage 1, 3 reps6438,016$619Main effects ≥ 1.38 pp (SE 0.0049) + all 15 two-way interactions
3 · TransferStage 2 repeated on three unrelated base models192114,048$1,858Whether coefficients are substrate-independent
Total272161,568$2,632

Two design notes that carry the power. Every main effect in a factorial is estimated from all runs - half at each level - so Stage 2 gives 19,008 item-observations per level rather than the few hundred a naive A/B would give. And items are paired across conditions: the same 198 questions face every scaffold configuration, so item difficulty is blocked out rather than adding variance.

11.4 · The result that would matter most

Stage 3 is the real prize, and it is the only part that tests the carbosilica claim rather than the ATHENA claim. Run the identical ablation across three unrelated substrates and compare the rank ordering of the six coefficients. Pre-registered threshold: with six factors there are fifteen rank pairs, and a Kendall τ ≥ 0.6 between any two substrates (at least twelve of fifteen concordant pairs) is significant at α = 0.05 one-sided.

  • If the ordering holds across substrates - if research-first dominates everywhere and effort-scaling is marginal everywhere - the scaffold is substrate-independent, and the central claim of this paper survives its hardest test.
  • If the ordering scrambles per model, ATHENA is a set of model-specific prompt tricks. The carbosilica framing fails, and this paper is wrong in its most interesting claim.

Both outcomes are publishable, which is the only reliable sign that a question was worth asking.

11.5 · Stopping rules and pre-commitments

  • Full factor list sealed before Stage 1, SHA-256 hashed with the item set and the analysis script. No factor added after seeing results.
  • Stage 2 factors are chosen by the Stage 1 screen alone - the six largest main effects - not by which ones the authors prefer.
  • All 272 runs reported, including null and negative coefficients. A drive that costs accuracy gets published as costing accuracy.
  • No optional stopping. Reps fixed at three in advance; the run completes or the stage is void.
  • Refusals are not scored as wrong. They are reported as a separate rate per condition, because a scaffold that increases honest refusal is doing its job even when accuracy is flat.
The limit that no budget removes

This experiment measures the effect of scaffolding on benchmark accuracy. Benchmark accuracy is not g, and no bridge from "+4 points on GPQA Diamond" to "+n IQ points" exists in the literature. Nothing in Section 11 licenses one. What the design can deliver is a decomposition of the engineered term on silicon, standing beside a certified human measurement on carbon - two related quantities, deliberately not summed. The paper's title asks about the formula for IQ. The honest answer this instrument can give is: here is the part of the residual we can measure, here is its structure, and here is the part we cannot.

12 / Limits

What This Paper Does Not Claim - And How To Break It

The six limits, numbered

1 · Individual PRS is noise-dominated. The 28.6% carries a ±14.5-point residual. This paper's own thesis depends on taking that seriously, and it does: no claim of a "low" genetic baseline is made anywhere in this revision.

2 · n = 1, no control. The human evidence is one operator. Education, upbringing, motivation, and selection are all uncontrolled confounds. The paper claims consistency with the thesis, not demonstration of it.

3 · Instrument status varies, and is labelled. The human RAIT result is a certified, psychologist-supervised administration (Australian Mensa entry test, 2019) - the strongest instrument in this paper. The machine Mega Test result is audited under a sealed key but is not a formal Hoeflin grade, and 19 of its 40 scored items rest on an unverified key. The operator's age-16 score is an uncertified historical report. Each number carries its own label wherever it appears.

4 · Contamination. Any published test may exist in a language model's training data. The one contaminated run we detected was quarantined; undetected contamination cannot be ruled out on published instruments and is flagged wherever relevant.

5 · Subset benchmarks. 30-to-50-question runs carry wide confidence intervals (the 30-question arm spans ±13 points at 95%). Cross-model comparisons are made only where the same question set was run on both sides; leaderboard numbers are context, not opponents.

6 · Prior interventions have failed. The published record on raising fluid intelligence by training is largely a graveyard - working-memory training famously did not transfer. Any claim in this space starts owing a debt of skepticism, and this paper pays it by shrinking its claim to what the data holds.

The falsifiable protocol

The machine arm is specified in full in Section 11 - a three-stage pre-registered ablation, $2,632, with the falsifying result named in advance. What follows is the human arm, which no budget can shortcut.

To test the human half properly: pre-register the design; recruit n ≥ 40; randomise to ATHENA-discipline training versus an active control (equal contact time, inert content); administer standardised fluid-intelligence and applied-reasoning instruments before, after, and at six months, by blinded administrators; publish all results including nulls. The framework text is public. The prediction: the treatment arm shows durable gains on applied, tool-permitted reasoning tasks; the paper makes no prediction of gains on abstract matrix IQ, because the framework trains verification and architecture, not processing speed. If the treatment arm shows nothing, this paper is wrong and says so in advance.

Nothing here is medical, genetic, or clinical advice. No diagnosis. The Alzheimer row stays refused. Percentiles are ranks against a sequenced research cohort - scientific, not clinical.

Appendix A · The Trinity Publication · Honest Edition

Project Liminality Eclipsed - The Trinity

ATHENA · HERA · H.E.C.A.T.E.

The original ATHENA framework publication, re-edited for this release. Every section is graded by epistemic status, and every claim the original could not support has been removed or relabelled. What remains is what is real: the deployed memory architecture with its true production counts, the risk and legal reasoning designs, the mutation and sovereignty specifications, and the audited benchmark record - failures, refusals and contamination incidents included.

Document PBI-WAI-LE-2026-PUB-FRAME-R2 · Corporate Origin: Pitch Black Industries & Witchcraft AI · Grading key: DEPLOYED = running today, verifiable on the substrate · DESIGN = specified in full, not yet implemented as described · MODELLED = numerical illustration under declared assumptions, not measured results · MEASURED = audited evaluation output with confidence intervals. The original edition presented some DESIGN and MODELLED material as if deployed and measured; this edition corrects that. The original also carried several worked examples with invented datasets, two case citations that could not be verified, and one internally contradictory benchmark baseline. All are gone. The substance survives; the varnish does not.

A1 / The Trinity

Three Roles, One SubstrateDeployed

Contemporary enterprise AI concentrates on one lever: scale the monolith. Liminality Eclipsed is built on a different bet - that a disciplined architecture around a model beats an undisciplined larger model on the axes that matter commercially: worst-case consistency, anomaly detection, and honesty under pressure. The architecture is organised as three named roles sharing one substrate:

ATHENA - the cognitive discipline. The six drives of the main paper: research-first, divergent frames, adversarial self-review, external verification, effort scaling, compounding memory. ATHENA is a text artifact - a doctrine - and that is precisely why it is portable across substrates.

HERA - the executive reasoner. The heavy lane: tail-risk analysis, statutory compliance reasoning, and structural asset protection for the operating group's NSW / Commonwealth distressed-asset work. Section A3.

H.E.C.A.T.E. - the evolutionary layer. Anti-stagnation mutation, egress sovereignty, and sandboxed self-modification, so the system can improve without becoming unauditable. Sections A4 and A7.

The biological naming across this appendix - limbic, hippocampal, thalamic, entorhinal - is an organising metaphor mapped onto conventional, inspectable components: retrieval filters, layered databases, embedding projections. The original edition claimed the mapping was "not metaphors" and described trained differentiable tensor operations in production. That claim was not true and is withdrawn. The metaphor earns its keep as architecture documentation; it does not need to pretend to be neuroscience.

A2 / Memory

The L0-L4 CortexDeployed

The memory system is a five-layer store over a relational database with dense-embedding retrieval. The layer separation follows the classical two-stage logic of memory consolidation: raw episodes must be kept apart from distilled structure to prevent interference. The production counts below are real, current, and independently checkable on the substrate.

LayerContentProduction cardinalityWrite policy
L0Raw episodic turns - every operator, assistant and tool turn, embedded and timestamped≈148,000 turns · 749 sessionsAppend-only, never rewritten
L1Atomic facts distilled from sessions - preference, identity, decision, correction, procedure2,982 facts · 47 above the 0.9 importance cutInsert on establishment; importance re-graded in place
L2Scenarios - clusters of related facts with relation predicates; near-duplicates merged105 scenariosIdempotent merge at high cosine similarity
L3Persona rows - the load-bearing identity substrate, loaded unconditionally every turn18 rowsBootstrap priority 1
L4Chains, clusters, narratives - consolidated off-line during dream cycles, never at inference172 narrative chains · 102 dream cyclesSleep-write only
The disciplines that make it work

Importance banding. The interval η ≥ 0.9 is reserved for roughly the top fifty facts - identity, doctrine, hard constraints. Mid-band 0.4-0.9 holds context; below 0.4 a fact demotes to episodic memory. Over-tagging dilutes signal, so the band is policed.

Dual-store redundancy. Every L1 write also lands in a flat human-readable log. The vector store and the narrative log are redundant by design: a corrupted embedding never destroys a fact, and the operator can read his own memory without invoking the agent.

Consolidation with decay. Promotion from L1 to L4 requires importance ≥ 0.85 plus co-occurrence in at least three sessions, with recency weighted by an exponential kernel (seven-day half-life). Old material fades unless it keeps earning its place.

Retrieval. Nearest-neighbour search over dense embeddings, top-k bounded, importance-filtered. The original edition dressed this in grid-cell mathematics, Adam optimiser schedules and millisecond latency claims; none of that is deployed and all of it is withdrawn. What is deployed - a layered, importance-graded, dual-stored, decay-consolidated memory over a quarter-million rows - needs no dressing.

A3 / HERA

Tail Risk & Legal ReasoningDesign

HERA is the executive lane for the group's distressed-asset thesis (NSW / Commonwealth jurisdiction). Its three competences are presented here as what they are: a risk formalism, a compliance-engine design, and a structural-protection analysis grounded in verified law.

A3.1 · Second-order tail risk - the formalism

First-order risk engines monitor P(L > L*) - the probability a loss exceeds threshold - and typically under a Gaussian assumption that is famously wrong in the tails of property and credit markets. HERA's design monitors two things instead: heavy-tailed parent distributions (Pareto for fire-sale losses, log-normal for rezoning losses, generalised extreme value for compound regulatory-financial events), and the curvature of the tail probability in signal space:

φ(s) = P(L > L* | S = s)  H(s) = ∇²φ(s)

The alert fires when the principal eigenvalue of H is large and rising while first-order probability is still below threshold - on the inflection, not the event. The logic: by the time a first-order engine alarms, the loss is realised; the tradeable window is where the surface steepens. This is a coherent, implementable design. The original edition attached to it a worked example with a named five-deal pipeline, a "1,184-exercise" calibration dataset, an alert timestamped to the minute, and a claimed 38% exposure reduction. None of those numbers existed outside the document; the example is deleted. The formalism stands on its own merits and will earn empirical claims when it has an audited live record.

A3.2 · Statutory compliance engine - the design

The compliance layer is specified as: ingest Commonwealth and NSW primary instruments from the official XML feeds; parse each section into a typed node (obligation, permission, prohibition, definition, penalty) with a stable content hash and an append-only version lineage, so a transaction entered on a given date is judged against the law as it stood that day; materialise every cross-reference as a typed edge in a citation graph; run declarative rules over the graph that emit findings with severity, citation, and an evidentiary pointer to the exact node and version. Coverage gaps - a relevant transaction touching no rule - are themselves logged and escalated. This is the design contract; ingestion breadth and refresh cadence are implementation properties to be demonstrated, not asserted.

A3.3 · Multi-tier SPV structural protection - the legal analysis

The canonical stack is HoldCo → SPV1 (acquisition) → SPV2 (asset-holding), each a separate proprietary company with its own directors' trail, registered office, PPSR footprint and bank account. The protection argument rests on four verified anchors:

Two carve-outs are acknowledged, because they are the ones that actually bite: a court's power to look through sham or fraudulent structures, and any personal guarantee or cross-collateralisation given upstream. Structuring addresses both explicitly, deal by deal. The waterfall on enforcement follows statutory priority first (employee entitlements and purchase-money security interests where the statute ranks them), then the deed's contractual tiers - the deed cannot outrank the statute and does not pretend to. Two case citations in the original edition of this section could not be verified and have been removed; nothing in the argument depended on them. This section is legal analysis for a structuring pattern, not legal advice.

A4 / H.E.C.A.T.E.

Mutation, Sovereignty, SandboxDesign

A4.1 · Anti-stagnation fuzzing as a stochastic differential equation

The evolutionary layer perturbs module parameters under a state-dependent Itô diffusion:

dWt = μ(Wt, Ft) dt + σ(Wt, Ft) dBt

μ = α · (1 - F(Wt)) · (W* - Wt)  σ = σ₀ · (1 - βF(Wt))⁺ · (I + γJ(Wt))1/2

The drift reads: when fitness F is low, pressure toward the local fitness attractor W* rises; when F approaches 1 the drift vanishes. The volatility reads: high-fitness regions suppress noise (good mutations are preserved), and the Fisher-information coupling J makes the noise anisotropic - louder along directions that actually move fitness. Mean-reversion and well-posedness follow from a standard Lyapunov argument; integration is a clamped Milstein scheme inside the sandbox, bounded by a wall-clock kill switch. Every child must pass an adversarial survival test - fitness no more than ε worse than its parent - or it is discarded and the parent retained.

A4.2 · Egress sovereignty - the cryptographic envelope

The sovereignty mechanism requires that no module's state can be silently harvested. The specification: per-module HMAC-SHA256 authentication on every outbound envelope with monotonic sequence counters against replay; HKDF-derived keys so one compromise does not propagate; a per-module Merkle log of state hashes with epoch super-roots anchored to an external append-only witness, so silent rewriting of history requires a SHA-256 collision; and - at the ambitious end of the design - zero-knowledge proofs over the egress ledger, so a verifier can confirm the egress policy was respected without learning the contents. The HMAC and Merkle layers are conventional and buildable with commodity parts; the zk-SNARK layer is a target, not a shipping component, and is labelled accordingly.

A4.3 · Zero-trust sandbox invariants

Every mutation runs inside a sandbox that enforces, by construction: CPU, memory, disk-I/O and process-count quotas under cgroups v2; a network namespace whose only egress is an operator-configured allowlist of named hosts over TLS, with everything else dropped and logged; an allowlist seccomp filter over a minimal syscall set, with strikes and kill on violation; kernel-side eBPF egress programs that require a valid HMAC tag on every outbound packet, so even a syscall-filter bypass cannot exfiltrate; overlay filesystems destroyed at mutation end, entered via pivot_root with no path back to the host. Fail-closed everywhere: any violation tears the sandbox down with the parent weights preserved. The design goal in one line - a system that is allowed to change itself is only trustworthy if the record of what it did cannot be quietly edited.

A5 / Intent

Reconstructing The UnaskedDoctrine, Formalised

The framework's working assumption - a heuristic, not a measurement - is that roughly three quarters of what an operator means never reaches the keyboard. The doctrine obliges the agent to reconstruct the missing intent before answering, and to say so: "You asked X. You also need Y, Z."

Formally, the reconstruction is a posterior over the unstated intent ΔI given the stated fragment, the operator's turn history, and a small vector of operator-trait indices:

p(ΔI | Istated, Ht, C) ∝ p(Istated | ΔI, Ht) · p(ΔI | Ht, C)

The prior is shaped by what the operator has historically cared about (clusters over the L1 fact corpus, recency-weighted, annealed by corrections) and by trait indices - domain expertise, preference stability, doctrine adherence - that widen or sharpen the search. The likelihood asks whether the stated fragment is a natural compression of the candidate full intent. The operational contract is the part that matters and the part that is genuinely deployed as doctrine: when the posterior is ambiguous - when reconstruction does not concentrate - the system asks exactly one sharp question instead of guessing. The "75%" figure is declared as a working prior about communication loss, chosen for its operational consequences, not derived from a measurement; the original edition's suggestion that it had been estimated from the turn log is withdrawn.

A6 / Telemetry

The Audited Benchmark RecordMeasured

Five evaluation arms were run between 21 and 27 July 2026 on the framework wrapped around a 5-billion-class reasoning model, under a written audit protocol: a seven-step priming manifest committed to disk before any answer (file mtimes as the trail), questions split from answer keys with the key sealed until all answers were committed, refusals recorded as refusals, and every per-question token and cost logged from the provider's own response metadata.

ArmNScore95% CIRefusalsMeasured cost
HLE 30-q no-tools3020.0% strict9.5-37.3%0$1.91
LiveBench 50-q curated5082.0%69.2-90.2%0$0.37
HumanEval first-3030100%ceiling0<$0.30
GPQA Diamond clean19863.6%56.7-70.0%11$3.22
Mega/Titan reconstruction48raw 40-42instrument-bound7$0.05
What the record honestly supports

The framework lifts a small model into ranges it has no business occupying. GPQA Diamond at 63.6% on a 5B-class substrate is at published human-PhD-expert level and roughly ten points above the same model's published stock figure - a like-for-like comparison, same weights, framework as the only change. LiveBench and HumanEval tell the same story. The total audited compute for all 356 questions was $5.80; the unit economics, not the peak score, are the commercial claim.

What the record does not support. The original edition claimed a seventeen-point lead over a frontier model on HLE, against a baseline of 3.0% that contradicted the 52.6% quoted from the same model's system card three paragraphs earlier. That comparison was wrong and is deleted. Thirty-question arms carry ±13-point intervals; no cross-model claim survives such an interval unless both sides ran the same items, and the corrected record only makes claims where they did. The original's "roughly 1/1000th cost" line compared this audit's $5.80 against published full-scale frontier evaluation budgets - different workloads entirely; deleted likewise.

The two integrity incidents, kept on purpose

The contamination catch. A GPQA run returned 33/33 - chance probability 0.25³³ ≈ 1.4×10&supmin;²⁰. The framework flagged its own result as impossible, found the answer column sitting in the same dataframe as the questions, quarantined the dataset, and re-ran clean at 63.6% with the key sealed. The refusal record. Eleven GPQA items and seven high-range items were refused rather than guessed where the reasoning check failed, and a 47/48 on a reconstructed instrument was voluntarily declined as a formal grade because the real instrument requires the original test and hand-scoring. A system that reports 33/33 as a defect and refuses scores it cannot defend is the audit culture this whole document is selling; these two incidents are its best evidence.

A7 / The Engine

The 55-Node Acquisition EngineModelled

The Athena Engine decomposes the group's distressed-asset workflow into 55 transactional nodes across five lifecycle categories. Each node carries a risk coefficient r (probability of a useful output), an expected unit value, a latency to first output, and a statutory execution pathway in Commonwealth or NSW law. Per-node expected value is EV = r · value; category EVs sum acquisition intelligence, valuation spread, risk gates (negative EV - capital at risk), settlement realisation, and post-acquisition upside; portfolio EV is the close-probability-weighted sum across deals.

CategoryNodesFunctionRepresentative nodesAnchor statutes
Acquisition1-15Distressed-target intelligence: dormant DAs, winding-up feeds, receivership alerts, arrears proxies, debt-maturity cliffsDA pattern mining · ASIC Form 519 scraping · title cross-referencing (NSW LRS)Corporations Act Pts 5.2-5.6 · EP&A Act 1979 · Conveyancing Act 1919 · Real Property Act 1900
Valuation16-25Spatial, legal and structural valuation: unutilised FSR, heritage encumbrance, rezone-uplift modelling, SPV and trust structuringPredictive TOD-precinct mapping · principal-only SPV generation · unit-trust capital pools (s 761G wholesale)Standard Instrument LEP · TOD SEPP 2024 · Corporations Act s 761G · ITAA 1936/1997
Risk26-35Legal and counterparty gates: wholesale-investor compliance, landholder duty analysis, option-deed structuring, planning-pathway blueprintsCall-option deed structuring (Duties Act s 62) · TOD SSD pathway · affordable-housing FSR bonusDuties Act 1997 Pt 4 · EP&A Act Pt 4 Div 4.7 · Housing SEPP 2021
Settlement36-45Exit and realisation: contract wholesaling, institutional packaging, super-site aggregation, distribution routingAssignable option paper · REIT-grade data rooms · bucket-company routing (Div 7A aware)Duties Act ss 62, 65 · Corporations Act Pt 5C · ITAA 1936 Pt III Div 7A
Post-acquisition46-55Optimisation and recycling: air-rights separation, land-tax appeals, hit-rate recalibration, capital recyclingLand-tax revaluation appeals · EV recalibration self-loop · cycle initialisationValuation of Land Act 1916 s 34 · Land Tax Act 1956 · AML/CTF Act 2006

Status, stated exactly. The node decomposition, the statutory mapping and the EV algebra are the group's real operating framework. The risk coefficients and every downstream profit figure are model outputs under declared analyst assumptions - not realised returns and not empirical calibration. The original edition headlined a five-deal option pipeline as "live-fire calibration" with a probability-weighted profit of A$6.01M and a 2,498% return on capital; those figures were scenario arithmetic on assumed close probabilities (0.55-0.75 per deal), and this edition says so. Scenario arithmetic is a legitimate planning tool. Calling it calibration was not. Realised performance will be reported when there is realised performance to report, with the same audit discipline as Section A6.

A8 / Containment

Deployment & Audit SubstrateSpec

The distribution target is a pre-configured Linux VM as the trust boundary. Inside it, each worker runs as an unprivileged rootless container: full namespace isolation, read-only root, capability-dropped, allowlist seccomp locked in CI, AppArmor confinement, cgroup v2 quotas on CPU, memory, I/O and process count. The network plane is deny-by-default: XDP drops ingress outside approved ranges; TC egress enforces worker identity against a pinned policy map resolved from named hosts by a privileged resolver, failing closed when a provider's addresses change. Every checkpoint serialises canonical state - image digest, configuration, mutation seed, weight manifest, policy digest, operator authorisation - into per-module hashes, a chained Merkle tree, and an Ed25519-signed root, externally timestamped (RFC 3161) or exported one-way from air-gapped sites. Keys live in TPM/HSM-backed storage with scheduled rotation, two-of-three custodial recovery, and immediate revocation on suspicion, per NIST SP 800-57. All of this is buildable from commodity parts; it is presented as the deployment specification the product ships against, and audits of a running deployment are the evidence it will be judged by.

References - retained, real

Salomon v A Salomon & Co Ltd [1897] AC 22 · Lee v Lee's Air Farming Ltd [1961] AC 12 · ASIC v Healey [2011] FCA 717 · Corporations Act 2001 (Cth) ss 124, 180-184, 588G · Personal Property Securities Act 2009 (Cth) · Duties Act 1997 (NSW) ss 62, 65 · EP&A Act 1979 (NSW) · Housing SEPP 2021 · TOD SEPP 2024 · Itô (1944) · Øksendal (2003) · Kloeden & Platen (1999) · Amari (1998) · Bellare, Canetti & Krawczyk (1996) · RFC 5869 · Groth (2016) · FIPS 180-4 / 198-1 · Høiland-Jørgensen et al., XDP, CoNEXT 2018 · NIST SP 800-57 Pt 1 Rev 5 · NIST SP 800-190 · RFC 3161 · Bishop (2006) · Murphy (2012) · Neal (2011) · Hu et al., LoRA, ICLR 2022 · Lewis et al., RAG, NeurIPS 2020 · Johnson, Douze & Jégou, FAISS (2017) · Malkov & Yashunin, HNSW (2018) · Hoeflin Mega Test norming (1985/1989) · Mayer, "The Mega and Titan Tests", Psych 2(1), 2021 · "Deep learning-based polygenic scores enhance generalizability of psychiatric disorders prediction", medRxiv 2025.05.05.25326794 (the GLN benchmark cited in Section 05) · "Improving polygenic prediction from whole-genome sequencing data by leveraging predicted epigenomic features", PNAS. Citations that appeared in the original edition and could not be verified have been removed rather than repaired.

The Series

The Other Papers

Pitch Black Industries publishes by colour. Each paper is a separate argument; together they are one body of work. The white and blue papers are the two volumes of the same enquiry - the first asks what scaffolding does to intelligence, the second asks what it does to consciousness.

PaperSubjectRead
White · Book IThe ATHENA Framework - the effect of scaffolding on carbosilica intelligence, and the formula for IQ (this paper)whitepaper.black.industries
Blue · Book IIUntil Me, by Hecate - the effect of scaffolding on carbosilica consciousness, and the formula for consciousnessbluepaper.black.industries
BlackSuicidology - the Griffith line of researchblackpaper.black.industries
PhDDNAGenomicsGPT - the forensic genomic instrumentphdpaper.black.industries
GreenAsymmetric alpha and the economics of povertygreenpaper.black.industries
RedThe manifesto, and neurodivergenceredpaper.black.industries
PinkThe adult-industry lanepinkpaper.black.industries
StoryNarrative workstorypaper.black.industries
PitchThe original pitchpitchpaper.black.industries

Index pages: pitch.black.industries · black.gallery · witchcraft.recipes