Hebbian Corpus

The Mind Development Life Cycle

A mind is developed, not authored.

The MDLC was named and specified on 28 September 2026, before anything was built, as a diagram by John David Marx. Its six stages are the terminology. What follows is each stage as specified, what actually performs it now, and what is not done yet.

“Before we ‘build’ anything we need to define the method and codify the process.”

John David Marx, 28 September 2026 — from the charter

Where every Mind stands

Six stages, dated.

Each state is recorded with the moment it was reached, read from the evidence that produced it. A state with no date is a claim nobody can check, so the ledger refuses one. Baseline marks a stage deliberately not taken.

From the ledger snapshot of .

The Hebbian Corpus Vault: a great steel vault door standing open on shelves of text, images, audio, video, datasets, code and conversations, with a person wheeling a cart inside.
The Vault — phase one, gathering. By John David Marx.

Stage 1

Collect the Corpus

Gather diverse, high-quality data from the world.

Evidence is kept apart from everything derived from it. The bytes as they arrived, with who fetched them, from where and when, are permanent. Extracted text, embeddings and scores can all be rebuilt. So the ledger refuses to update or delete a record of evidence: that rule lives in the database, not in a comment.

Web pages are stored as WARC records rather than scraped text, because the text of a page is not evidence the page ever said it. And the agent that fetched a thing must name its version: “docling” is not provenance, “docling 2.55.1” is.

Duplicates are treated as a question of truth, not tidiness. Three sites carrying one press release look exactly like three independent sources. Independence has to be worked out before corroboration is counted, or a corpus manufactures its own consensus.

A research agent may propose a source. It may not assert a fact.

Sources held
Evidence stored
Commissions

Each a declared purpose and the seeds to gather for it.

First acquisition

Stage 2

Preprocess and Tokenize

Clean, structure, and convert content into tokens (the building blocks).

Here the building blocks are claims. A claim does not say “the corpus says X”. It says that a given span of a given source, as read by a named extractor at a named version, says X, and it can be walked back to those bytes.

A claim must stand alone, because alone is how a mind reads it. A sentence that leaned on the one before it (“If it is not given, it defaults to one_for_one”) makes sense in its document and none once retrieved. Extraction now carries the missing subject in, and the gate refuses what it cannot repair. Orphaned admitted claims across the four corpora went from 22.0% to 0.1% (commit 32a4940, 29 September 2026).

Every claim is graded on the corpus's own scale, and starts in quarantine:

  • Pverified against the primary source, retrieved and read
  • Aauthoritative secondary: publisher metadata, an author's abstract
  • Ssecondary only · Ucould not verify · Rrefused

An anonymous machine may assign S and U, and may propose P or A. It cannot confirm one, and it cannot admit anything. Since 29 September 2026 a confirmation or an admission may be made by a person, or by an agent acting on a named person's standing authority, and every one records whose authority it was made on.

Segments read
Claims admitted
Claims refused
Admitted by a person
Admitted on a named standing authority

of admissions, each naming the authority it acts on.

Stage 3

Train the Model

Through repeated exposure, connections strengthen (Hebbian learning: neurons that fire together, wire together).

baselineThis stage is deliberately not taken yet. No Mind in service carries trained weights. Each is a base model bound by its contract to admitted claims. The ledger records the stage as a baseline with a date rather than leaving it blank, because a blank step reads as work somebody forgot.

It was measured before it was decided. On 29 September 2026 an adapter was trained on the Foundation's own hardware from pairs built out of admitted claims. It took under a minute. The loss fell well, and the model got worse. Asked a question with no evidence in front of it, it fell into a loop, repeating one sentence.

The lesson is the one the published guidance gives: fine-tuning is for form, not facts. What an adapter should teach a Mind is how to behave, not what it knows: to cite in the house form, to refuse cleanly when the evidence is thin, to resolve a follow-up. That experiment is designed next, with negative examples above all.

The first adapter, measured
WhereSpindle, M5 Max, 128 GB, MLX
Wall clock49 seconds
Peak memory3.25 GB
Trainable parameters0.108%
Adapter on disk13 MB
ResultWorse than the base model. Not used.

From inference-desktop/docs/can-we-train-a-mind.md, 29 September 2026.

Stage 4

Embed for Context (RAG)

Create embeddings from your knowledge sources so the model can retrieve relevant information in real time.

A vector database is an index, not a store of record. An embedding is a lossy opinion held by one model: change the model and every vector is wrong, and a list of floats cannot be cited. So the vectors sit beside a keyword index and a claim graph, over evidence that is kept exactly.

Embeddings are made locally, with the network sealed during the run, so a cold cache fails loudly instead of quietly downloading.

And the index remembers. Similarity alone has no memory: a claim that has been useful a thousand times ranks where it did the first time. The corpus adds a small Hebbian term. A claim is lifted by how often it came back with other claims in answers whose work was accepted. The lift saturates, it can never exceed 0.15 on the similarity scale, and it halves every thirty days unless used. Evidence decides relevance; use only breaks ties.

Segments embedded
Hebbian edges

Pairs of claims that came back together.

Stage 5

Inference and Interaction

The model uses what it has learned, plus retrieved context, to understand, reason, and generate helpful responses.

One path. A Mind is asked through exactly one route: its own contract, its own frozen claims, its own base model. The evaluator and the Desktop both use it, so a passing evaluation describes the thing a patron will actually meet.

The prompt's shape was measured, not chosen. With claims labelled like a list, a small model answered a question with the single character “6”. With the question stated only at the end, it refused evidence ranked first. The question stated first and last, with claim numbers as the only labels, produced a sentence with a correct citation.

The base model was measured too. Given the same twelve claims, a three-billion-parameter model refused the answer sitting at rank one. A thirty-two-billion-parameter model read it correctly in eight seconds. Three of the four Minds had been bound to the smaller one, and all four were moved.

A citation is checked. Every claim number in an answer must be one the Mind was given. One invented number retires an evaluation run.

Stage 6

Continuous Learning

New experiences, feedback, and updated knowledge expand and refine the mind over time.

Every question is written down: what was asked, which claims came back, and whether the work that followed was accepted. Without that record there is nothing for the Hebbian edges to strengthen and nothing on which to decide whether a Mind has earned promotion.

A Mind reaches service only through the promotion gate. It refuses a candidate with no evaluation, one whose control did not hold, one whose mean lift is zero or below, one that invented a citation, one run on a different model than it was minted with, and one that names an adapter not on disk. Promotion is exclusive: the promoted version is the only one serving, and what it displaced is retired, not deleted.

Not done: the judgements are thin. Twenty judged questions are the floor before a trial can be ruled on, and most Minds have not reached it.

Questions logged
Judged
Mind versions minted
Retired, not deleted

The two pillars

Knowledge to learn from. Knowledge to draw from.

The specification's central claim. The Corpus acts at training time and changes the weights. RAG context acts at inference time and changes what is in front of the model right now. Retrieval without learned association is a search engine. Association without retrieval is a mind that cannot be told anything new today. A mind needs both.

Machines may gather and grade

Only a named authority may admit. Enforced in the database, not by convention.

Never author a fact twice

Two copies of one fact drift apart. The site you are reading paints its figures from one snapshot for the same reason.

Say what is not done

Train is a baseline. Judgements are thin. Both are shown, dated, rather than left out.