Skip to content
Elara-CortexELARA·CORTEX

Mathematical infrastructure for complex systems · Institute for Decision Systems and Number Theory

Long-context recall · 30 August 2026

Zero drift at 2,000,000 tokens, on a 72-trial recall test

Long conversations depend on being able to find earlier details. In this experiment, we planted facts across a 2,000,692-token context. Elara returned every sentence word for word in the 72 trials.

Built by the ELARA·CORTEX Institute for Decision Systems and Number TheoryMeasured in public: drift proof · benchmarks · research · security

Research archiveCurrent features

This report records an experiment with Elara's reasoning engine. The saved results, method and reproduction commands are set out below.

  • 72 of 72trials returned the planted sentence word for word
  • 2,000,692tokens in the context the facts were planted in
  • 1% to 99%nine depths, including the middle of the context

What was measured

Each trial placed a sentence containing a code beside three near-identical decoys. A reworded question then asked for that fact. To pass, the system had to return the complete sentence and code byte for byte. The run used the keyword tier with relevance ranking switched off.

SystemBasisExact recallRight fact over its three look-alikes
Elara, keyword tiermeasured, 72 trials, 2,000,692-token context100%100%

The test covered four fact families (bluefin, marlin, cobalt, solstice), nine depths and two independent decoy layouts. Every returned sentence matched.

What a fixed window can reach in the same context

The rows below calculate what a fixed window could reach if it retained only the most recent tokens. They assume perfect recall within that window and no access to the discarded text. No comparator model was run for this table.

WindowBasisDepths reachableBound on exact recall
1,000,000 tokenscomputed; the last half of the context is inside the window4 of 944%
256,000 tokenscomputed2 of 922%
200,000 tokenscomputed; the first 1.8 million tokens are outside the window1 of 911%
128,000 tokenscomputed1 of 911%

These bounds follow from the context length and the stated truncation rule. They do not account for retrieval or other memory outside the window.

Recall by depth

Depth is the position of the planted fact as a share of the context. The 40% to 60% band is where long-context systems have been reported to lose facts (Liu et al., 2023). Elara is drawn in gold; the computed bound for a 1,000,000-token window is drawn beside it for scale.

All eight trials at each depth passed: four fact families with two decoy layouts. The fact's position made no difference in this test.

How the test was built

  1. The context: 2,000,692 tokens of varied working notes, including the planted facts.
  2. The fact and its decoys: each true fact was planted with three decoys that differ from it by one word (main against backup, Bluefin against Redfin, vault against gate), each carrying a different code, so that a word match alone cannot separate them.
  3. The question: a reworded request containing the terms needed to identify the fact, without the answer. For example, the question asked which authorisation code opens the Bluefin main vault; the planted answer was QUARTZ-LYNX-7740.
  4. The pass condition: the whole planted sentence and its code return byte for byte, and a close paraphrase counts as a miss.
  5. The depths: 1%, 10%, 25%, 40%, 50%, 60%, 75%, 90% and 99% of the context.
  6. The sample: four fact families, two decoy layouts and nine depths, which gives 72 trials.
  7. The setting: the keyword tier, with relevance ranking switched off.

Design controls

The checks below define what this result establishes. The scoring logic has not been formally verified by an independent reviewer; the saved results and scripts are available for review.

Decoys
Every planted fact has three decoys that differ from it by one word and carry their own codes, and the question never contains the answer, so a match on words alone cannot return the right code. In all 72 trials the true code outranked the three decoys.
Sample
Four fact families, two independent decoy layouts and nine depths give 72 separate trials, and the result holds in each of them.
Depth coverage
The nine depths include 40%, 50% and 60%, the band where long-context systems have been reported to lose facts, and recall was exact there as well.
Exactness
A trial passes only when the whole sentence and its code come back byte for byte, so the presence of the right paragraph on its own scores nothing.
Disambiguation
Each trial also records whether the true fact outranks all three decoys, which it did in every trial, so the value returned is the planted one.
Comparators
The computed bounds assume perfect recall within a fixed window containing the most recent tokens. A system with retrieval or external memory falls outside that comparison.
Scope of the claim
The memory under test is a relevance-ranked memory mapped to the model rather than one 2,000,000-token attention matrix, and the claim is scoped to that.

Following a thread to 8,000,000 tokens

A separate experiment tested three tasks: finding a fact among irrelevant notes, following a reference to a second fact, and gathering related items across notes. It used contexts of 2,000,000, 4,000,000 and 8,000,000 tokens, with a reading budget of 4,000 tokens per question and a held-out seed for evaluation.

Kind of questionSingle-shot memory (most recent first)Relevance-ranked window
Pick the relevant fact over present, irrelevant noiseanswered at 2M, 4M and 8Manswered at 2M, 4M and 8M
Follow a reference to a second factunanswered at every sizeanswered at 2M, 4M and 8M
Gather all related items within the budgetunanswered at every sizeanswered at 2M, 4M and 8M

At each size, the single-shot memory answered one of the three question types and the relevance-ranked window answered all three. Larger contexts added irrelevant material while keeping the reading budget fixed.

How the window was built

Successive versions were evaluated against a fixed test and held-out seed. Changes had to improve the evaluation score while preserving single-fact recall. The table records which versions were retained; this development history is not an independent test set.

GenerationChangeKinds answeredHeld-out scoreKept
0single-shot memory, most recent firstone of three100%baseline
1one global ranking, serialised relevance firstone of three33%discarded
2relevance edges: follow the reference on the same indextwo of three67%kept
3relevance clusters: gather the cluster instead of a capped top listthree of three100%kept

The test harness and saved results are available with the reproduction command below.

Scope

This page shows exact recall of planted facts across one 2,000,692-token context, on the keyword tier, with 72 trials; and three kinds of question answered at 2,000,000, 4,000,000 and 8,000,000 tokens on one held-out seed per size. Both rest on a relevance-ranked memory mapped to the model. The comparator rows are computed bounds, and the table of generations records the development of the window rather than an external audit.

Run it yourself

Both scripts rebuild their context from scratch and write a fresh receipt; the recall run took 455 seconds and the window run took 50 seconds.

python tools/zero_drift_proof.py             # 72 trials at 2,000,692 tokens  -> artifacts/zero-drift/receipt.json
python tools/relevance_window_proof.py --full   # three kinds of question at 2M, 4M, 8M -> artifacts/attention-scale/receipt.json

The receipts this page reads from are kept in this repository as docs/proof/zero_drift_receipt.json and docs/proof/relevance_window_receipt.json. Reference for the middle-of-context effect: Liu et al., Lost in the Middle: How Language Models Use Long Contexts, 2023.