Mathematical infrastructure for complex systems · Institute for Decision Systems and Number Theory
Long-context recall · 30 August 2026
Zero drift at 2,000,000 tokens, on a 72-trial recall test
Long conversations depend on being able to find earlier details. In this experiment, we planted facts across a 2,000,692-token context. Elara returned every sentence word for word in the 72 trials.
Built by the ELARA·CORTEX Institute for Decision Systems and Number TheoryMeasured in public: drift proof · benchmarks · research · security
Research archiveCurrent features
This report records an experiment with Elara's reasoning engine. The saved results, method and reproduction commands are set out below.
- 72 of 72trials returned the planted sentence word for word
- 2,000,692tokens in the context the facts were planted in
- 1% to 99%nine depths, including the middle of the context
What was measured
Each trial placed a sentence containing a code beside three near-identical decoys. A reworded question then asked for that fact. To pass, the system had to return the complete sentence and code byte for byte. The run used the keyword tier with relevance ranking switched off.
| System | Basis | Exact recall | Right fact over its three look-alikes |
|---|---|---|---|
| Elara, keyword tier | measured, 72 trials, 2,000,692-token context | 100% | 100% |
The test covered four fact families (bluefin, marlin, cobalt, solstice), nine depths and two independent decoy layouts. Every returned sentence matched.
What a fixed window can reach in the same context
The rows below calculate what a fixed window could reach if it retained only the most recent tokens. They assume perfect recall within that window and no access to the discarded text. No comparator model was run for this table.
| Window | Basis | Depths reachable | Bound on exact recall |
|---|---|---|---|
| 1,000,000 tokens | computed; the last half of the context is inside the window | 4 of 9 | 44% |
| 256,000 tokens | computed | 2 of 9 | 22% |
| 200,000 tokens | computed; the first 1.8 million tokens are outside the window | 1 of 9 | 11% |
| 128,000 tokens | computed | 1 of 9 | 11% |
These bounds follow from the context length and the stated truncation rule. They do not account for retrieval or other memory outside the window.
Recall by depth
Depth is the position of the planted fact as a share of the context. The 40% to 60% band is where long-context systems have been reported to lose facts (Liu et al., 2023). Elara is drawn in gold; the computed bound for a 1,000,000-token window is drawn beside it for scale.
All eight trials at each depth passed: four fact families with two decoy layouts. The fact's position made no difference in this test.
How the test was built
- The context: 2,000,692 tokens of varied working notes, including the planted facts.
- The fact and its decoys: each true fact was planted with three decoys that differ from it by one word (main against backup, Bluefin against Redfin, vault against gate), each carrying a different code, so that a word match alone cannot separate them.
- The question: a reworded request containing the terms needed to identify the fact, without the answer. For example, the question asked which authorisation code opens the Bluefin main vault; the planted answer was QUARTZ-LYNX-7740.
- The pass condition: the whole planted sentence and its code return byte for byte, and a close paraphrase counts as a miss.
- The depths: 1%, 10%, 25%, 40%, 50%, 60%, 75%, 90% and 99% of the context.
- The sample: four fact families, two decoy layouts and nine depths, which gives 72 trials.
- The setting: the keyword tier, with relevance ranking switched off.
Design controls
The checks below define what this result establishes. The scoring logic has not been formally verified by an independent reviewer; the saved results and scripts are available for review.
- Decoys
- Every planted fact has three decoys that differ from it by one word and carry their own codes, and the question never contains the answer, so a match on words alone cannot return the right code. In all 72 trials the true code outranked the three decoys.
- Sample
- Four fact families, two independent decoy layouts and nine depths give 72 separate trials, and the result holds in each of them.
- Depth coverage
- The nine depths include 40%, 50% and 60%, the band where long-context systems have been reported to lose facts, and recall was exact there as well.
- Exactness
- A trial passes only when the whole sentence and its code come back byte for byte, so the presence of the right paragraph on its own scores nothing.
- Disambiguation
- Each trial also records whether the true fact outranks all three decoys, which it did in every trial, so the value returned is the planted one.
- Comparators
- The computed bounds assume perfect recall within a fixed window containing the most recent tokens. A system with retrieval or external memory falls outside that comparison.
- Scope of the claim
- The memory under test is a relevance-ranked memory mapped to the model rather than one 2,000,000-token attention matrix, and the claim is scoped to that.
Following a thread to 8,000,000 tokens
A separate experiment tested three tasks: finding a fact among irrelevant notes, following a reference to a second fact, and gathering related items across notes. It used contexts of 2,000,000, 4,000,000 and 8,000,000 tokens, with a reading budget of 4,000 tokens per question and a held-out seed for evaluation.
| Kind of question | Single-shot memory (most recent first) | Relevance-ranked window |
|---|---|---|
| Pick the relevant fact over present, irrelevant noise | answered at 2M, 4M and 8M | answered at 2M, 4M and 8M |
| Follow a reference to a second fact | unanswered at every size | answered at 2M, 4M and 8M |
| Gather all related items within the budget | unanswered at every size | answered at 2M, 4M and 8M |
At each size, the single-shot memory answered one of the three question types and the relevance-ranked window answered all three. Larger contexts added irrelevant material while keeping the reading budget fixed.
How the window was built
Successive versions were evaluated against a fixed test and held-out seed. Changes had to improve the evaluation score while preserving single-fact recall. The table records which versions were retained; this development history is not an independent test set.
| Generation | Change | Kinds answered | Held-out score | Kept |
|---|---|---|---|---|
| 0 | single-shot memory, most recent first | one of three | 100% | baseline |
| 1 | one global ranking, serialised relevance first | one of three | 33% | discarded |
| 2 | relevance edges: follow the reference on the same index | two of three | 67% | kept |
| 3 | relevance clusters: gather the cluster instead of a capped top list | three of three | 100% | kept |
The test harness and saved results are available with the reproduction command below.
Scope
This page shows exact recall of planted facts across one 2,000,692-token context, on the keyword tier, with 72 trials; and three kinds of question answered at 2,000,000, 4,000,000 and 8,000,000 tokens on one held-out seed per size. Both rest on a relevance-ranked memory mapped to the model. The comparator rows are computed bounds, and the table of generations records the development of the window rather than an external audit.
Run it yourself
Both scripts rebuild their context from scratch and write a fresh receipt; the recall run took 455 seconds and the window run took 50 seconds.
python tools/zero_drift_proof.py # 72 trials at 2,000,692 tokens -> artifacts/zero-drift/receipt.json python tools/relevance_window_proof.py --full # three kinds of question at 2M, 4M, 8M -> artifacts/attention-scale/receipt.json
The receipts this page reads from are kept in this repository as docs/proof/zero_drift_receipt.json and docs/proof/relevance_window_receipt.json. Reference for the middle-of-context effect: Liu et al., Lost in the Middle: How Language Models Use Long Contexts, 2023.