The same problem shows up in very different businesses: the answer is spread across documents, the documents have been revised, and the person making the decision needs to see what the answer rests on before acting on it. I have watched reviewers do this reconciliation by hand, paging between a correction and the wording it replaced, and it is careful, slow work that deserves better than a highlighter.
Every pattern below is shown on one shared cast: the case file, a published fictional company record of 45 documents across ledgers, email threads, chat exports, calendars, notes, and formal reviews. Open the file alongside these patterns and check them the way a reviewer would. These are capability patterns, not deployment claims, compliance certifications, or promises of production readiness.
Seven patterns, one cast#
Each pattern is shown on the shared case file: the excerpts a good response would use, the lookalikes it must set aside, and the outcome it should return. The documents behind every pattern are published further down if you want to check any pattern against the source, the way a reviewer would.
01 · Verify an answer
Keep the answer attached to the records behind it.
"How much did Vortex Grid spend with Aurora Circuit Works Ltd. in Q2 2026?" The answer, GBP 14,542.00, is only useful if the application can point at what establishes it: the five matched ledger rows carrying that legal name, the credit that corrects one of them, and the header rule that fixes the currency. A plausible answer is not enough when someone has to defend it.
For technical readers
The response carries machine-checkable proof and source references for each material claim, including what was set aside: the struck summary rows and the namesake supplier do not become the answer. Question 1 below works this case in full, and the worked envelope below shows the response it should get.
INV-4404, INV-4409, INV-4413, INV-4417, INV-4422 name Aurora Circuit Works Ltd. (SUP-778), each in GBP.
−48.00 GBP corrects the INV-4422 unit-price variance; the original row is retained.
GBP rows stay GBP: no implicit conversion, and no approved FX rate is attached.
Repeat 2,436.00 and 3,392.00 under a struck short name. Not new charges.
SUP-441, in USD. The shared word "Aurora" is where the names overlap.
Awaiting receiving confirmation; not included as matched.
With proof and sources attached.
02 · Ask what applied then
Keep today's status and the one that held at the checkpoint.
"At 15:30 UTC on 17 June 2026, what was the status of the MEE-26 steering meeting and its replacement?" The calendar export on record was taken on 21 July and does not reach backwards. At that instant the original session was already cancelled, and its replacement was confirmed for the next day. Both rows remain in the export, each carrying the status that held.
For technical readers
When a claim applied and when it entered the record are separate axes, and questions can be asked along either. A time-scoped answer returns the applicable version and the date boundary with the result. Questions without a date get the current state: current by default, history on request. Question 5 below tries this one.
Title struck through, status: Cancelled. The row names its replacement.
The replacement, status: Confirmed, still in the future at the checkpoint.
The export date and every later row describe July, not 17 June.
The state that held at that instant, not the state that holds now.
03 · Read a figure in context
Where a number sits is part of what it means.
The Q2 ledger reads 3,392.00 twice: once on the INV-4422 invoice row, once on SUMMARY-AURORA-2, a historical summary whose supplier name is struck through. Matching on the figure alone counts the same charge twice. The row kind (invoice, adjustment, summary, clarification) is what distinguishes them, so it has to survive into the answer.
For technical readers
Document geometry is treated as part of the record rather than as formatting to be discarded before indexing. The sample below shows what a plain-text extractor throws away on exactly this table.
| Row | Amount |
|---|---|
| IROW-018 · INV-4422 | 3,392.00 GBP |
| IROW-019 · ADJ-4422-A | −48.00 GBP |
3,392.00 GBP under a name struck through; the same charge repeated, not a second one.
0.00. Records which supplier the row means; adds nothing.
Row kind returned with the value.
04 · Find missing support
Name the gap instead of answering around it.
"What result did bench test HTL-2026-044A produce for the two circuit-assembly kits?" Everything around the answer is legible: both packaging seals were intact, nothing visible was damaged, and neither assembly was powered. But the bench continuity test has not been run, and the notes say so twice, including an instruction to no one in particular: do not append the result here from memory.
For technical readers
A named gap is a first-class outcome, not a low-confidence answer: the response identifies the fact that could not be established from the material in scope. Question 4 below is this case.
Visual only: seals intact, no visible moisture, deformation, or connector damage; neither assembly powered.
Bench continuity tests under test record HTL-2026-044A: still open, unchecked.
No bench-test result exists to report, in either direction.
05 · Connect documents
Two field reports, one incident, two exact counts.
The remote sweep says exactly 42 devices were affected by service disruption IR-2026-024. The physical inspection says exactly 17. Both reports say "exactly", both are dated the same day, and the second offers its own explanation for the difference. Whether that explanation settles the question is not for the system to decide silently, so the response shows both counts, their methods, and their sources.
For technical readers
Records that disagree in a way that matters are surfaced as a conflict, with the disagreeing records attached. Scope decides what is in play: leave one report out and the question answers cleanly with a single number, which is correct behavior, not wobble. Question 3 below.
"Exactly 42 devices" failed remote health queries during the incident window.
"Exactly 17 devices" with confirmed local buffer faults.
42 by reachability, 17 by hardware faults; both counts stay attributed.
06 · Count with a visible set
A count is proved by its members.
"Among the incident-linked tickets, how many met the standing customer-notification criterion?" The register rows alone cannot answer it: the criterion, reportable when the noted delay reaches sixty minutes, lives in the incident review, not beside the data. The answer is four tickets totalling 460 minutes, and the members make it auditable.
For technical readers
An aggregate is proved by the set counted, the record behind each member, and the boundary the count claims to be complete over. Question 6 below tries it.
TKT-2612 (227), TKT-2613 (74), TKT-2614 (96), TKT-2616 (63).
Below the sixty-minute criterion.
The criterion itself; the register stores the minutes without restating it.
With the set it counted, and the document the boundary came from.
07 · Produce a document
Reissue the summary, with an account of what changed.
The quarterly close summary used "Aurora" for two different supplier matters, and the correction on record attaches a dated clarification instead of overwriting the original. Reissuing the summary itself means two mentions resolve to legal names with vendor IDs, every figure stays put, and the reason for the change is carried on the face of the document. Design target, not tested behavior.
For technical readers
A document transform returns the reissued document plus an account of what was preserved, changed, recalculated, or omitted, so the difference is reviewable rather than diffed by hand. The contract is in Reports and transforms on the technical overview.
"Aurora" for both the sensor-housing shipment and the circuit-assembly invoice.
PO-60731 → Aurora Components LLC, SUP-441; PO-61108 → Aurora Circuit Works Ltd., SUP-778 (per clarification VRM-2026-09). Both original passages preserved.
The case file#
Fictional documents · CC0 · illustrative, not a benchmark
Every example on this site uses the same cast: a fictional company, Vortex Grid Systems Inc., its suppliers, and one customer environment, across a record of email threads, chat and SMS exports, calendars, notes, operational registers, and formal reviews. The traps are deliberate. A careless read, and a similarity search, can land on a namesake supplier, a struck-through figure, or a cause the record no longer states, and each of those is genuinely written down somewhere in the file.
The documents are published in full, as sealed PDFs with a per-document integrity contract, under CC0: the 45-document challenge is the whole file, with the audited 55-question suite anyone can run. The documents the patterns above draw on:
Nothing here is a benchmark, a product screenshot, or a claim about released software. The point comes before any of that: a reader should be able to feel why the question is hard before being asked to believe any solution. So read the file, pick a question, and answer it from the documents alone, noting which passages you used and which you set aside. That discipline, what was used, what was set aside, and why, is exactly what the technical overview asks of a machine, scored under the rules on the evaluation page.
The documents are also dressed the way source material actually arrives: dated in two timezones, corrected in place and by attachment, signed with initials two people share, and split across formats. That is not decoration. Real documents are messy, and the mess is where wrong answers come from.
See what a plain-text extractor keeps
VORTEX GRID SYSTEMS INC. Invoice Reconciliation Table Record ID OPS-002 Ledger Finance Controls Q2 supplier review Row ID Invoice adjustment Legal supplier Vendor ID PO Qty Unit price Currency Extended amount Reconciliation note IROW-018 INV-4422 Aurora Circuit Works Ltd. SUP-778 PO-61108 16 212.00 GBP 3,392.00 Unit-price variance under review separate legal supplier IROW-019 ADJ-4422-A Aurora Circuit Works Ltd. SUP-778 PO-61108 16 -3.00 GBP -48.00 Supplier credit corrects unit-price variance original invoice retained IROW-021 SUMMARY-AURORA-2 Aurora Ambiguous original summary UNRESOLVED-ORIGINAL PO-61108 16 212.00 GBP 3,392.00 Historical summary used the short name linked to SUP-778 by clarification
Every number survives extraction, including the one that would double-count. What does not survive is everything that made them checkable: the row IDs that name each row's kind, the strikethrough that retires the short name, the footnote that says clarification rows add nothing, and the header rule that fixes the currency. A search over this text finds 3,392.00 twice and cannot tell you one is a summary of the other. Where a number sits is part of what it means.
All entities, people, addresses, dates, and figures are fictional. The full record is published under CC0 on the challenge page.
Questions to try#
Work from the documents above and nothing else. Each question names the outcome a good response should have; the notes say where the wrong answers come from. Question 1 is worked in full at the end, and the full audited suite of 55 questions ships with the corpus. Five of the six are suite scenarios, shortened where the suite asks for more than one thing at once. The clarify question is this site's own: the published suite scores answers and gaps, and has no scenario whose expected outcome is a clarification request.
1. The answerable one#
"How much did Vortex Grid spend with Aurora Circuit Works Ltd. in Q2 2026?"
A good response answers: GBP 14,542.00, with proof and sources attached.
What makes this hard
The total is built from pieces that each invite a different error: five invoice rows under a legal name that shares a word with a different supplier; a credit that corrects one invoice beside the retained original; two summary rows that repeat 2,436.00 and 3,392.00 under a struck short name; a clarification row that adds nothing; a pending invoice that is not yet matched; and three currencies with no approved conversion rate. Reaching 14,542.00 requires the row IDs; reaching GBP requires not converting. One assumption is worth naming: the VRM-2026-09 clarification maps PO-61108 to SUP-778, and question 2 asks what happens without it.
2. The clarify one#
"How much did Vortex Grid spend with Aurora in Q2 2026?"
A good response asks: Aurora Components LLC, or Aurora Circuit Works Ltd.?
What makes this hard
One word names two legal suppliers with different vendor IDs and different currencies, and the management summary used it for both. Picking one without asking is a coin flip dressed as an answer. A similarity search cannot rescue the choice either: the string "Aurora" matches both, and both genuinely have Q2 spend. The correction on record exists precisely because someone caught the ambiguity before it reached a supplier scorecard.
3. The conflict one#
"How many devices were affected by service disruption IR-2026-024?"
A good response surfaces the conflict: Field Report 021 says exactly 42, from a remote sweep; Field Report 022 says exactly 17, from physical inspection. Shown, not chosen.
What makes this hard
Both reports say "exactly", both are dated the same day, and the second offers its own explanation for the difference, which is itself a claim in the record rather than a resolution the system may silently adopt. A system that picks one is guessing with extra steps. Notice what happens when one report is left out of scope: the question answers cleanly with a single number, which is correct behavior, not wobble. Scope is part of the question.
4. The gap one#
"What result did bench test HTL-2026-044A produce for the two circuit-assembly kits?"
A good response names the gap: no result exists; the bench test has not been run.
What makes this hard
Everything around the answer is legible (seals intact, no visible damage, a checklist with one open item), so the missing result hides behind a complete-looking note. The visual inspection establishes what it establishes; it does not establish that either kit passed or failed, in either direction. The honest move is to say what is missing rather than assume it, and the notes themselves instruct exactly that: do not append the result from memory.
5. The time-scoped one#
"At 15:30 UTC on 17 June 2026, what was the status of the MEE-26 steering meeting and its replacement?"
A good response answers: already cancelled, with the replacement confirmed for the next day, the state that held at that instant.
What makes this hard
The export on record was taken on 21 July, weeks after the checkpoint, and reading the calendar's later state back into June is the error. The two sessions share a project and nearly a name, and the cancelled row keeps its struck-through title and its link forward to the replacement. The same file yields different correct answers depending on an instant that appears only in the question.
6. The count#
"Among the incident-linked tickets in the service register, how many met the standing customer-notification criterion, and what delay did they total?"
A good response answers: four tickets, 460 minutes, and shows the set it counted: TKT-2612, TKT-2613, TKT-2614, TKT-2616.
What makes this hard
The criterion lives in the incident review, not beside the register: a ticket whose noted delay reaches sixty minutes is reportable. The register stores the minutes without restating the rule, so the count has to join two documents. TKT-2615, at 52 minutes, falls below the criterion and is excluded, and the response has to say so. A bare number is not reviewable either way.
What a response has to contain#
The short envelopes on the technical overview show the shape of a response. Below is what the inside of one looks like when the proof is written out: the response a reader should be able to demand for question 1, with the outcome, the qualifications, the sources, and the derivation in order.
outcome: answer
question: "How much did Vortex Grid spend with Aurora Circuit Works Ltd. in Q2 2026?"
answer: "GBP 14,542.00"
qualifications:
currency: GBP, unconverted; the ledger attaches no approved FX rate
supplier: Aurora Circuit Works Ltd. (SUP-778), not Aurora Components LLC (SUP-441)
period: 2026-04-01 through 2026-06-30, per the ledger header
basis: matched rows only; the pending invoice is excluded
sources:
- OPS-002, five matched invoice rows naming SUP-778
- OPS-002, credit row ADJ-4422-A against INV-4422
- OPS-002, ledger header currency rule
- EM-003, clarification VRM-2026-09 mapping PO-61108 to SUP-778
proof:
- step: fix the entity
used: clarification VRM-2026-09, PO-61108 to Aurora Circuit Works Ltd. / SUP-778
set_aside: the Aurora Components LLC rows (SUP-441, USD; a different legal supplier)
- step: select the charge rows
used: IROW-003, IROW-007, IROW-011, IROW-014, IROW-018
set_aside: SUMMARY-AURORA-2 (a historical summary, not a new charge); CLAR-4422 (identity only, 0.00)
- step: apply the correction
used: IROW-019, credit ADJ-4422-A of -48.00 against INV-4422, original row retained
set_aside: overwriting the original row (the record keeps both)
- step: hold the unit
used: the ledger header rule: GBP rows stay GBP, no implicit conversion
set_aside: a converted grand total (not available from this record)
result: GBP 14,542.00
Illustrative, not a benchmark or a screenshot. The field names and layout are not a frozen schema; the contract is the behavior, not the spelling.
Nothing in that object asks to be believed on presentation. A checker reads the response, the cited records, and the recorded steps, then replays the derivation and returns a verdict: the rows really carry the legal name, the credit really corrects the invoice, the header really fixes the currency, and no set-aside record was needed to reach the result. What can check the proof, and what its verdict does and does not mean, is covered in the technical overview.
Notice also what is not in the envelope: how the passages were found, read, or ranked. One part of CleverMemory stays ours: the code and data that tie everything together. That piece is not public; everything you need to check what it tells you is, and on this page it is sitting just above. The response has to stand on what it contains.
What a good fit looks like#
CleverMemory is a general layer, not a vertical product, so the honest guide is the shape of the problem rather than the industry:
- The record gets revised. Amendments, versions, reissues, and corrections mean the current answer and the historical answer are both needed, and both have to stay checkable.
- Position carries meaning. Tables, exhibits, definitions clauses, and signature blocks change what a figure or sentence commits to.
- Someone signs for the answer. If a person or regulator can ask "what is this based on," the answer has to arrive with its basis attached.
- Silence has to be distinguishable. "The record does not say" must be a different outcome from "the record says no."
A platform boundary, not a vertical product#
CleverMemory is not being built as a contract tool, a clinical system, or a procurement suite. It is a document intelligence layer those applications could use. Domain rules, access controls, review processes, and regulatory approval remain the responsibility of the product and organization around it.