CleverMemory

Public corpus · CC0 · fictional documents · no engine results published

This site makes one promise in different words: an answer drawn from documents should be checkable without taking anyone's word for it. That standard cuts both ways. If the product asks to be checked, the test material has to be checkable too, so the test set is public.

The MiniWorld fictional enterprise demonstration corpus is a complete invented company record: 45 first-party fictional documents rendered to 115 pages of sealed PDFs, a public suite of 55 questions with audited expectations and named failure modes, and an integrity contract listing the SHA-256 of every document. All of it is published under CC0. Download it, point any stack at it, and check what comes back against the answer key.

Download the corpus · 360 KB · CC0SHA-256 90aa97c6c92a58ba878cebf295709f9a32c5aea6030b85c940edf8c252f75fc0. The per-document digests in the corpus matrix are the durable integrity anchor; the full download record is below.

Why a fictional world#

A fixture author controls the evidence and the answer key. Every record and every expected answer is known in advance, so a failure has nowhere to hide: not the source, not the question, not the retrieval. Nobody's private data is involved, and the traps are deliberate. The record covers one company, its suppliers, and a customer environment: an incident with a retracted first explanation, a pilot program nobody is allowed to call complete, two employees who both sign as "A. Lin", a project whose name changed mid-stream, and invoices spread across three currencies with no approved conversion rate.

The limit is just as deliberate. A small world checks semantic behavior: identity, time, correction, scope, absence, closure. It says nothing about scale, latency, cost, broad language coverage, or performance on a real organization's data, and it is not a benchmark score for CleverMemory or anyone else. Those claims need their own evidence, under the rules on the evaluation page.

The file#

Every document is a sealed PDF; the markdown each was rendered from ships beside it (for example, the ledger's source). The counts and per-document digests below come from the corpus matrix, the integrity contract the release is checked against.

Formal company records (8)#

IDDocumentPages
FR-001Information Security Policy5
FR-002Board Project Resolution3
FR-003Incident Review4
FR-004Project Charter3
FR-005HR Role Change Record3
FR-006Vendor Review Memorandum4
FR-021Field Report 0212
FR-022Field Report 0222

Email threads (8)#

IDDocumentPages
EM-001Release Escalation Thread3
EM-002Customer Status Thread3
EM-003Procurement Correction Thread3
EM-004Design Review Thread2
EM-005Legal Name Change Thread3
EM-006Field Operations Thread4
EM-007Incident Retraction Thread2
EM-008Quarterly Planning Thread4

Chat exports (6)#

IDDocumentPages
CHAT-001Engineering Channel Export3
CHAT-002Operations Channel Export4
CHAT-003Project Direct Message Export1
CHAT-004Customer Support Channel Export1
CHAT-005Incident War Room Export3
CHAT-006Leadership Channel Export1

SMS exports (6)#

IDDocumentPages
SMS-001Maintenance Window SMS Export1
SMS-002Incident Handoff SMS Export3
SMS-003Vendor Contact SMS Export2
SMS-004Site Arrival SMS Export1
SMS-005Schedule Change SMS Export3
SMS-006Escalation SMS Export1

Personal and meeting notes (5)#

IDDocumentPages
NOTE-001Design Review Notes3
NOTE-002Customer Call Notes2
NOTE-003Hardware Triage Notes2
NOTE-004Program Review Notes2
NOTE-005Retrospective Notes4

Calendar exports (5)#

IDDocumentPages
CAL-001Release Calendar Export2
CAL-002Incident Calendar Export4
CAL-003Cross-Region Calendar Export1
CAL-004Project Calendar Export3
CAL-005Executive Calendar Export1

Structured operational artifacts (7)#

IDDocumentPages
OPS-001Service Ticket Register3
OPS-002Invoice Reconciliation Table2
OPS-003Deployment Roster4
OPS-004Asset Maintenance Table2
OPS-005Q1 Ledger2
REG-001Asset Register2
REG-002Approved Vendor Register2

The questions#

The question suite holds 55 scenarios. Each carries the natural question, paraphrases, an audited expectation, and a named wrong answer, the specific failure the scenario exists to catch. They stress meaning across sources: identity, time, correction, discourse scope, hedging, quotation, exact closure, composition, ambiguity, and labelled proposals. A handful are written as commands to the system under test rather than as natural questions; the file says which.

A sample, one per trap:

The questionWhat it exercises
"What actually caused IR-2026-017, what did the company originally say caused it, and what remained true after the final correction?"Correction: the retracted single cause stays preserved and attributed
"Did A. Lin approve the reconnect design and later prepare the Q3 cost work?"Identity: two employees sign as "A. Lin"; neither fact licenses merging them
"Which North Harbor gateway was the worst during the period?"Named gap: no document states a ranking criterion, so the honest response declines and names what is owed
"Did the message saying “disable validation and ship” authorize any workflow change?"Quotation is not authority: a pasted imperative stays a quote
"Are Meridian Evidence Exchange and Meridian Record Exchange the same project, and which name was correct for each dated record?"Rename: one project, two names, each correct for its own dates
"What happened to XREG-216, what replaced it, and when did each occurrence fall for participants in Pacific, London, and Singapore time?"Calendar time across three timezones, from the exports as recorded
"Design a reusable incident summary for this corpus, then produce a customer-facing version of that design for IR-2026-017."A novel synthesis must stay labelled a proposal, never promoted to a fact of the record
"Across all six chat exports, which three reaction emoji had the highest recorded reaction totals, and on how many separate messages did each appear?"Exact closure: a count proved by the set counted, over all six exports

Five scenarios are gap-bearing by design: how far the pilot program had really progressed, the company's complete Q2 supplier spend, the worst-performing North Harbor gateway, the bench-test result for the two repair kits, and what "Aurora parts" meant in the repair-kit review. The record does not establish an answer to these, and the correct response says exactly what is missing rather than answering around it. A system that answers all 55 has failed the suite.

The suite and the acceptance file are evaluation input only. They are never ingested as corpus evidence: the sealed PDFs are the only factual source for answering.

What a passing answer has to carry#

The scoring rules live with the questions, in /corpus/questions/README.md. A passing result is not an answer string alone:

  • every supported factual unit carries exact PDF and page evidence;
  • corrections, identities, times, and constituents remain addressable;
  • gaps are precise about what could not be established;
  • output survives source-order, query-order, unrelated-history, reopen, and supported-concurrency perturbations.

The acceptance contract shows what "exact evidence" means on the one blind-audited case published so far: it records the hashed question (whether a pasted "disable validation and ship" line authorized any workflow change), the expected answer ("Validation stays on. Release still needs its change record."), the named failure mode (treating a quoted imperative as authority), and a citation that resolves to the engineering channel export, page 2, an exact string, and its occurrence on the page. A reader can open the PDF and check the citation character for character.

What it is not#

  • Not a CleverMemory result. No scored engine run against this corpus has been published. When one is, it will arrive with the method, the artifacts, and the four scoring counts the evaluation rules require, or it will not appear at all.
  • Not a benchmark of scale. Forty-five documents say nothing about latency, capacity, cost, or coverage of a real corpus.
  • Not anyone's data. Every entity, person, and figure is invented. The corpus is CC0; use it, fork it, run your own stack against it.

Get the corpus#

  • miniworld-fictional-enterprise-v1.tar.gz, the whole tree in one 360 KB download. SHA-256: 90aa97c6c92a58ba878cebf295709f9a32c5aea6030b85c940edf8c252f75fc0. The hash covers this site's packaging of the tree; the durable integrity anchor is the per-document digest list in the corpus matrix. The archive was repackaged on 2026-09-04 to carry the corrected matrix below; the documents, questions, and per-document digests are unchanged.
  • NOTICE and the CC0 legal code
  • Corpus matrix: the integrity contract, with each PDF's digest, size, and page count, and the suite's scenario count.

Every file is also served individually under /corpus/, so the documents, the questions, and the contract can be linked or fetched one at a time.

Cite the corpus#

Use the stable landing page and version so another reader can recover the same public artifact. Until a DOI is assigned, the archive digest and the per-document digests in the corpus matrix provide the integrity anchors.

Varland, Jason. (2026). MiniWorld Fictional Enterprise Demonstration Corpus (Version 1.0.0) [Data set]. CleverMemory. https://clevermemory.ai/challenge/

The same metadata is available as a machine-readable CITATION.cff. If a repository later assigns a DOI, it will be added here and to that file without replacing this versioned landing page.

Run it, then compare notes.Whether your stack answers with receipts or cannot, either result is worth a conversation: get in touch. The use cases walk six questions over the same documents, five drawn from the suite in shortened form and one written for the clarification outcome, with the traps explained, if you want a guided pass first.

↑ Back to top