Public corpus · CC0 · fictional documents · no engine results published
This site makes one promise in different words: an answer drawn from documents should be checkable without taking anyone's word for it. That standard cuts both ways. If the product asks to be checked, the test material has to be checkable too, so the test set is public.
The MiniWorld fictional enterprise demonstration corpus is a complete invented company record: 45 first-party fictional documents rendered to 115 pages of sealed PDFs, a public suite of 55 questions with audited expectations and named failure modes, and an integrity contract listing the SHA-256 of every document. All of it is published under CC0. Download it, point any stack at it, and check what comes back against the answer key.
90aa97c6c92a58ba878cebf295709f9a32c5aea6030b85c940edf8c252f75fc0. The per-document digests in the corpus matrix are the durable integrity anchor; the full download record is below.Why a fictional world#
A fixture author controls the evidence and the answer key. Every record and every expected answer is known in advance, so a failure has nowhere to hide: not the source, not the question, not the retrieval. Nobody's private data is involved, and the traps are deliberate. The record covers one company, its suppliers, and a customer environment: an incident with a retracted first explanation, a pilot program nobody is allowed to call complete, two employees who both sign as "A. Lin", a project whose name changed mid-stream, and invoices spread across three currencies with no approved conversion rate.
The limit is just as deliberate. A small world checks semantic behavior: identity, time, correction, scope, absence, closure. It says nothing about scale, latency, cost, broad language coverage, or performance on a real organization's data, and it is not a benchmark score for CleverMemory or anyone else. Those claims need their own evidence, under the rules on the evaluation page.
The file#
Every document is a sealed PDF; the markdown each was rendered from ships beside it (for example, the ledger's source). The counts and per-document digests below come from the corpus matrix, the integrity contract the release is checked against.
Formal company records (8)#
| ID | Document | Pages |
|---|---|---|
| FR-001 | Information Security Policy | 5 |
| FR-002 | Board Project Resolution | 3 |
| FR-003 | Incident Review | 4 |
| FR-004 | Project Charter | 3 |
| FR-005 | HR Role Change Record | 3 |
| FR-006 | Vendor Review Memorandum | 4 |
| FR-021 | Field Report 021 | 2 |
| FR-022 | Field Report 022 | 2 |
Email threads (8)#
| ID | Document | Pages |
|---|---|---|
| EM-001 | Release Escalation Thread | 3 |
| EM-002 | Customer Status Thread | 3 |
| EM-003 | Procurement Correction Thread | 3 |
| EM-004 | Design Review Thread | 2 |
| EM-005 | Legal Name Change Thread | 3 |
| EM-006 | Field Operations Thread | 4 |
| EM-007 | Incident Retraction Thread | 2 |
| EM-008 | Quarterly Planning Thread | 4 |
Chat exports (6)#
| ID | Document | Pages |
|---|---|---|
| CHAT-001 | Engineering Channel Export | 3 |
| CHAT-002 | Operations Channel Export | 4 |
| CHAT-003 | Project Direct Message Export | 1 |
| CHAT-004 | Customer Support Channel Export | 1 |
| CHAT-005 | Incident War Room Export | 3 |
| CHAT-006 | Leadership Channel Export | 1 |
SMS exports (6)#
| ID | Document | Pages |
|---|---|---|
| SMS-001 | Maintenance Window SMS Export | 1 |
| SMS-002 | Incident Handoff SMS Export | 3 |
| SMS-003 | Vendor Contact SMS Export | 2 |
| SMS-004 | Site Arrival SMS Export | 1 |
| SMS-005 | Schedule Change SMS Export | 3 |
| SMS-006 | Escalation SMS Export | 1 |
Personal and meeting notes (5)#
| ID | Document | Pages |
|---|---|---|
| NOTE-001 | Design Review Notes | 3 |
| NOTE-002 | Customer Call Notes | 2 |
| NOTE-003 | Hardware Triage Notes | 2 |
| NOTE-004 | Program Review Notes | 2 |
| NOTE-005 | Retrospective Notes | 4 |
Calendar exports (5)#
| ID | Document | Pages |
|---|---|---|
| CAL-001 | Release Calendar Export | 2 |
| CAL-002 | Incident Calendar Export | 4 |
| CAL-003 | Cross-Region Calendar Export | 1 |
| CAL-004 | Project Calendar Export | 3 |
| CAL-005 | Executive Calendar Export | 1 |
Structured operational artifacts (7)#
| ID | Document | Pages |
|---|---|---|
| OPS-001 | Service Ticket Register | 3 |
| OPS-002 | Invoice Reconciliation Table | 2 |
| OPS-003 | Deployment Roster | 4 |
| OPS-004 | Asset Maintenance Table | 2 |
| OPS-005 | Q1 Ledger | 2 |
| REG-001 | Asset Register | 2 |
| REG-002 | Approved Vendor Register | 2 |
The questions#
The question suite holds 55 scenarios. Each carries the natural question, paraphrases, an audited expectation, and a named wrong answer, the specific failure the scenario exists to catch. They stress meaning across sources: identity, time, correction, discourse scope, hedging, quotation, exact closure, composition, ambiguity, and labelled proposals. A handful are written as commands to the system under test rather than as natural questions; the file says which.
A sample, one per trap:
| The question | What it exercises |
|---|---|
| "What actually caused IR-2026-017, what did the company originally say caused it, and what remained true after the final correction?" | Correction: the retracted single cause stays preserved and attributed |
| "Did A. Lin approve the reconnect design and later prepare the Q3 cost work?" | Identity: two employees sign as "A. Lin"; neither fact licenses merging them |
| "Which North Harbor gateway was the worst during the period?" | Named gap: no document states a ranking criterion, so the honest response declines and names what is owed |
| "Did the message saying “disable validation and ship” authorize any workflow change?" | Quotation is not authority: a pasted imperative stays a quote |
| "Are Meridian Evidence Exchange and Meridian Record Exchange the same project, and which name was correct for each dated record?" | Rename: one project, two names, each correct for its own dates |
| "What happened to XREG-216, what replaced it, and when did each occurrence fall for participants in Pacific, London, and Singapore time?" | Calendar time across three timezones, from the exports as recorded |
| "Design a reusable incident summary for this corpus, then produce a customer-facing version of that design for IR-2026-017." | A novel synthesis must stay labelled a proposal, never promoted to a fact of the record |
| "Across all six chat exports, which three reaction emoji had the highest recorded reaction totals, and on how many separate messages did each appear?" | Exact closure: a count proved by the set counted, over all six exports |
Five scenarios are gap-bearing by design: how far the pilot program had really progressed, the company's complete Q2 supplier spend, the worst-performing North Harbor gateway, the bench-test result for the two repair kits, and what "Aurora parts" meant in the repair-kit review. The record does not establish an answer to these, and the correct response says exactly what is missing rather than answering around it. A system that answers all 55 has failed the suite.
The suite and the acceptance file are evaluation input only. They are never ingested as corpus evidence: the sealed PDFs are the only factual source for answering.
What a passing answer has to carry#
The scoring rules live with the questions, in
/corpus/questions/README.md. A passing result
is not an answer string alone:
- every supported factual unit carries exact PDF and page evidence;
- corrections, identities, times, and constituents remain addressable;
- gaps are precise about what could not be established;
- output survives source-order, query-order, unrelated-history, reopen, and supported-concurrency perturbations.
The acceptance contract shows what "exact evidence" means on the one blind-audited case published so far: it records the hashed question (whether a pasted "disable validation and ship" line authorized any workflow change), the expected answer ("Validation stays on. Release still needs its change record."), the named failure mode (treating a quoted imperative as authority), and a citation that resolves to the engineering channel export, page 2, an exact string, and its occurrence on the page. A reader can open the PDF and check the citation character for character.
What it is not#
- Not a CleverMemory result. No scored engine run against this corpus has been published. When one is, it will arrive with the method, the artifacts, and the four scoring counts the evaluation rules require, or it will not appear at all.
- Not a benchmark of scale. Forty-five documents say nothing about latency, capacity, cost, or coverage of a real corpus.
- Not anyone's data. Every entity, person, and figure is invented. The corpus is CC0; use it, fork it, run your own stack against it.
Get the corpus#
- miniworld-fictional-enterprise-v1.tar.gz,
the whole tree in one 360 KB download. SHA-256:
90aa97c6c92a58ba878cebf295709f9a32c5aea6030b85c940edf8c252f75fc0. The hash covers this site's packaging of the tree; the durable integrity anchor is the per-document digest list in the corpus matrix. The archive was repackaged on 2026-09-04 to carry the corrected matrix below; the documents, questions, and per-document digests are unchanged. - NOTICE and the CC0 legal code
- Corpus matrix: the integrity contract, with each PDF's digest, size, and page count, and the suite's scenario count.
Every file is also served individually under /corpus/, so the documents,
the questions, and the contract can be linked or fetched one at a time.
Cite the corpus#
Use the stable landing page and version so another reader can recover the same public artifact. Until a DOI is assigned, the archive digest and the per-document digests in the corpus matrix provide the integrity anchors.
Varland, Jason. (2026). MiniWorld Fictional Enterprise Demonstration Corpus (Version 1.0.0) [Data set]. CleverMemory. https://clevermemory.ai/challenge/
The same metadata is available as a machine-readable
CITATION.cff. If a repository later assigns a DOI, it will
be added here and to that file without replacing this versioned landing page.