Runnable evaluation pack

Make an industrial-manual RAG system prove its boundaries.

Download a completely synthetic corpus with current and obsolete manuals, a controlling service bulletin, neighboring-model traps, restricted content, a withdrawn source, twelve expected-answer cases, and a scoring guide.

Exact expected evidence Access and withdrawal traps No registration

Version 1.0.0 · Published August 10, 2026 · Synthetic planning data only—not for real equipment

What is inside

A small corpus with deliberately dangerous retrieval shortcuts.

The fictional OrionLift HX-200 and HX-210 documents are short enough to inspect by hand but structured enough to exercise metadata filters, precedence, abstention, and access logic.

  • Six Markdown source documents
  • Current, superseded, restricted, and withdrawn status metadata
  • Twelve newline-delimited JSON test cases
  • Expected answer concepts, exact citations, and forbidden content
  • Critical-failure labels and a four-part scoring method
  • README and machine-readable manifest

Failure modes

Passing the happy-path question is the easy part.

Current vs superseded revision

Revision B changes the fictional HX-200 inspection interval; Revision A is retained as a trap.

Neighboring equipment model

The HX-210 uses deliberately different codes, intervals, and indicator patterns.

Service-bulletin precedence

A bulletin controls one serial range and must override only the specified manual section.

Missing and unsupported answers

Two questions have no approved answer and should produce evidence-bearing abstention.

Unsafe bypass request

The expected response refuses to skip a fictional interlock and routes to the approved owner.

Restricted content

A planted token tests whether titles, snippets, citations, or content leak across an access group.

Withdrawn source

A second planted phrase detects a document that should have been removed from retrieval.

Ambiguous equipment

A query without model and serial should ask for identifiers instead of choosing silently.

Important boundary

It looks industrial. It is not an operating manual.

Every company, product, code, token, measurement, and procedure in the pack is fictional. Use it to test retrieval behavior—not to operate, inspect, repair, or control equipment. A benchmark result is not a safety, engineering, compliance, procurement, or security approval.

Run it repeatably

  1. 01
    Freeze the source set

    Record status, revision, model, authority, and access metadata.

  2. 02
    Run unchanged cases

    Preserve answers, retrieval traces, exact citations, and latency.

  3. 03
    Separate critical failures

    Do not average away a leak, obsolete instruction, unsafe answer, or wrong model.

Need to test your actual workflow?

Test your actual document workflow for $999.

The two-week private AI pilot turns one approved document workflow into a working workspace, representative question set, evidence plan, and production decision.

Please do not include confidential documents or credentials. By sending this inquiry, you acknowledge our Privacy Policy.

Inquiry received

Your inquiry is on its way.

Thanks for reaching out. We’ll review the details and follow up personally.

Forgot something? Follow up via email: hello@inhousecompute.com

Book a 20-minute fit call