MillerKnoll AI
AI Leverage

The retrieval layer is the product.

SAM’s hard problem was never the model — it was making 2,191 pages of documentation verifiable. Parsing it into structured commands, retrieving them two ways, and grading the result with evals is what turns a big PDF into an assistant you can trust.

extract_adminguide.py
2,191 → 1,491

Parsing at scale

A parsing script read the entire 2,191-page admin guide and restructured it into 1,491 discrete, verifiable commands — work no manual indexing effort could absorb on top of everything else.

Dataverse + search
two strategies

Exact + semantic retrieval

Exact lookup for “what does this command do,” semantic search for “which command does this” — so a rigid index doesn’t miss the fuzzier questions developers actually ask.

Control set + gates
80% → 100%

Evaluation engineering

A documented eval set (by Neha Choudhary) grades every build: ≥80% to promote to TEST, 100% the goal before PROD, and any fabricated command is an automatic fail — turning “is it accurate?” into a number.

v2 → v3
73/75 parity

A retrieval layer that survived a replatform

When SAM moved onto the new skills-based experience, the Dataverse retrieval was rebuilt natively as a Workflow — and the same evals caught the two remaining gaps in the parity run, proving the layer held up.