The retrieval layer is the product.
SAM’s hard problem was never the model — it was making 2,191 pages of documentation verifiable. Parsing it into structured commands, retrieving them two ways, and grading the result with evals is what turns a big PDF into an assistant you can trust.
Parsing at scale
A parsing script read the entire 2,191-page admin guide and restructured it into 1,491 discrete, verifiable commands — work no manual indexing effort could absorb on top of everything else.
Exact + semantic retrieval
Exact lookup for “what does this command do,” semantic search for “which command does this” — so a rigid index doesn’t miss the fuzzier questions developers actually ask.
Evaluation engineering
A documented eval set (by Neha Choudhary) grades every build: ≥80% to promote to TEST, 100% the goal before PROD, and any fabricated command is an automatic fail — turning “is it accurate?” into a number.
A retrieval layer that survived a replatform
When SAM moved onto the new skills-based experience, the Dataverse retrieval was rebuilt natively as a Workflow — and the same evals caught the two remaining gaps in the parity run, proving the layer held up.
