2,191 pages, made answerable — and grounded.
SmartAssembly developers needed an assistant grounded in the official admin guide — but that guide runs 2,191 pages, where keyword search returns a page number, not an answer. SAM parses it into a structured, verifiable knowledge layer and puts an orchestrative multi-agent front end on it: ask in natural language, get a cited, grounded answer — and never a fabricated command.
What got measurably better
A structured knowledge layer, a multi-agent front end, and an evaluation gate that turns “grounded” into a number.
The eval gate is the whole point
SAM answers nothing from the model’s own memory. A documented control set grades every build — and any fabricated command is an automatic failure.
(97%)
Any fabricated command name is an automatic eval failure. SAM flags the invalid name and returns the real alternatives — the definition of cite-or-abstain.
Three rules hold it together
Grounded, never from memory
The orchestrator answers no technical question itself. Every claim comes from a specialist reading authoritative SmartAssembly documentation — the model coordinates, the sources decide.
Cite-or-abstain
A fabricated command name is an automatic evaluation failure. Ask for a command that doesn’t exist and SAM says so, then offers the documented alternatives — it never invents syntax.
Separation of concerns
Reference looks up, Guide teaches, Debug diagnoses. Each specialist owns one job against one source of truth, so the system stays deterministic and debuggable.
From a parsing script to a replatform
Parse the 2,191-page guide
A parsing script (extract_adminguide.py) restructures the admin guide into 1,491 discrete, structured commands.
Dataverse retrieval layer
The commands load into Dataverse for exact lookup, paired with semantic search for domain-discovery questions.
Orchestrative multi-agent (v2)
An Orchestrator routes intent to three specialists — Reference, Guide, Debug — each grounded in the authoritative sources.
v3 — skills-based, Sonnet 5
Replatformed onto the new experience: specialists become inline skills, retrieval rebuilt as a native Workflow, eval parity confirmed at 73/75.
Live in PROD as SAM
UAT signed off by Doug Brandt; promoted DEV → TEST → PROD as a clean managed solution (SmartAssemblyMentorSAM) and renamed Smart Assembly Mentor (SAM), shared with one small-team security group across all three environments.
Where it stands
SAM is live in Copilot Agents PROD as of September 18, 2026 — the v2 multi-agent build replatformed onto the new skills-based experience on Claude Sonnet 5, with the Dataverse retrieval flow rebuilt as a native Workflow. It cleared the parity eval at 73/75, passed UAT with Doug Brandt, and was promoted DEV → TEST → PROD as a clean managed solution, renamed from SAM v3 RP to Smart Assembly Mentor (SAM). The evaluation discipline — a documented control set where any fabricated command is an automatic failure, 80% to reach TEST and 100% the goal before PROD — is what makes “grounded, not guessed” a measured claim rather than a hope.
