MillerKnoll AI
Executive Summary

2,191 pages, made answerable — and grounded.

SmartAssembly developers needed an assistant grounded in the official admin guide — but that guide runs 2,191 pages, where keyword search returns a page number, not an answer. SAM parses it into a structured, verifiable knowledge layer and puts an orchestrative multi-agent front end on it: ask in natural language, get a cited, grounded answer — and never a fabricated command.

By the numbers

What got measurably better

A structured knowledge layer, a multi-agent front end, and an evaluation gate that turns “grounded” into a number.

2,191 → 1,491
Admin-guide pages parsed into structured commands
Exact + semantic
Dual retrieval strategy over Dataverse
73 / 75
Replatform parity eval
Orchestrator + 3
Reference · Guide · Debug specialists
80% → 100%
Eval gate: TEST → PROD
Sonnet 5
Skills-based replatform (RP)
Grounded, measured

The eval gate is the whole point

SAM answers nothing from the model’s own memory. A documented control set grades every build — and any fabricated command is an automatic failure.

73/75parity eval
(97%)
│ 80% → TEST100% → PROD goal
Cite-or-abstain · from the eval set
▸ What does USER_SELECT_MODEL do?
“USER_SELECT_MODEL” is not a documented SmartAssembly command.
You likely want one of these:
LOAD_MDL — Assembly Description Block · opens a file dialog to select a model from disk.
USER_SELECT — selects a geometry reference.

Any fabricated command name is an automatic eval failure. SAM flags the invalid name and returns the real alternatives — the definition of cite-or-abstain.

Why it works

Three rules hold it together

01

Grounded, never from memory

The orchestrator answers no technical question itself. Every claim comes from a specialist reading authoritative SmartAssembly documentation — the model coordinates, the sources decide.

02

Cite-or-abstain

A fabricated command name is an automatic evaluation failure. Ask for a command that doesn’t exist and SAM says so, then offers the documented alternatives — it never invents syntax.

03

Separation of concerns

Reference looks up, Guide teaches, Debug diagnoses. Each specialist owns one job against one source of truth, so the system stays deterministic and debuggable.

Timeline

From a parsing script to a replatform

Early 2026

Parse the 2,191-page guide

A parsing script (extract_adminguide.py) restructures the admin guide into 1,491 discrete, structured commands.

2026

Dataverse retrieval layer

The commands load into Dataverse for exact lookup, paired with semantic search for domain-discovery questions.

2026

Orchestrative multi-agent (v2)

An Orchestrator routes intent to three specialists — Reference, Guide, Debug — each grounded in the authoritative sources.

Summer 2026

v3 — skills-based, Sonnet 5

Replatformed onto the new experience: specialists become inline skills, retrieval rebuilt as a native Workflow, eval parity confirmed at 73/75.

Sep 18 2026

Live in PROD as SAM

UAT signed off by Doug Brandt; promoted DEV → TEST → PROD as a clean managed solution (SmartAssemblyMentorSAM) and renamed Smart Assembly Mentor (SAM), shared with one small-team security group across all three environments.

Where it stands

SAM is live in Copilot Agents PROD as of September 18, 2026 — the v2 multi-agent build replatformed onto the new skills-based experience on Claude Sonnet 5, with the Dataverse retrieval flow rebuilt as a native Workflow. It cleared the parity eval at 73/75, passed UAT with Doug Brandt, and was promoted DEV → TEST → PROD as a clean managed solution, renamed from SAM v3 RP to Smart Assembly Mentor (SAM). The evaluation discipline — a documented control set where any fabricated command is an automatic failure, 80% to reach TEST and 100% the goal before PROD — is what makes “grounded, not guessed” a measured claim rather than a hope.