MillerKnoll AI
AI Leverage

We point the AI at our own work, too.

AutoBid isn’t only an AI product — it was built with AI. These are the accelerators that let a small team read thousands of documents, mine months of quotes, and engineer real evaluation suites in days instead of weeks.

Teamwork Graph / Rovo
3,300+ articles read

Knowledge audit at scale

Read 3,300+ Specials articles to surface rule conflicts, complexity-label mismatches, trends, and gaps — distilled into the de-facto spec and the determinism-first article template.

Jarvis quote mining
~9,349 quote lines

Pricing & upcharge analysis

Analyzed a two-month, ~9,349-line Jarvis quote export to validate the 10% rule, the ~82%-Simple distribution, and the baseline upcharges — the controls and pricing-inference ground truth.

Control suites + eval passes
40 → 100 → 400 → ~1,000

Evaluation engineering

Built control-question evaluation suites — grown from an original 40, to 100 at UAT, to 400 today — targeting nearly 1,000 by end of July to test against the real breadth of Specials volume, not a sample.

The agent improves its own KB
closed-loop refinement

Self-diagnosis loop

When an answer was wrong, unclear, or slow, testers triggered a prompt asking the agent what would have made the decision easier — feeding straight back into knowledge-base fixes.

Team-maintained rules
7,039 edits · 979 pages

Knowledge hardened at 2–4× pace

The engine gave the Specials team a reason to tighten the rules behind their work: 7,039 edits across 979 pages, with 465 distinct pages updated in 2026 alone.

Canonical search-term anchoring
487 titles analyzed

Deterministic retrieval

Mined the article-title corpus to make retrieval repeatable: a bare deviation term matches 20–30+ articles across unrelated lines, so a fixed prefix + canonical Search 0 anchor now runs first — the same bid returns the same article instead of drifting run-to-run.

Claude (production) × Copilot Studio (Azure twin)
4 defects, both sides

Two builds, graded against each other

The same engine on two platforms, run on the same bids and compared response by response with both reasoning traces read side by side. Disagreement between two independent builds catches defects that testing inside one build can’t. It found four, including a precedent search that was turning quotable orders into engineering referrals, and the comparison is why the team could move production to the cheaper platform with confidence.

Grounding assertions per environment
green ≠ correct

Known-answer verification

A contract-only smoke test passed on five consecutive runs where the engine retrieved nothing and answered anyway — well-formed, confident, wrong. Every environment is now probed with a bid whose correct answer is already known, several passes at a time, before anyone claims it works.