MillerKnoll AI
Executive Summary

We chose to build, and it compounded.

A vendor wanted ~$200–240K a year to evaluate MillerKnoll specials with AI. The Lean AI team built it internally instead, and it went into production on August 22, 2026. In its first 30 days it evaluated 5,700 real dealer requests, agreed with the quoters’ actual decisions 91.7% of the time, and helped cut quote response time from about 72 hours to 26.1.

The build-vs-buy moment

At the October 2025 VP Summit, Larry Kallio (VP, Product Development & Custom) was ready to license Naya Studio — ~$50K to implement plus $200–240K/year — to give the Specials team AI pricing support. On October 27, Derek Torrey reframed it as a possible internal “medium bet,” and the Lean AI team took the challenge.

The proof of concept was stood up in days. The internal build didn’t just match the vendor’s promise — it became a platform the team owns, extends, and reuses across the agent portfolio.

From hedge to production

The Claude Managed Agents build started as insurance: the same orchestrator instructions and five skills, carried over word for word and graded on the same verified bids, so that a business-critical engine wouldn’t depend on one vendor’s pricing or roadmap. It came out ahead. Claude’s pricing made it about 30% cheaper to run, and it went into production at cutover on August 22. The Copilot Studio build is kept in step as the Azure digital twin, so the choice stays open.

Running two builds also found problems nobody went looking for. Comparing them on the same bids surfaced four real defects in the production logic before launch. The most serious one was quietly turning quotable orders into engineering-review referrals while a matching prior special was sitting in the data.

Improvements

What got measurably better

Production results from the first 30 days, alongside what the build avoided and the volume in scope.

91.7%
Agreement with the quoters’ actual decisions, first 30 days
72h → 26.1h
Quote response time (goal: under 24)
5,700
Real dealer requests evaluated since the Aug 22 cutover
$200K/yr
Vendor license avoided by building internally
~30%
Lower run cost on Claude than the Copilot Studio build
100k+
Specials RFQs a year in scope
By the numbers

Accuracy climbed as the knowledge deepened

Two curves that move together: evaluation accuracy on one side, and the Specials team hardening the knowledge base on the other — exactly the Phase-1 thesis, that putting the engine in their hands would make them own the rules.

Single-deviation accuracy

By build, toward the 90% UAT target

50%75%100%target 90%61%73%83%88%POC0.21.01.4

Specials KB activity per month

Pages created vs. updated · Jul 2025 – Jun 2026 (June, full month)

UpdatedCreated
0651292523311927335680629512985141017111722223526486235Jul’25Jan’26Jun
Success at scale

Solving multi-deviation is the whole game

Most real specials carry more than one custom change — multi-deviation success, not raw single-string accuracy, is the number that tells us this is real. It has climbed with every architecture generation, and the control-question set that grades it is growing to match.

Multi-deviation success · separate control suite

Confirmed by version · 1.4 → 2.0 → 2.0RP

20%60%100%38%55%72%81%1.42.02.0RP2.1

2.1 testing2.1’s multi-deviation cases are now graded through the Product Portal API itself — the same path the business will use — and run several passes per case, because retrieval and complexity can still vary on identical input and a single run proves nothing either way. The numbers shown are confirmed results per version.

Control-question coverage

Growing the suite to match the real breadth of Specials volume — not a sample.

40 original
100 at UAT
400 today
~1,000 target
Why it works

Three principles held throughout

01

Determinism first

The engine cites documented rules and never invents an approval or a price. When the knowledge isn’t there, it abstains to Engineering Review rather than guess.

02

Separation of concerns

Each agent owns one job — enrich, classify, score complexity, price. The orchestrator coordinates and makes no decision itself.

03

Grounded in real precedent

“Have we done this before?” The engine searches prior completed specials, so a documented precedent becomes a first-class approval signal.

Key moments

The decisions that shaped it

Oct 27 2025

Build vs. buy

Redirect the $200K/yr Naya spend into an internal Lean AI build.

Dec 2025

Determinism-first KB

Cite-or-abstain; tables over prose; complexity encoded in article titles.

Jan 2026

Ship it to the team

Phase-1 goal: put it in Specials’ hands in TEST so they maintain the rules behind their work.

Apr 2026

RAG → multi-agent

RAG alone was too limiting; move to an orchestrator + children and expand pricing/complexity logic.

Jun 2026

Rovo + Dataverse

Deviations Knowledge moves to Rovo/Teamwork Graph; pricing falls back to a deterministic Dataverse lookup.

Jun 13 2026

Sonnet over GPT-5

Chose Sonnet for search accuracy + instruction-following over raw speed.

Jul 2026

Replatform to 2.1

Rebuild on skills-based Copilot Studio: five skills under one orchestrator, native Workflow tools, deterministic retrieval — on Sonnet 5.

Jul 2026

Real-dollar pricing

The Dynamic Pricing API resolves the starting reference’s list price, so the engine computes a real total (list + specials 10% + upcharges).

Jul 23 2026

Finish codes, no spreadsheet

A live spike proved Pub itself resolves a finish request to its MillerKnoll stain code (match “Classic Oak 143 laminate” → EJ25). Shipped into the Enrichment skill and tested the same day — no new tool, and the unreachable SharePoint spreadsheet dependency is gone.

Jul 24 2026

Rebuild clean, don’t untangle

The agent had drifted across two publishers and could no longer be exported. Two sessions of patching it in place had produced a regression; rebuilding it from scratch under one publisher took an afternoon and removed the whole class of problem. It is what made promotion possible.

Aug 1 2026

Store the payload, parse it later

The return trip stores what the business actually quoted verbatim rather than parsing it into typed columns. The payload shape isn’t known yet, and a parser written against a guess fills an accuracy record with plausible-but-wrong values — worse than leaving it empty.

Aug 4 2026

A green response is not a right answer

The engine answered through the API path confidently and wrongly in TEST — well-formed, complete, and ungrounded, because promoted retrieval tools silently pointed back at the environment they came from. Fixed, and the test suite now asserts a known answer instead of an HTTP contract.

Aug 5 2026

Complexity only for what we’re building

A business ruling: a deviation that isn’t approved has no complexity to assess, because nobody is making it. Denials had been scoring Complex and dragging whole requests with them. Paired with a second ruling — a documented limit on an approved change is a condition the dealer must meet, not a reason for refusal.

Jul–Aug 2026

A second build as a hedge

The same engine carried over to Claude Managed Agents and graded on the same bids. Insurance against platform pricing and roadmap shifts, and unexpectedly the best test harness on the project: four real defects found by comparing two independent builds.

Aug 2026

Claude takes production

Claude’s pricing made the second build about 30% cheaper to run, so it became the production engine. The Copilot Studio build stays in step as the Azure digital twin.

Aug 22 2026

Cutover

Live on Saturday at noon. First weekend: every request evaluated, none lost, no manual intervention.

Aug 22 – Sep 2026

Hypercare with the business

Daily testing with the Specials and Product Portal teams, and more than 40 fixes promoted DEV → TEST → PROD in the first month, each announced in a plain-language changelog. Five production pushes landed in the first six days alone.

Aug 30 2026

The right model for each job

Catalog and precedent work moved to a Haiku sub-agent. Same decisions, 58% cheaper on the same input.

Sep 2 2026

Speak the Portal’s language

The engine’s decisions now use the Portal’s own statuses, so accuracy is scored directly against the quoters with no translation layer.

Sep 2026

The Solutions Desk

With the engine taking the first pass on research, the team launched a Solutions Desk: dealers can book time to talk a special through with a person and find a way to make it work.

Where it stands

AutoBid is in production. It went live on August 22, 2026 on Claude Managed Agents, which came in about 30% cheaper to run than the Copilot Studio build. That build is kept in step as the Azure digital twin. In its first 30 days the engine evaluated 5,700 real dealer requests and agreed with the quoters’ actual decision type 91.7% of the time, above the pre-launch control-testing estimate. Quote response time is down from about 72 hours to 26.1, with under 24 as the goal. Hypercare ran as a daily collaboration with the Specials and Product Portal teams, and more than 40 fixes shipped through DEV, TEST and PROD in the first month. The time the engine frees up is going back to dealers: the Solutions Desk now takes appointments to work through specials in person. Next is selective autonomy on patterns the quoters have already approved, and writing the knowledge-base articles the engine has shown are missing.