We chose to build, and it compounded.
A vendor wanted ~$200–240K a year to evaluate MillerKnoll specials with AI. The Lean AI team built it internally instead, and it went into production on August 22, 2026. In its first 30 days it evaluated 5,700 real dealer requests, agreed with the quoters’ actual decisions 91.7% of the time, and helped cut quote response time from about 72 hours to 26.1.
At the October 2025 VP Summit, Larry Kallio (VP, Product Development & Custom) was ready to license Naya Studio — ~$50K to implement plus $200–240K/year — to give the Specials team AI pricing support. On October 27, Derek Torrey reframed it as a possible internal “medium bet,” and the Lean AI team took the challenge.
The proof of concept was stood up in days. The internal build didn’t just match the vendor’s promise — it became a platform the team owns, extends, and reuses across the agent portfolio.
The Claude Managed Agents build started as insurance: the same orchestrator instructions and five skills, carried over word for word and graded on the same verified bids, so that a business-critical engine wouldn’t depend on one vendor’s pricing or roadmap. It came out ahead. Claude’s pricing made it about 30% cheaper to run, and it went into production at cutover on August 22. The Copilot Studio build is kept in step as the Azure digital twin, so the choice stays open.
Running two builds also found problems nobody went looking for. Comparing them on the same bids surfaced four real defects in the production logic before launch. The most serious one was quietly turning quotable orders into engineering-review referrals while a matching prior special was sitting in the data.
What got measurably better
Production results from the first 30 days, alongside what the build avoided and the volume in scope.
Accuracy climbed as the knowledge deepened
Two curves that move together: evaluation accuracy on one side, and the Specials team hardening the knowledge base on the other — exactly the Phase-1 thesis, that putting the engine in their hands would make them own the rules.
Single-deviation accuracy
By build, toward the 90% UAT target
Specials KB activity per month
Pages created vs. updated · Jul 2025 – Jun 2026 (June, full month)
Solving multi-deviation is the whole game
Most real specials carry more than one custom change — multi-deviation success, not raw single-string accuracy, is the number that tells us this is real. It has climbed with every architecture generation, and the control-question set that grades it is growing to match.
Multi-deviation success · separate control suite
Confirmed by version · 1.4 → 2.0 → 2.0RP
2.1 testing2.1’s multi-deviation cases are now graded through the Product Portal API itself — the same path the business will use — and run several passes per case, because retrieval and complexity can still vary on identical input and a single run proves nothing either way. The numbers shown are confirmed results per version.
Control-question coverage
Growing the suite to match the real breadth of Specials volume — not a sample.
Three principles held throughout
Determinism first
The engine cites documented rules and never invents an approval or a price. When the knowledge isn’t there, it abstains to Engineering Review rather than guess.
Separation of concerns
Each agent owns one job — enrich, classify, score complexity, price. The orchestrator coordinates and makes no decision itself.
Grounded in real precedent
“Have we done this before?” The engine searches prior completed specials, so a documented precedent becomes a first-class approval signal.
The decisions that shaped it
Build vs. buy
Redirect the $200K/yr Naya spend into an internal Lean AI build.
Determinism-first KB
Cite-or-abstain; tables over prose; complexity encoded in article titles.
Ship it to the team
Phase-1 goal: put it in Specials’ hands in TEST so they maintain the rules behind their work.
RAG → multi-agent
RAG alone was too limiting; move to an orchestrator + children and expand pricing/complexity logic.
Rovo + Dataverse
Deviations Knowledge moves to Rovo/Teamwork Graph; pricing falls back to a deterministic Dataverse lookup.
Sonnet over GPT-5
Chose Sonnet for search accuracy + instruction-following over raw speed.
Replatform to 2.1
Rebuild on skills-based Copilot Studio: five skills under one orchestrator, native Workflow tools, deterministic retrieval — on Sonnet 5.
Real-dollar pricing
The Dynamic Pricing API resolves the starting reference’s list price, so the engine computes a real total (list + specials 10% + upcharges).
Finish codes, no spreadsheet
A live spike proved Pub itself resolves a finish request to its MillerKnoll stain code (match “Classic Oak 143 laminate” → EJ25). Shipped into the Enrichment skill and tested the same day — no new tool, and the unreachable SharePoint spreadsheet dependency is gone.
Rebuild clean, don’t untangle
The agent had drifted across two publishers and could no longer be exported. Two sessions of patching it in place had produced a regression; rebuilding it from scratch under one publisher took an afternoon and removed the whole class of problem. It is what made promotion possible.
Store the payload, parse it later
The return trip stores what the business actually quoted verbatim rather than parsing it into typed columns. The payload shape isn’t known yet, and a parser written against a guess fills an accuracy record with plausible-but-wrong values — worse than leaving it empty.
A green response is not a right answer
The engine answered through the API path confidently and wrongly in TEST — well-formed, complete, and ungrounded, because promoted retrieval tools silently pointed back at the environment they came from. Fixed, and the test suite now asserts a known answer instead of an HTTP contract.
Complexity only for what we’re building
A business ruling: a deviation that isn’t approved has no complexity to assess, because nobody is making it. Denials had been scoring Complex and dragging whole requests with them. Paired with a second ruling — a documented limit on an approved change is a condition the dealer must meet, not a reason for refusal.
A second build as a hedge
The same engine carried over to Claude Managed Agents and graded on the same bids. Insurance against platform pricing and roadmap shifts, and unexpectedly the best test harness on the project: four real defects found by comparing two independent builds.
Claude takes production
Claude’s pricing made the second build about 30% cheaper to run, so it became the production engine. The Copilot Studio build stays in step as the Azure digital twin.
Cutover
Live on Saturday at noon. First weekend: every request evaluated, none lost, no manual intervention.
Hypercare with the business
Daily testing with the Specials and Product Portal teams, and more than 40 fixes promoted DEV → TEST → PROD in the first month, each announced in a plain-language changelog. Five production pushes landed in the first six days alone.
The right model for each job
Catalog and precedent work moved to a Haiku sub-agent. Same decisions, 58% cheaper on the same input.
Speak the Portal’s language
The engine’s decisions now use the Portal’s own statuses, so accuracy is scored directly against the quoters with no translation layer.
The Solutions Desk
With the engine taking the first pass on research, the team launched a Solutions Desk: dealers can book time to talk a special through with a person and find a way to make it work.
Where it stands
AutoBid is in production. It went live on August 22, 2026 on Claude Managed Agents, which came in about 30% cheaper to run than the Copilot Studio build. That build is kept in step as the Azure digital twin. In its first 30 days the engine evaluated 5,700 real dealer requests and agreed with the quoters’ actual decision type 91.7% of the time, above the pre-launch control-testing estimate. Quote response time is down from about 72 hours to 26.1, with under 24 as the goal. Hypercare ran as a daily collaboration with the Specials and Product Portal teams, and more than 40 fixes shipped through DEV, TEST and PROD in the first month. The time the engine frees up is going back to dealers: the Solutions Desk now takes appointments to work through specials in person. Next is selective autonomy on patterns the quoters have already approved, and writing the knowledge-base articles the engine has shown are missing.
