We point the AI at our own work, too.
AutoBid isn’t only an AI product — it was built with AI. These are the accelerators that let a small team read thousands of documents, mine months of quotes, and engineer real evaluation suites in days instead of weeks.
Knowledge audit at scale
Read 3,300+ Specials articles to surface rule conflicts, complexity-label mismatches, trends, and gaps — distilled into the de-facto spec and the determinism-first article template.
Pricing & upcharge analysis
Analyzed a two-month, ~9,349-line Jarvis quote export to validate the 10% rule, the ~82%-Simple distribution, and the baseline upcharges — the controls and pricing-inference ground truth.
Evaluation engineering
Built control-question evaluation suites — grown from an original 40, to 100 at UAT, to 400 today — targeting nearly 1,000 by end of July to test against the real breadth of Specials volume, not a sample.
Self-diagnosis loop
When an answer was wrong, unclear, or slow, testers triggered a prompt asking the agent what would have made the decision easier — feeding straight back into knowledge-base fixes.
Knowledge hardened at 2–4× pace
The engine gave the Specials team a reason to tighten the rules behind their work: 7,039 edits across 979 pages, with 465 distinct pages updated in 2026 alone.
Deterministic retrieval
Mined the article-title corpus to make retrieval repeatable: a bare deviation term matches 20–30+ articles across unrelated lines, so a fixed prefix + canonical Search 0 anchor now runs first — the same bid returns the same article instead of drifting run-to-run.
Two builds, graded against each other
The same engine on two platforms, run on the same bids and compared response by response with both reasoning traces read side by side. Disagreement between two independent builds catches defects that testing inside one build can’t. It found four, including a precedent search that was turning quotable orders into engineering referrals, and the comparison is why the team could move production to the cheaper platform with confidence.
Known-answer verification
A contract-only smoke test passed on five consecutive runs where the engine retrieved nothing and answered anyway — well-formed, confident, wrong. Every environment is now probed with a bid whose correct answer is already known, several passes at a time, before anyone claims it works.
