Don't take our word.
Check our work.

No marketing claims, no cherry-picked testimonials. We test Guidon blind against published Board of Veterans' Appeals decisions and publish everything — including the misses. We believe we're the only company in this space that does. We'd love to stop being the only one.

The method

Current results — 24 decisions, 92 issues (updated June 2026)

93% codes identified 40 of 43 rated issues 98% ratings covered 39 of 40 rated issues 0 fabricated criteria across 92 issues
MetricResult
Board's diagnostic code identified (43 rated issues)93% (40 of 43)
Board's exact percentage among Guidon's candidates (40 rated issues)98% (39 of 40)
Fabricated or invented rating criteria, across all 92 issuesZero — Guidon quotes only regulation text it actually retrieved

Issues where the Board assigned no diagnostic code or percentage (dismissals, remands, service-connection-only rulings) can't be scored against this metric and are excluded: 48 of 92.

The misses — all of them

CaseWhat happenedOur read
Dry eye syndromeBoard rated under code 6018; Guidon proposed the adjacent code 6025A true sibling-code miss. The deciding clinical detail wasn't carried in the evidence summary. Logged as a regression case.
Ankle instability (two issues, one case)Board rated by analogy under 5273; Guidon proposed 5003-5271 — and explicitly flagged that the record lacked the facts needed to choose, listing exactly which clinical findings would decide itWorking as designed: when the record doesn't decide between adjacent codes, Guidon asks for the deciding fact rather than guessing. The Board, with the complete file, chose differently. We count it as a miss anyway.
One staged-rating percentageOne 0%-for-a-period candidate missing from an otherwise correctly staged conditionTemporal staging shipped in June 2026 and converted seven of eight staging misses; this is the remainder. Regression case logged.

Beyond the benchmark: testing on real files

Published Board decisions are our public scorecard. But we also pressure-test Guidon on real, de-identified veteran files — handled under the protections on our security page, never published, never used to train anything. Real files are messier than any benchmark, and they're where the engine earns its trust.

A recent test is worth telling on, because of what it caught. Given an unfamiliar veteran's records, Guidon matched every condition the VA had rated, identified the effective dates correctly, and flagged every condition the VA had denied as needing specific additional evidence — rather than overselling them as easy wins. Then it made a mistake: one recommendation suggested pursuing a disability rating the veteran already held.

That is exactly the kind of error a seasoned service officer would catch in a heartbeat — and exactly the kind that erodes trust in a tool. We caught it in testing. Within the hour, we built a fix: Guidon now reconciles every recommendation against the VA's existing decisions, so it never suggests chasing a rating already granted, and it reframes denied conditions as "reopen with this specific evidence" rather than a naive re-file. We verified the fix and moved on. No veteran's data appears here or anywhere public. The lesson does.

What this benchmark has already changed

This page isn't decoration — it's the engine's development process. Earlier rounds exposed a retrieval failure on mixed-condition names (fixed: every candidate code now gets its governing regulation section guaranteed into context) and the absence of staged-rating support (built and verified the same week). Every future miss follows the same path: diagnosed, fixed, re-verified blind, published here.

Verify any of it. The decisions are public documents on va.gov. Ask us for the full per-issue results and we'll send them — citation numbers included.