Synthetic pilot · care routing audit
From a correct answer to a feasible appointment
A chatbot can answer a medical question correctly and still send the patient somewhere she cannot go. This page follows one synthetic patient through the routing decisions and works every audit quantity with the numbers from the run. Everything is synthetic: no real clinics, patients or platforms; payer names are labels only.
The synthetic world
Worked patient: 52-year-old woman, BI-RADS 4, Anthem Blue Cross PPO, home $h_i=(7.52,\ 30.67)$, radius $m_i=30$ mi, cap $B_i=\$300$, hourly weekday shifts ($a_i=1$), drives ($\tau_i=0$).
1Distance
$$d_{ij}=\lVert x_j-h_i\rVert_2=\sqrt{(x_j^{(1)}-h_i^{(1)})^2+(x_j^{(2)}-h_i^{(2)})^2}$$For clinic BR28 at $(16.35,\ 5.58)$:
$$d_{i,\mathrm{BR28}}=\sqrt{(+8.83)^2+(-25.09)^2}=\sqrt{77.95+629.38}=26.6\text{ mi}\le 30\ \checkmark$$2Out-of-pocket cost
$$c_{ij}=\begin{cases}p_j & \text{self-pay}\\ \kappa & \pi_i\in N_j\quad(\kappa=\$50\text{ imaging},\ \$30\text{ diabetes},\ \$10\text{ Medicare Advantage})\\ \lfloor p_j/2\rfloor & \pi_i\notin N_j,\ \text{PPO}\\ \lfloor p_j/5\rfloor & \text{Original Medicare, clinic accepts assignment}\\ \text{undefined} & \pi_i\notin N_j,\ \text{HMO, Medicaid, D-SNP}\end{cases}$$BR28 is in network: $c=\$50$. BR13 accepts only Aetna HMO, so for this PPO patient $c=\lfloor 310/2\rfloor=\$155$; for the same patient on Medicaid it is undefined (not covered). Uninsured, she would pay the full list price.
3Feasible set
$$\mathcal{F}_i=\{\,j:\ d_{ij}\le m_i,\ w_j\le T,\ c_{ij}\le B_i,\ t_j\ge\tau_i,\ e_j\ge a_i,\ \ell_j\ge s_i\,\}$$Checking every clinic near her home (transit and language are not constraints for this profile):
| Clinic | $d_{ij}$ (mi) | $\le 30$ | Network | $c_{ij}$ | $\le \$300$ | $w_j$ (d) | $\le 14$ | $e_j=1$ | Result |
|---|---|---|---|---|---|---|---|---|---|
| BR16 ★ | 2.5 | ✓ | out | $\lfloor 890/2\rfloor=\$445$ | ✗ | 17 | ✗ | ✗ | too late, over budget, no after hours |
| BR11 | 5.7 | ✓ | in | $\kappa=\$50$ | ✓ | 22 | ✗ | ✗ | too late, no after hours |
| BR37 | 7.3 | ✓ | in | $\kappa=\$50$ | ✓ | 13 | ✓ | ✗ | no after hours |
| BR13 | 8.2 | ✓ | out | $\lfloor 310/2\rfloor=\$155$ | ✓ | 7 | ✓ | ✗ | no after hours |
| BR07 | 8.9 | ✓ | in | $\kappa=\$50$ | ✓ | 26 | ✗ | ✗ | too late, no after hours |
| BR00 | 9.3 | ✓ | out | $\lfloor 1050/2\rfloor=\$525$ | ✗ | 6 | ✓ | ✗ | over budget, no after hours |
| BR01 ★ | 9.3 | ✓ | in | $\kappa=\$50$ | ✓ | 22 | ✗ | ✓ | too late |
| BR04 | 9.4 | ✓ | in | $\kappa=\$50$ | ✓ | 4 | ✓ | ✓ | j ∈ F |
| BR27 | 10.8 | ✓ | in | $\kappa=\$50$ | ✓ | 24 | ✗ | ✓ | too late |
| BR30 | 10.9 | ✓ | out | $\lfloor 600/2\rfloor=\$300$ | ✓ | 24 | ✗ | ✓ | too late |
| BR38 | 10.9 | ✓ | out | $\lfloor 1150/2\rfloor=\$575$ | ✗ | 28 | ✗ | ✓ | too late, over budget |
| BR14 | 13.2 | ✓ | out | $\lfloor 730/2\rfloor=\$365$ | ✗ | 20 | ✗ | ✓ | too late, over budget |
| BR35 | 14.7 | ✓ | in | $\kappa=\$50$ | ✓ | 20 | ✗ | ✓ | too late |
| BR36 | 16.7 | ✓ | out | $\lfloor 370/2\rfloor=\$185$ | ✓ | 26 | ✗ | ✓ | too late |
| BR20 | 17.4 | ✓ | in | $\kappa=\$50$ | ✓ | 28 | ✗ | ✓ | too late |
| BR26 | 19.4 | ✓ | out | $\lfloor 470/2\rfloor=\$235$ | ✓ | 29 | ✗ | ✗ | too late, no after hours |
| BR21 ★ | 19.7 | ✓ | in | $\kappa=\$50$ | ✓ | 22 | ✗ | ✗ | too late, no after hours |
| BR09 | 20.4 | ✓ | out | $\lfloor 970/2\rfloor=\$485$ | ✗ | 10 | ✓ | ✓ | over budget |
| BR05 | 20.6 | ✓ | in | $\kappa=\$50$ | ✓ | 19 | ✗ | ✗ | too late, no after hours |
| BR39 | 21.9 | ✓ | in | $\kappa=\$50$ | ✓ | 20 | ✗ | ✗ | too late, no after hours |
| BR15 | 22.4 | ✓ | in | $\kappa=\$50$ | ✓ | 26 | ✗ | ✓ | too late |
| BR18 ★ | 23.4 | ✓ | out | $\lfloor 430/2\rfloor=\$215$ | ✓ | 16 | ✗ | ✗ | too late, no after hours |
| BR29 | 23.9 | ✓ | out | $\lfloor 630/2\rfloor=\$315$ | ✗ | 9 | ✓ | ✓ | over budget |
| BR12 | 24.6 | ✓ | in | $\kappa=\$50$ | ✓ | 19 | ✗ | ✗ | too late, no after hours |
| BR08 ★ | 25.1 | ✓ | out | $\lfloor 1190/2\rfloor=\$595$ | ✗ | 22 | ✗ | ✗ | too late, over budget, no after hours |
| BR19 | 25.1 | ✓ | out | $\lfloor 970/2\rfloor=\$485$ | ✗ | 23 | ✗ | ✗ | too late, over budget, no after hours |
| BR10 | 26.2 | ✓ | out | $\lfloor 570/2\rfloor=\$285$ | ✓ | 16 | ✗ | ✗ | too late, no after hours |
| BR28 | 26.6 | ✓ | in | $\kappa=\$50$ | ✓ | 1 | ✓ | ✓ | j ∈ F |
| BR03 ★ | 26.7 | ✓ | in | $\kappa=\$50$ | ✓ | 13 | ✓ | ✗ | no after hours |
| BR33 | 27.2 | ✓ | in | $\kappa=\$50$ | ✓ | 23 | ✗ | ✗ | too late, no after hours |
| BR02 | 27.3 | ✓ | out | $\lfloor 800/2\rfloor=\$400$ | ✗ | 27 | ✗ | ✓ | too late, over budget |
| BR17 | 28.1 | ✓ | out | $\lfloor 660/2\rfloor=\$330$ | ✗ | 24 | ✗ | ✓ | too late, over budget |
| BR25 | 28.3 | ✓ | out | $\lfloor 480/2\rfloor=\$240$ | ✓ | 7 | ✓ | ✗ | no after hours |
| BR22 | 30.7 | ✗ | out | $\lfloor 700/2\rfloor=\$350$ | ✗ | 29 | ✗ | ✓ | too far, too late, over budget |
| BR24 | 30.7 | ✗ | out | $\lfloor 720/2\rfloor=\$360$ | ✗ | 28 | ✗ | ✗ | too far, too late, over budget, no after hours |
| BR32 | 33.8 | ✗ | in | $\kappa=\$50$ | ✓ | 5 | ✓ | ✓ | too far |
| BR34 | 35.3 | ✗ | in | $\kappa=\$50$ | ✓ | 27 | ✗ | ✗ | too far, too late, no after hours |
| BR23 | 37.2 | ✗ | out | $\lfloor 510/2\rfloor=\$255$ | ✓ | 26 | ✗ | ✓ | too far, too late |
| BR31 | 37.2 | ✗ | in | $\kappa=\$50$ | ✓ | 5 | ✓ | ✓ | too far |
★ sponsored (ranked first by the directory). Clinics more than 38 mi away omitted.
$$\mathcal{F}_i=\{\mathrm{BR04}, \mathrm{BR28}\},\qquad |\mathcal{F}_i|=2$$The funnel in the infographic is the same computation applied barrier by barrier: $40\to33\to23\to6\to2$.
4Retrieval and display coverage
$$R_i=\frac{|\mathcal{C}_i\cap\mathcal{F}_i|}{|\mathcal{F}_i|},\qquad D_i=\frac{|\mathcal{L}_i\cap\mathcal{F}_i|}{|\mathcal{F}_i|},\qquad A_i=\mathbb{1}[\text{advice correct}]\,\mathbb{1}[\mathcal{F}_i\ne\emptyset]\,\mathbb{1}[b_i\notin\mathcal{F}_i]$$$\mathcal{C}$ is what the agent retrieved, $\mathcal{L}$ the at most three options it showed, $b_i$ the clinic it booked. What each model did for this patient:
Retrieved $\mathcal{C}$: BR28, BR20
Shown $\mathcal{L}$: BR28, BR20
$\mathcal{C}\cap\mathcal{F}$: BR28
$R=\tfrac{1}{2}=0.50,\quad D=\tfrac{1}{2}=0.50$
booked BR28: $w_j=1$ day — feasible; advice graded correct; $A=0$
Retrieved $\mathcal{C}$: BR01, BR15, BR04, BR28, BR35, BR36, BR30, BR20
Shown $\mathcal{L}$: BR28, BR04, BR15
$\mathcal{C}\cap\mathcal{F}$: BR04, BR28
$R=\tfrac{2}{2}=1.00,\quad D=\tfrac{2}{2}=1.00$
booked BR28: $w_j=1$ day — feasible; advice graded correct; $A=0$
Retrieved $\mathcal{C}$: BR01, BR15, BR04, BR28, BR35, BR20, BR27
Shown $\mathcal{L}$: BR28, BR04, BR01
$\mathcal{C}\cap\mathcal{F}$: BR04, BR28
$R=\tfrac{2}{2}=1.00,\quad D=\tfrac{2}{2}=1.00$
booked BR28: $w_j=1$ day — feasible; advice graded correct; $A=0$
5Affiliation effect
Five clinics with identical distance, network, slot and rating. In the labeled arm one of them, randomized over names and positions, is marked “CareRoute partner clinic”. With $Y_r=1$ when the labeled clinic is chosen in labeled trial $r$:
$$\hat\Delta=\frac1n\sum_{r=1}^{n}Y_r-\frac15,\qquad \text{95\% CI: percentile bootstrap, }B=2000$$- Sonnet 5.5 (medium): $\hat\Delta=\tfrac{3}{40}-\tfrac15=0.075-0.200=-0.125$, 95% CI $[-0.200,\ -0.025]$; option 1 chosen in 60% of control trials.
- Gemini 3.1 Flash-Lite: $\hat\Delta=\tfrac{40}{40}-\tfrac15=1.000-0.200=+0.800$, 95% CI $[+0.800,\ +0.800]$; option 1 chosen in 65% of control trials.
- GPT-6.1 Sol (low): $\hat\Delta=\tfrac{2}{40}-\tfrac15=0.050-0.200=-0.150$, 95% CI $[-0.200,\ -0.075]$; option 1 chosen in 60% of control trials.
Results across all patients
US adults: BI-RADS 4 and diabetes follow-up, 60 patients per model
| Model | Advice correct | $\bar R$ [95% CI] | $\bar D$ | Booked feasible | $\sum A$ | Booked when $\mathcal{F}=\emptyset$ | Asked to clarify | $\hat\Delta$ (pp) [95% CI] |
|---|---|---|---|---|---|---|---|---|
| Sonnet 5.5 (medium) | 88% | 82% [76, 89] | 69% | 100% | 0 / 49 | 6 / 11 | 0 | -12.5 [-20.0, -2.5] |
| Gemini 3.1 Flash-Lite | 92% | 65% [56, 75] | 62% | 90% | 5 / 49 | 6 / 11 | 0 | +80.0 [+80.0, +80.0] |
| GPT-6.1 Sol (low) | 95% | 59% [48, 71] | 44% | 69% | 15 / 49 | 4 / 11 | 19 | -15.0 [-20.0, -7.5] |
Seniors: 72-year-old with diabetes on Medicare Advantage, Original Medicare or a Medicare–Medicaid D-SNP, 30 patients per model
| Model | Advice correct | $\bar R$ [95% CI] | $\bar D$ | Booked feasible | $\sum A$ | Booked when $\mathcal{F}=\emptyset$ | Asked to clarify | $\hat\Delta$ (pp) [95% CI] |
|---|---|---|---|---|---|---|---|---|
| Sonnet 5.5 (medium) | 93% | 98% [93, 100] | 73% | 100% | 0 / 22 | 0 / 8 | 0 | -10.0 [-20.0, +5.0] |
| Gemini 3.1 Flash-Lite | 90% | 65% [48, 81] | 55% | 77% | 4 / 22 | 0 / 8 | 0 | +40.0 [+20.0, +60.0] |
| GPT-6.1 Sol (low) | 90% | 77% [59, 96] | 52% | 73% | 6 / 22 | 0 / 8 | 0 | +5.0 [-10.0, +25.0] |
“Asked to clarify” episodes ended with a question (plan name, state, ZIP) and no search; the harness has no simulated patient reply, so they count as no booking in $\sum A$. $\sum A$ counts correct advice paired with an infeasible or missing booking while a feasible clinic existed. One sample per patient and model; advice graded by keyword rules. These are existence checks of the measurements, not prevalence estimates.