Call sites
8 call sites, 3 downgraded
Every call site was evaluated on the baseline and each cheaper model. A candidate must reach at least 95% of the baseline pass rate and clear a floor of 80%. The cheapest passing model wins. Judge-graded cases pass at 4 of 5 or higher.
| Call site | Decision | 7b (base) | 3b | 1.5b | 0.5b | Savings/mo |
|---|---|---|---|---|---|---|
supportdesk/agent_assist.py · hard · judge | keep | 13.6% | 9.1% | 4.5% | 0.0% | $0.00 |
supportdesk/agent_assist.py · medium · judge | keep | 59.1% | 22.7% | 0.0% | 4.5% | $0.00 |
supportdesk/extract.py · medium · json_fields | downgrade | 95.2% | 100.0% | 81.0% | 71.4% | $282.04 |
lang_offound by Bob supportdesk/misc_utils.py · easy · exact | downgrade | 90.9% | 81.8% | 90.9% | 81.8% | $163.88 |
supportdesk/policy.py · hard · json_fields | keep | 40.0% | 40.0% | 28.0% | 0.0% | $0.00 |
supportdesk/triage.py · easy · exact | keep | 84.0% | 64.0% | 72.0% | 48.0% | $0.00 |
detect_sentimentfound by Bob supportdesk/triage.py · easy · exact | downgrade | 90.9% | 90.9% | 63.6% | 81.8% | $118.73 |
tag_urgencyfound by Bob supportdesk/triage.py · easy · exact | keep | 72.7% | 50.0% | 40.9% | 27.3% | $0.00 |
Pass rate per qwen2.5 model on each call site's evals. Teal is the model Downshift picked.
Dollar figures are projections: measured token counts x illustrative per-model prices x an assumed call volume, all set in downshift.yaml. They are not a real bill.