Downshift

Call sites

8 call sites, 3 downgraded

Every call site was evaluated on the baseline and each cheaper model. A candidate must reach at least 95% of the baseline pass rate and clear a floor of 80%. The cheapest passing model wins. Judge-graded cases pass at 4 of 5 or higher.

Call siteDecision7b (base)3b1.5b0.5bSavings/mo
supportdesk/agent_assist.py · hard · judge
keep13.6%9.1%4.5%0.0%$0.00
supportdesk/agent_assist.py · medium · judge
keep59.1%22.7%0.0%4.5%$0.00
supportdesk/extract.py · medium · json_fields
downgrade95.2%100.0%81.0%71.4%$282.04
lang_offound by Bob
supportdesk/misc_utils.py · easy · exact
downgrade90.9%81.8%90.9%81.8%$163.88
supportdesk/policy.py · hard · json_fields
keep40.0%40.0%28.0%0.0%$0.00
supportdesk/triage.py · easy · exact
keep84.0%64.0%72.0%48.0%$0.00
detect_sentimentfound by Bob
supportdesk/triage.py · easy · exact
downgrade90.9%90.9%63.6%81.8%$118.73
tag_urgencyfound by Bob
supportdesk/triage.py · easy · exact
keep72.7%50.0%40.9%27.3%$0.00

Pass rate per qwen2.5 model on each call site's evals. Teal is the model Downshift picked.

Dollar figures are projections: measured token counts x illustrative per-model prices x an assumed call volume, all set in downshift.yaml. They are not a real bill.