CASE STUDY / Returns-policy decisions
Better returns decisions. 40.9% lower inference cost.
Decision accuracy jumped from 60.5% to 97.5%. Inference cost fell by 40.9%.
The task: apply the policy, down to the refund.
A returns decision can require a denial, a wait, a request for information, manual review or a resolution. The model must choose the correct outcome and produce the exact required fields, quantities, reasons and refund total.
The baseline answered 121 of 200 decisions correctly. The optimization priorities were cost, then quality, then latency, with a minimum quality requirement of 95%.
Sharper instructions. Better reasoning. A smaller model.
Isoquant optimized the model configuration, policy instructions and reasoning settings. GPT-5.4 nano crossed the quality threshold without fine-tuning.
The selected configuration correctly answered 195 of 200 decisions. It preserved all 121 baseline successes and fixed 74 failures. On the resolution subset, correct answers increased from 6 of 40 to 38 of 40.
A 37-point accuracy gain at lower cost.
Accuracy rose from 60.5% to 97.5%, while estimated inference cost fell from $1.20 to $0.71 per 1,000 decisions. The result cleared the 95% quality requirement and reduced cost by 40.9%.
The more deliberate configuration took longer: mean completion rose from 0.853 to 2.654 seconds, and p95 rose from 1.204 to 6.664 seconds. There was no hard latency limit in the optimization objective.
By the numbers.
| Metric | Baseline | With Isoquant |
|---|---|---|
| Decision accuracy | 60.5% | 97.5% |
| Correct / total | 121 / 200 | 195 / 200 |
| Inference cost / 1,000 tasks | $1.20 | $0.708 |
| Mean completion latency | 0.853 s | 2.65 s |