Skip to content
All case studies

CASE STUDY / Returns-policy decisions

Better returns decisions. 40.9% lower inference cost.

Decision accuracy jumped from 60.5% to 97.5%. Inference cost fell by 40.9%.

97.5%decision accuracy
40.9%lower inference cost
74more correct decisions

The task: apply the policy, down to the refund.

A returns decision can require a denial, a wait, a request for information, manual review or a resolution. The model must choose the correct outcome and produce the exact required fields, quantities, reasons and refund total.

The baseline answered 121 of 200 decisions correctly. The optimization priorities were cost, then quality, then latency, with a minimum quality requirement of 95%.

Sharper instructions. Better reasoning. A smaller model.

Isoquant optimized the model configuration, policy instructions and reasoning settings. GPT-5.4 nano crossed the quality threshold without fine-tuning.

The selected configuration correctly answered 195 of 200 decisions. It preserved all 121 baseline successes and fixed 74 failures. On the resolution subset, correct answers increased from 6 of 40 to 38 of 40.

A 37-point accuracy gain at lower cost.

Accuracy rose from 60.5% to 97.5%, while estimated inference cost fell from $1.20 to $0.71 per 1,000 decisions. The result cleared the 95% quality requirement and reduced cost by 40.9%.

The more deliberate configuration took longer: mean completion rose from 0.853 to 2.654 seconds, and p95 rose from 1.204 to 6.664 seconds. There was no hard latency limit in the optimization objective.

MATCHED COMPARISON

By the numbers.

BASELINEGPT-5.4 mini
WITH ISOQUANTGPT-5.4 nano + Isoquant optimization
Baseline and optimized results on 200 decisions
MetricBaselineWith Isoquant
Decision accuracy60.5%97.5%
Correct / total121 / 200195 / 200
Inference cost / 1,000 tasks$1.20$0.708
Mean completion latency0.853 s2.65 s
YOUR WORKLOAD. YOUR SUCCESS CRITERIA.

Let’s build your next performance advantage.

Talk to us
Banking intent classification91.7%lower inference costRead the case study Coding and repair64.4%lower inference costRead the case study Multi-step order resolution98.7%task successRead the case study