CASE STUDY / Coding and repair
Same coding success. 64.4% lower inference cost.
A smaller model matched 95.9% task success across 512 coding and repair tasks, at 64.4% lower inference cost.
The task: write, test and repair working code.
These tasks require more than a plausible code snippet. The model implements Python, runs checks and repairs failures over a multi-turn episode. Success is determined by executable final checks.
The original GPT-5.4 mini configuration passed 491 of 512 tasks. Isoquant searched for a lower-cost configuration that preserved that level of task success.
A smaller model with a better execution strategy.
Isoquant compared model and instruction combinations, then selected GPT-5.4 nano with a specification-focused checklist. No fine-tuning was needed.
The selected configuration also passed 491 of 512 tasks. Complete-task inference cost, including repair calls, dropped from $4.34 to $1.55 per 1,000 tasks.
The same success rate, with a different cost profile.
The result was 64.4% lower estimated inference cost at the same 95.9% aggregate task success. Mean task completion took 9.31 seconds rather than 6.60 seconds. The objective favored cost while preserving quality, with no latency ceiling.
The task-level results were not identical: 20 previous failures were fixed, and 20 previous successes regressed. Isoquant also retained a higher-quality alternative that reached 98.8% success while still reducing cost by 51.5%.
By the numbers.
| Metric | Baseline | With Isoquant |
|---|---|---|
| Task success | 95.9% | 95.9% |
| Correct / total | 491 / 512 | 491 / 512 |
| Inference cost / 1,000 tasks | $4.34 | $1.55 |
| Mean completion latency | 6.60 s | 9.31 s |
More than one way to improve.
These are separate retained configurations, each with its own cost, quality and latency profile.
GPT-5.4 nano + Isoquant optimization, higher-quality alternative
- Task success
- 98.8% (506 / 512)
- Cost / 1,000 tasks
- $2.11
- Mean latency
- 12.16 s
GPT-4.1 mini + Isoquant optimization, retained alternative
- Task success
- 98.2% (503 / 512)
- Cost / 1,000 tasks
- $2.47
- Mean latency
- 10.70 s