Models / OpenAI
GPT-6.1 Sol
OpenAI · Proprietary
Decision brief
Start by investigating Multimodal and Knowledge workloads. These are observed strengths in the dataset; validate them on your own tasks.
Coverage counts measured dimensions, not confidence. Benchmark conditions and public source links may be missing. Missing: Mathematics, Instructions, Multilingual, swe, terminal, gpqa, hle, mcp.
View evidence ↓Measured task dimensions
Normalized dimensions. A dash means unknown, not zero.
Input / output
$2 / $10 · USD / 1M tokens
Listed dataset prices; provider terms and dates need verification. Zero without explicit free-price evidence is treated as unknown.
Cache, batch, tool fees and self-hosting costs: not verified.
Index variability stability
100.0 · 3 days of history
short_band
This is index variability, not API uptime or reliability.
Before deployment
- Confirm provider endpoint, license and version.
- Test your prompts, tool calls and failure recovery.
- Verify prices and latency with real usage.
Evidence
These are source categories published by the dataset. Original public links, sample sizes and harness details remain incomplete.
| Source category | Provenance type | As of | Original public link |
|---|---|---|---|
| Coding | core / official_json | 2026-10-03 | Not published |
| Agent workflows | core / official_json | 2026-10-03 | Not published |
| depth | core / official_json | 2026-10-03 | Not published |
Published index inputs · cf-1.3
The method includes overall, but this snapshot does not expose that input. The total cannot be fully reproduced from these displayed axes.
Method: median of available axes, with an 8% Ops blend when present. Source scores are displayed without recalculation or imputation.
Method →Index history
| As of | Cheese Index |
|---|---|
| 2026-10-01 | 67.8 |
| 2026-10-02 | 67.8 |
| 2026-10-03 | 67.8 |
Only recorded dates are shown. Missing days are not filled.