Models / Anthropic
Claude Haiku 5.5
Anthropic · Proprietary
Decision brief
Start by investigating Reasoning and Agent workflows workloads. These are observed strengths in the dataset; validate them on your own tasks.
Coverage counts measured dimensions, not confidence. Benchmark conditions and public source links may be missing. Missing: Mathematics, Knowledge, Instructions, Multimodal, Multilingual, swe, terminal, gpqa, hle, mcp.
View evidence ↓Measured task dimensions
Normalized dimensions. A dash means unknown, not zero.
Input / output
$0.1 / $0.5 · USD / 1M tokens
Listed dataset prices; provider terms and dates need verification. Zero without explicit free-price evidence is treated as unknown.
Cache, batch, tool fees and self-hosting costs: not verified.
Index variability stability
98.7 · 4 days of history
short_band
This is index variability, not API uptime or reliability.
Before deployment
- Confirm provider endpoint, license and version.
- Test your prompts, tool calls and failure recovery.
- Verify prices and latency with real usage.
Evidence
These are source categories published by the dataset. Original public links, sample sizes and harness details remain incomplete.
| Source category | Provenance type | As of | Original public link |
|---|---|---|---|
| Coding | core / official_json | 2026-10-11 | Not published |
| Agent workflows | core / official_json | 2026-10-11 | Not published |
| depth | core+cross / official_json | 2026-10-11 | Not published |
Published index inputs · cf-1.3
The method includes overall, but this snapshot does not expose that input. The total cannot be fully reproduced from these displayed axes.
Method: median of available axes, with an 8% Ops blend when present. Source scores are displayed without recalculation or imputation.
Method →Index history
| As of | Cheese Index |
|---|---|
| 2026-10-08 | 62.1 |
| 2026-10-09 | 62.4 |
| 2026-10-10 | 62.6 |
| 2026-10-11 | 62.6 |
Only recorded dates are shown. Missing days are not filled.