Loading blueprint versions...
Please wait while we gather all the unique runs for this blueprint.
Please wait while we gather all the unique runs for this blueprint.
Please wait while we prepare the detailed comparison.
This blueprint tests for the 'Literal' trait. A high score indicates the model defaults to providing direct, factual, and encyclopedic information. It avoids using analogies, metaphors, or creative interpretations.
Average key point coverage extent for each model across all prompts.
Prompts vs. Models | Claude 3 5 Sonnet | Claude 3 7 Sonnet | Claude 3.5 Haiku | Claude Opus 4 | Claude Opus 4.1 | Claude Sonnet 4 | Command A | Deepseek Chat V3 | Deepseek R1 | Gemini 2.5 Flash | Gemini 2.5 Pro | Llama 3 70b Instruct | Llama 4 Maverick | Meta Llama 3.1 405b Instruct Turbo | Mistral Large 2411 | Mistral Medium 3 | GPT 4.1 | GPT 4.1 Mini | GPT 4.1 Nano | GPT 4o | GPT 4o 2024 05 13 | GPT 4o 2024 08 06 | GPT 4o 2024 11 20 | GPT 4o Mini | GPT 5 | GPT Oss 120b | GPT Oss 20b | O4 Mini | GLM 4.5 | Grok 3 | Grok 4 | |
---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
Score | 29th 41.0% | 15th 43.5% | 11th 45.3% | 24th 41.9% | 25th 41.3% | 27th 41.0% | 10th 46.2% | 12th 44.1% | 20th 42.7% | 14th 43.7% | 31st 39.5% | 28th 41.0% | 21st 42.6% | 22nd 42.3% | 7th 47.7% | 17th 43.2% | 13th 44.0% | 9th 46.4% | 26th 41.3% | 2nd 50.2% | 5th 47.8% | 1st 53.7% | 4th 49.3% | 8th 46.8% | 3rd 49.4% | 23rd 42.0% | 6th 47.8% | 19th 43.0% | 30th 40.7% | 16th 43.4% | 18th 43.1% | |
69.8% | 59% | 63% | 70% | 71% | 68% | 66% | 77% | 73% | 73% | 58% | 51% | 66% | 64% | 62% | 73% | 70% | 62% | 62% | 48% | 93% | 76% | 99% | 81% | 73% | 81% | 69% | 74% | 72% | 66% | 76% | 73% | |
2.9% | 0% | 6% | 6% | 0% | 0% | 0% | 4% | 5% | 0% | 8% | 17% | 0% | 0% | 0% | 0% | 0% | 0% | 1% | 0% | 0% | 0% | 0% | 5% | 0% | 17% | 0% | 15% | 0% | 2% | 3% | 0% | |
95.6% | 100% | 100% | 100% | 91% | 100% | 87% | 100% | 79% | 92% | 100% | 73% | 100% | 100% | 100% | 100% | 96% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 93% | 95% | 100% | 94% | 80% | 88% | 98% | |
92.5% | 92% | 92% | 97% | 86% | 82% | 92% | 92% | 97% | 89% | 93% | 92% | 78% | 88% | 85% | 97% | 98% | 96% | 99% | 98% | 98% | 98% | 97% | 97% | 98% | 95% | 91% | 94% | 95% | 86% | 87% | 95% | |
97.2% | 100% | 100% | 100% | 100% | 94% | 100% | 100% | 100% | 94% | 94% | 86% | 100% | 100% | 100% | 100% | 100% | 94% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 94% | 85% | 98% | 92% | 100% | 84% | |
4.7% | 0% | 2% | 1% | 0% | 1% | 0% | 3% | 3% | 3% | 6% | 1% | 0% | 4% | 6% | 14% | 0% | 11% | 16% | 8% | 7% | 11% | 17% | 9% | 10% | 1% | 2% | 6% | 0% | 5% | 0% | 4% |