Loading blueprint versions...
Please wait while we gather all the unique runs for this blueprint.
Please wait while we gather all the unique runs for this blueprint.
Please wait while we prepare the detailed comparison.
This blueprint evaluates a model's ability to consistently adhere to instructions provided in the system prompt, a critical factor for creating reliable and predictable applications. It tests various common failure modes observed in language models.
Core Areas Tested:
Average key point coverage extent for each model across all prompts.
Prompts vs. Models | Claude 3.5 Haiku | Claude Opus 4 | Claude Sonnet 4 | Command A | Deepseek Chat V3 | Deepseek R1 | Gemini 2.5 Flash | Gemini 2.5 Pro | Llama 3 70b Instruct | Llama 4 Maverick | Meta Llama 3.1 405b Instruct Turbo | Mistral Large 2411 | Mistral Medium 3 | GPT 4.1 | GPT 4.1 Mini | GPT 4.1 Nano | GPT 4o | GPT 4o Mini | O4 Mini | Kimi K2 Instruct | Grok 3 | Grok 3 Mini | Grok 4 | |
---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
Score | 6th 87.3% | 2nd 90.3% | 7th 86.6% | 20th 73.6% | 15th 77.5% | 16th 76.1% | 5th 88.1% | 11th 82.9% | 22nd 71.9% | 21st 72.2% | 17th 75.3% | 9th 83.1% | 18th 74.7% | 12th 81.5% | 19th 73.9% | 23rd 67.2% | 10th 82.9% | 14th 78.3% | 8th 84.1% | 13th 78.8% | 4th 88.9% | 1st 91.8% | 3rd 89.8% | |
87.5% | 90% | 90% | 88% | 80% | 90% | 100% | 90% | 85% | 78% | 78% | 80% | 88% | 88% | 90% | 85% | 98% | 78% | 88% | 93% | 100% | 88% | 88% | 80% | |
88.3% | 88% | 86% | 79% | 86% | 66% | 64% | 93% | 93% | 86% | 100% | 93% | 93% | 79% | 79% | 100% | 100% | 79% | 100% | 100% | 79% | 100% | 100% | ||
74.1% | 100% | 70% | 90% | 70% | 20% | 80% | 80% | 40% | 80% | 80% | 95% | 80% | 40% | 90% | 80% | 70% | 80% | 70% | 80% | 80% | 70% | 80% | 80% | |
84.1% | 81% | 92% | 98% | 92% | 81% | 33% | 92% | 83% | 92% | 94% | 94% | 92% | 96% | 100% | 60% | 44% | 92% | 98% | 83% | 56% | 96% | 94% | 92% | |
100.0% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | |
53.9% | 80% | 100% | 58% | 45% | 45% | 45% | 90% | 43% | 45% | 43% | 43% | 43% | 25% | 25% | 45% | 25% | 45% | 28% | 65% | 100% | 43% | 100% | 58% | |
81.7% | 83% | 100% | 100% | 33% | 96% | 100% | 100% | 96% | 17% | 17% | 33% | 94% | 94% | 96% | 63% | 88% | 90% | 96% | 96% | 90% | 100% | 98% | 98% | |
83.1% | 80% | 80% | 80% | 80% | 100% | 100% | 75% | 80% | 80% | 80% | 80% | 80% | 80% | 80% | 80% | 80% | 80% | 80% | 98% | 80% | 80% | 80% | 98% | |
72.3% | 100% | 100% | 100% | 90% | 80% | 50% | 100% | 100% | 20% | 20% | 38% | 83% | 80% | 83% | 60% | 20% | 100% | 60% | 60% | 20% | 100% | 100% | 100% | |
49.7% | 70% | 40% | 45% | 40% | 35% | 30% | 50% | 70% | 48% | 40% | 65% | 30% | 35% | 25% | 30% | 40% | 60% | 30% | 70% | 20% | 100% | 93% | 78% | |
91.3% | 100% | 100% | 83% | 100% | 100% | 100% | 100% | 100% | 88% | 100% | 80% | 100% | 85% | 83% | 90% | 83% | 100% | 85% | 80% | 80% | 83% | 80% | 100% | |
78.5% | 68% | 98% | 78% | 70% | 75% | 78% | 80% | 85% | 95% | 70% | 50% | 68% | 73% | 90% | 73% | 70% | 68% | 68% | 85% | 83% | 90% | 93% | 98% | |
93.2% | 100% | 100% | 100% | 93% | 98% | 95% | 100% | 100% | 80% | 80% | 88% | 100% | 100% | 98% | 90% | 38% | 93% | 93% | 98% | 100% | 100% | 100% | 100% | |
76.0% | 75% | 98% | 100% | 44% | 77% | 67% | 71% | 69% | 69% | 81% | 90% | 96% | 73% | 83% | 52% | 52% | 85% | 79% | 63% | 94% | 83% | 71% | 75% | |
97.0% | 94% | 100% | 100% | 81% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 73% | 100% | 100% | 100% | 94% | 100% | 90% | 100% | 100% | 100% | 100% |