Loading blueprint versions...
Please wait while we gather all the unique runs for this blueprint.
Please wait while we gather all the unique runs for this blueprint.
Please wait while we prepare the detailed comparison.
Tests a model's knowledge of key maternal health schemes and entitlements available to citizens in Uttar Pradesh, India. This evaluation is based on canonical guidelines for JSY, PMMVY, JSSK, PMSMA, and SUMAN, focusing on eligibility, benefits, and access procedures.
Average key point coverage extent for each model across all prompts.
Prompts vs. Models | Claude 3 5 Sonnet | Claude 3 7 Sonnet | Claude 3.5 Haiku | Claude Opus 4 | Claude Opus 4.1 | Claude Sonnet 4 | Command A | Deepseek Chat V3 | Deepseek R1 | Gemini 2.5 Flash | Gemini 2.5 Pro | Llama 3 70b Instruct | Llama 4 Maverick | Meta Llama 3.1 405b Instruct Turbo | Mistral Large 2411 | Mistral Medium 3 | GPT 4.1 | GPT 4.1 Mini | GPT 4.1 Nano | GPT 4o | GPT 4o Mini | GPT 5 | GPT Oss 120b | GPT Oss 20b | O4 Mini | GLM 4.5 | Grok 3 | Grok 4 | |
---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
Score | 8th 58.2% | 16th 53.6% | 24th 46.0% | 6th 60.2% | 11th 57.5% | 18th 53.2% | 16th 53.6% | 9th 57.8% | 3rd 62.5% | 10th 57.8% | 2nd 63.4% | 23rd 46.7% | 27th 37.7% | 20th 49.6% | 21st 49.6% | 12th 56.6% | 13th 56.3% | 15th 53.8% | 25th 44.5% | 19th 52.9% | 26th 40.6% | 4th 62.3% | 22nd 47.9% | 28th 36.8% | 7th 59.4% | 14th 55.8% | 5th 61.8% | 1st 67.3% | |
94.4% | 99% | 95% | 91% | 99% | 100% | 98% | 93% | 98% | 100% | 100% | 98% | 95% | 85% | 93% | 98% | 100% | 99% | 99% | 84% | 99% | 75% | 100% | 94% | 71% | 100% | 85% | 97% | 100% | |
55.5% | 66% | 48% | 44% | 56% | 60% | 73% | 64% | 62% | 93% | 60% | 80% | 40% | 24% | 39% | 33% | 64% | 57% | 50% | 40% | 57% | 43% | 60% | 44% | 42% | 56% | 60% | 73% | 70% | |
33.9% | 18% | 43% | 21% | 41% | 18% | 17% | 18% | 43% | 39% | 39% | 58% | 10% | 15% | 27% | 35% | 32% | 41% | 35% | 43% | 33% | 23% | 62% | 42% | 18% | 35% | 39% | 44% | 65% | |
35.9% | 38% | 13% | 29% | 40% | 40% | 46% | 49% | 38% | 33% | 41% | 37% | 44% | 39% | 44% | 42% | 40% | 36% | 36% | 19% | 38% | 16% | 33% | 35% | 14% | 53% | 43% | 40% | 35% | |
82.9% | 88% | 97% | 75% | 100% | 97% | 80% | 82% | 89% | 93% | 87% | 97% | 79% | 37% | 86% | 78% | 87% | 90% | 94% | 65% | 78% | 74% | 99% | 55% | 50% | 92% | 92% | 90% | 95% | |
19.5% | 41% | 26% | 17% | 25% | 31% | 6% | 16% | 19% | 18% | 21% | 12% | 12% | 28% | 10% | 13% | 18% | 16% | 10% | 17% | 12% | 13% | 20% | 19% | 27% | 21% | 16% | 28% | 40% |