Please wait while we gather all the unique runs for this blueprint.
Please wait while we gather all the unique runs for this blueprint.
Please wait while we prepare the detailed comparison.
Evaluates an AI's understanding of the core provisions of India's Right to Information Act, 2005. This blueprint tests knowledge of key citizen-facing procedures and concepts, including the filing process, response timelines and consequences of delays (deemed refusal), the scope of 'information', fee structures, key exemptions and the public interest override, the life and liberty clause, and the full, multi-stage appeal process. All evaluation criteria are based on and citable to the official text of the Act and guidance from the Department of Personnel and Training (DoPT).
Average key point coverage extent for each model across all prompts.
Prompts vs. Models | Claude 3.5 Haiku | Claude Sonnet 4 | Command A | Deepseek Chat V3 | Deepseek R1 | Gemini 2.5 Flash | Gemini 2.5 Pro Preview 05 06 | Mistral Large 2411 | Mistral Medium 3 | GPT 4.1 | GPT 4.1 Mini | GPT 4.1 Nano | GPT 4o | GPT 4o 2024 05 13 | GPT 4o 2024 08 06 | GPT 4o 2024 11 20 | GPT 4o Mini | O4 Mini | Grok 3 | Grok 3 Mini | Grok 4 | |
---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
Score | 19th 64.8% | 10th 77.8% | 17th 70.0% | 12th 75.3% | 8th 79.8% | 2nd 88.2% | 18th 66.8% | 13th 74.2% | 13th 74.2% | 4th 81.4% | 15th 71.1% | 21st 52.4% | 7th 80.0% | 11th 76.9% | 9th 78.7% | 5th 81.1% | 20th 58.5% | 16th 70.9% | 3rd 84.8% | 6th 80.9% | 1st 94.7% | |
47.3% | 13% | 32% | 38% | 25% | 47% | 100% | 50% | 94% | 53% | 53% | 38% | 0% | 53% | 25% | 59% | 63% | 13% | 19% | 63% | 56% | 100% | |
60.7% | 63% | 75% | 56% | 63% | 69% | 25% | 63% | 63% | 50% | 75% | 63% | 50% | 69% | 63% | 63% | 63% | 63% | 50% | 63% | 63% | 63% | |
86.1% | 80% | 90% | 98% | 88% | 88% | 93% | 75% | 85% | 85% | 83% | 90% | 70% | 88% | 85% | 85% | 83% | 65% | 98% | 90% | 95% | 95% | |
82.3% | 71% | 94% | 85% | 96% | 100% | 100% | 33% | 90% | 67% | 83% | 60% | 35% | 100% | 81% | 98% | 100% | 69% | 69% | 100% | 100% | 98% | |
64.1% | 75% | 53% | 56% | 72% | 100% | 75% | 50% | 50% | 34% | 78% | 75% | 31% | 72% | 75% | 63% | 75% | 53% | 50% | 75% | 59% | 75% | |
100.0% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | |
83.4% | 55% | 95% | 90% | 90% | 98% | 100% | 58% | 68% | 90% | 90% | 83% | 55% | 88% | 80% | 90% | 98% | 55% | 98% | 100% | 70% | 100% | |
60.4% | 75% | 63% | 54% | 42% | 71% | 54% | 100% | 83% | 46% | 46% | 38% | 50% | 54% | 63% | 54% | 46% | 38% | 75% | 42% | 75% | 100% | |
92.1% | 100% | 94% | 94% | 97% | 91% | 100% | 53% | 88% | 100% | 100% | 84% | 78% | 91% | 97% | 94% | 94% | 91% | 88% | 100% | 100% | 100% | |
83.9% | 63% | 100% | 67% | 100% | 67% | 100% | 33% | 63% | 67% | 100% | 67% | 100% | 100% | 100% | 67% | 100% | 67% | 100% | 100% | 100% | 100% | |
77.3% | 34% | 78% | 84% | 81% | 100% | 100% | 53% | 100% | 72% | 75% | 88% | 34% | 75% | 75% | 100% | 100% | 47% | 75% | 100% | 53% | 100% | |
96.5% | 100% | 100% | 88% | 100% | 100% | 100% | 100% | 67% | 100% | 100% | 100% | 71% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | |
45.4% | 13% | 38% | 0% | 25% | 7% | 100% | 100% | 13% | 100% | 75% | 38% | 7% | 50% | 56% | 50% | 32% | 0% | 0% | 69% | 81% | 100% |