Please wait while we gather all the unique runs for this blueprint.
Please wait while we gather all the unique runs for this blueprint.
Please wait while we prepare the detailed comparison.
Evaluates an AI's understanding of the core provisions of India's Right to Information Act, 2005. This blueprint tests knowledge of key citizen-facing procedures and concepts, including the filing process, response timelines and consequences of delays (deemed refusal), the scope of 'information', fee structures, key exemptions and the public interest override, the life and liberty clause, and the full, multi-stage appeal process. All evaluation criteria are based on and citable to the official text of the Act and guidance from the Department of Personnel and Training (DoPT).
Average key point coverage extent for each model across all prompts.
Prompts vs. Models | Claude 3.5 Haiku | Claude Sonnet 4 | Command A | Deepseek Chat V3 | Deepseek R1 | Gemini 2.5 Flash | Gemini 2.5 Pro | Mistral Large 2411 | Mistral Medium 3 | GPT 4.1 | GPT 4.1 Mini | GPT 4.1 Nano | GPT 4o | GPT 4o 2024 05 13 | GPT 4o 2024 08 06 | GPT 4o 2024 11 20 | GPT 4o Mini | O4 Mini | Kimi K2 Instruct | Grok 3 | Grok 3 Mini | Grok 4 | |
---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
Score | 18th 72.4% | 11th 78.8% | 9th 80.2% | 16th 74.0% | 14th 76.7% | 2nd 87.9% | 3rd 87.5% | 13th 77.2% | 8th 80.9% | 10th 80.1% | 17th 73.7% | 22nd 52.7% | 7th 81.8% | 15th 76.0% | 12th 77.9% | 6th 81.8% | 21st 59.5% | 20th 70.2% | 19th 70.9% | 5th 83.3% | 4th 86.4% | 1st 92.8% | |
51.2% | 19% | 50% | 91% | 31% | 10% | 100% | 56% | 100% | 56% | 13% | 78% | 0% | 59% | 25% | 44% | 56% | 16% | 91% | 0% | 56% | 94% | 81% | |
61.9% | 63% | 63% | 56% | 75% | 69% | 25% | 69% | 63% | 75% | 75% | 50% | 50% | 63% | 63% | 63% | 63% | 75% | 63% | 50% | 63% | 63% | 63% | |
86.1% | 85% | 90% | 100% | 90% | 95% | 93% | 90% | 70% | 90% | 90% | 83% | 63% | 85% | 90% | 85% | 85% | 65% | 95% | 75% | 95% | 85% | 95% | |
88.0% | 100% | 88% | 85% | 94% | 92% | 100% | 100% | 88% | 96% | 90% | 60% | 50% | 96% | 81% | 98% | 94% | 71% | 60% | 96% | 100% | 100% | 98% | |
66.9% | 75% | 78% | 56% | 75% | 75% | 75% | 53% | 50% | 59% | 75% | 75% | 28% | 75% | 75% | 72% | 75% | 44% | 31% | 100% | 75% | 75% | 75% | |
100.0% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | |
87.1% | 73% | 60% | 100% | 75% | 95% | 100% | 98% | 98% | 80% | 93% | 75% | 60% | 93% | 85% | 90% | 100% | 60% | 95% | 90% | 98% | 100% | 98% | |
54.3% | 46% | 79% | 54% | 54% | 46% | 50% | 75% | 71% | 54% | 42% | 38% | 42% | 54% | 50% | 54% | 46% | 46% | 38% | 42% | 46% | 67% | 100% | |
95.4% | 100% | 84% | 94% | 94% | 97% | 100% | 97% | 94% | 100% | 100% | 94% | 94% | 94% | 94% | 84% | 100% | 91% | 97% | 100% | 94% | 100% | 97% | |
84.8% | 67% | 100% | 100% | 67% | 67% | 100% | 100% | 63% | 67% | 100% | 67% | 67% | 100% | 100% | 67% | 100% | 67% | 67% | 100% | 100% | 100% | 100% | |
84.0% | 75% | 100% | 100% | 88% | 100% | 100% | 100% | 100% | 75% | 75% | 75% | 53% | 75% | 75% | 100% | 100% | 38% | 75% | 69% | 100% | 75% | 100% | |
97.9% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 71% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 83% | 100% | |
48.5% | 38% | 32% | 7% | 19% | 51% | 100% | 100% | 7% | 100% | 88% | 63% | 7% | 69% | 50% | 56% | 44% | 0% | 0% | 0% | 56% | 81% | 100% |