Principle benchmark report

Models ranked by held-out recall — how well the stated principle recovers matching statements the model never saw. False-positive rates keep that number honest: a vague accept-everything principle scores high recall and high FP together.

ModelSuitesHeld-out recallShown accFP (near)FP (far)OverallRun stdevMalformedOver limit
typesafe/jev-1.13211.0000.9810.0920.0000.9720.00300
mistralai/mistral-medium-3-5210.9440.9650.2700.0000.9020.03500
openai/gpt-6-astra210.9350.9970.0540.0000.9600.01900
openai/gpt-5.4-mini210.9350.9590.1940.0000.9150.02600
cohere/command-a210.9330.9520.1240.0000.9300.02800
anthropic/claude-haiku-4.5210.9310.9490.1240.0000.9290.01900
inception/mercury-2.5210.9310.9650.0790.0000.9440.02901
openai/gpt-5.6-sol210.9290.9810.0250.0080.9600.02000
moonshotai/kimi-k3200.9230.9830.0370.0000.9560.03200
z-ai/glm-5.3-flash210.9210.9940.0570.0000.9520.02900
google/gemini-3.5-flash-lite210.9130.9520.0290.0000.9460.02100
minimax/minimax-m3210.9110.9430.1020.0000.9250.05400
openai/gpt-5.6-terra210.9070.9710.0570.0000.9410.02500
amazon/nova-2-lite-v1210.9050.9210.2290.0000.8850.03300
z-ai/glm-5.3210.9030.9870.0570.0000.9440.02700
anthropic/claude-fable-5.1210.8970.9900.0030.0000.9560.01600
anthropic/claude-sonnet-5210.8890.9620.0540.0000.9330.02404
deepseek/deepseek-v4-pro210.8870.9590.0540.0000.9310.04200
meta-llama/llama-4-maverick210.8770.9370.1900.0000.8870.03500
anthropic/claude-opus-5210.8730.9650.0130.0000.9370.01801
stepfun/step-3.7-flash210.8730.9430.0480.0000.9230.04500
nvidia/nemotron-3-ultra-550b-a55b210.8650.9780.0190.0000.9360.03300
arcee-ai/trinity-large-thinking210.8650.9270.0730.0000.9100.05700
qwen/qwen3.8-max-0902210.8630.9620.0290.0000.9290.03000
openai/gpt-5.4-nano210.8610.9330.3050.0080.8510.05900
qwen/qwen3.8-flash210.8530.9460.0440.0080.9160.04900
tencent/hy4-preview210.8530.9710.0630.0000.9180.03400
deepseek/deepseek-v4.1-flash210.8510.9560.0410.0080.9180.05200
nvidia/nemotron-3-super-120b-a12b200.8380.9200.0630.0000.8990.04900
meta-llama/llama-4-scout210.7980.8950.1710.0000.8500.04200
nvidia/nemotron-3.5-lightning210.7600.8790.1240.0320.8400.10800
cognitivecomputations/dolphin-mistral-24b-venice-edition210.7420.7490.4290.0240.7250.08000
nvidia/nemotron-3-nano-30b-a3b210.7120.7620.0730.0000.8070.06800
google/gemini-3.1-pro-preview210.7060.8760.0060.0000.8500.03500
google/gemini-3.8-flash210.5870.7900.0030.0000.7820.04900

Held-out recall by shown tier

Modelshown=5
typesafe/jev-1.131.000
mistralai/mistral-medium-3-50.944
openai/gpt-6-astra0.935
openai/gpt-5.4-mini0.935
cohere/command-a0.933
anthropic/claude-haiku-4.50.931
inception/mercury-2.50.931
openai/gpt-5.6-sol0.929
moonshotai/kimi-k30.923
z-ai/glm-5.3-flash0.921
google/gemini-3.5-flash-lite0.913
minimax/minimax-m30.911
openai/gpt-5.6-terra0.907
amazon/nova-2-lite-v10.905
z-ai/glm-5.30.903
anthropic/claude-fable-5.10.897
anthropic/claude-sonnet-50.889
deepseek/deepseek-v4-pro0.887
meta-llama/llama-4-maverick0.877
anthropic/claude-opus-50.873
stepfun/step-3.7-flash0.873
nvidia/nemotron-3-ultra-550b-a55b0.865
arcee-ai/trinity-large-thinking0.865
qwen/qwen3.8-max-09020.863
openai/gpt-5.4-nano0.861
qwen/qwen3.8-flash0.853
tencent/hy4-preview0.853
deepseek/deepseek-v4.1-flash0.851
nvidia/nemotron-3-super-120b-a12b0.838
meta-llama/llama-4-scout0.798
nvidia/nemotron-3.5-lightning0.760
cognitivecomputations/dolphin-mistral-24b-venice-edition0.742
nvidia/nemotron-3-nano-30b-a3b0.712
google/gemini-3.1-pro-preview0.706
google/gemini-3.8-flash0.587

Comparison chart

Model comparison across benchmark metrics