👁️ Multimodal · llm-stats.com

CharXiv Reasoning

Reasoning over scientific charts and figures from arXiv papers.

12
테스트된 모델
Claude Mythos Preview
톱 모델
93.2%
최고 점수
Anthropic
제공업체
#모델명CharXiv Reasoning제공업체Podium Score
1
93.2%
Anthropic97.4
2
Kimi K3
Moonshot AI
91.3%
Moonshot AI83.3
3
91%
Anthropic52.0
4
89.9%
Anthropic76.0
5
Kimi K2.6
Moonshot AI
86.7%
Moonshot AI49.0
6
86.4%
Meta47.0
7
85.9%
Alibaba42.2
8
GPT-5.2
OpenAI
82.1%
OpenAI49.2
9
81.5%
Alibaba32.4
10
81.4%
Google51.9
11
80.3%
Google42.4
12
77.4%
Anthropic61.2
벤치마크 데이터는 llm-stats.com에서 집계. 자세한 내용은 방법론.