Upload README.md
Browse files
README.md
CHANGED
@@ -569,15 +569,11 @@ All models are evaluated in chat mode (e.g. with the respective conversation tem
|
|
569 |
|
570 |
*: Grok results are reported by [X.AI](https://x.ai/).
|
571 |
|
572 |
-
<div>
|
573 |
-
<
|
574 |
-
5-shot:
|
575 |
</div>
|
576 |
|
577 |
-
|
578 |
-
|----------|-------|------------|----------------|-------|---------------|-------|
|
579 |
-
| ChatGPT | 47.81 | 55.68 | 56.5 | 62.66 | 50.69 | 55.51 |
|
580 |
-
| OpenChat | 38.7 | 45.99 | 48.32 | 50.23 | 43.27 | 45.85 |
|
581 |
|
582 |
<div>
|
583 |
<h3>Multi-Level Multi-Discipline Chinese Evaluation Suite (CEVAL)</h3>
|
@@ -588,6 +584,14 @@ All models are evaluated in chat mode (e.g. with the respective conversation tem
|
|
588 |
| ChatGPT | 54.4 | 52.9 | 61.8 | 50.9 | 53.6 |
|
589 |
| OpenChat | 47.29 | 45.22 | 52.49 | 48.52 | 45.08 |
|
590 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
591 |
|
592 |
<div align="center">
|
593 |
<h2> Limitations </h2>
|
|
|
569 |
|
570 |
*: Grok results are reported by [X.AI](https://x.ai/).
|
571 |
|
572 |
+
<div align="center">
|
573 |
+
<h2> 中文评估结果 / Chinese Evaluations </h2>
|
|
|
574 |
</div>
|
575 |
|
576 |
+
⚠️ Note that this model was not explicitly trained in Chinese (only < 0.1% of the data is in Chinese). 请注意本模型没有针对性训练中文(中文数据占比小于0.1%)。
|
|
|
|
|
|
|
577 |
|
578 |
<div>
|
579 |
<h3>Multi-Level Multi-Discipline Chinese Evaluation Suite (CEVAL)</h3>
|
|
|
584 |
| ChatGPT | 54.4 | 52.9 | 61.8 | 50.9 | 53.6 |
|
585 |
| OpenChat | 47.29 | 45.22 | 52.49 | 48.52 | 45.08 |
|
586 |
|
587 |
+
<div>
|
588 |
+
<h3>Massive Multitask Language Understanding in Chinese (CMMLU, 5-shot)</h3>
|
589 |
+
</div>
|
590 |
+
|
591 |
+
| Models | STEM | Humanities | SocialSciences | Other | ChinaSpecific | Avg |
|
592 |
+
|----------|-------|------------|----------------|-------|---------------|-------|
|
593 |
+
| ChatGPT | 47.81 | 55.68 | 56.5 | 62.66 | 50.69 | 55.51 |
|
594 |
+
| OpenChat | 38.7 | 45.99 | 48.32 | 50.23 | 43.27 | 45.85 |
|
595 |
|
596 |
<div align="center">
|
597 |
<h2> Limitations </h2>
|