leaderboard-pr-bot commited on
Commit
428c77d
1 Parent(s): f6e0343

Adding Evaluation Results

Browse files

This is an automated PR created with https://huggingface.co/spaces/Weyaxi/open-llm-leaderboard-results-pr

The purpose of this PR is to add evaluation results from the Open LLM Leaderboard to your model card.

If you encounter any issues, please report them to https://huggingface.co/spaces/Weyaxi/open-llm-leaderboard-results-pr/discussions

Files changed (1) hide show
  1. README.md +120 -3
README.md CHANGED
@@ -1,19 +1,122 @@
1
  ---
2
- base_model: LumiOpen/Poro-34B
3
  datasets:
4
  - cerebras/SlimPajama-627B
5
  - bigcode/starcoderdata
6
  - mc4
7
  - allenai/dolma
 
 
8
  inference: false
9
- license: apache-2.0
10
  model_creator: LumiOpen
11
- model_name: Poro 34B
12
  model_type: bloom
13
  prompt_template: '{prompt}
14
 
15
  '
16
  quantized_by: TheBloke
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
17
  ---
18
  <!-- markdownlint-disable MD041 -->
19
 
@@ -478,3 +581,17 @@ Poro is an advanced language model, primarily optimized for English, Finnish and
478
  ## License
479
 
480
  Poro is released under the Apache 2.0 license.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ license: apache-2.0
3
  datasets:
4
  - cerebras/SlimPajama-627B
5
  - bigcode/starcoderdata
6
  - mc4
7
  - allenai/dolma
8
+ model_name: Poro 34B
9
+ base_model: LumiOpen/Poro-34B
10
  inference: false
 
11
  model_creator: LumiOpen
 
12
  model_type: bloom
13
  prompt_template: '{prompt}
14
 
15
  '
16
  quantized_by: TheBloke
17
+ model-index:
18
+ - name: Poro-34B-GPTQ
19
+ results:
20
+ - task:
21
+ type: text-generation
22
+ name: Text Generation
23
+ dataset:
24
+ name: AI2 Reasoning Challenge (25-Shot)
25
+ type: ai2_arc
26
+ config: ARC-Challenge
27
+ split: test
28
+ args:
29
+ num_few_shot: 25
30
+ metrics:
31
+ - type: acc_norm
32
+ value: 47.01
33
+ name: normalized accuracy
34
+ source:
35
+ url: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=TheBloke/Poro-34B-GPTQ
36
+ name: Open LLM Leaderboard
37
+ - task:
38
+ type: text-generation
39
+ name: Text Generation
40
+ dataset:
41
+ name: HellaSwag (10-Shot)
42
+ type: hellaswag
43
+ split: validation
44
+ args:
45
+ num_few_shot: 10
46
+ metrics:
47
+ - type: acc_norm
48
+ value: 73.75
49
+ name: normalized accuracy
50
+ source:
51
+ url: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=TheBloke/Poro-34B-GPTQ
52
+ name: Open LLM Leaderboard
53
+ - task:
54
+ type: text-generation
55
+ name: Text Generation
56
+ dataset:
57
+ name: MMLU (5-Shot)
58
+ type: cais/mmlu
59
+ config: all
60
+ split: test
61
+ args:
62
+ num_few_shot: 5
63
+ metrics:
64
+ - type: acc
65
+ value: 32.47
66
+ name: accuracy
67
+ source:
68
+ url: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=TheBloke/Poro-34B-GPTQ
69
+ name: Open LLM Leaderboard
70
+ - task:
71
+ type: text-generation
72
+ name: Text Generation
73
+ dataset:
74
+ name: TruthfulQA (0-shot)
75
+ type: truthful_qa
76
+ config: multiple_choice
77
+ split: validation
78
+ args:
79
+ num_few_shot: 0
80
+ metrics:
81
+ - type: mc2
82
+ value: 38.37
83
+ source:
84
+ url: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=TheBloke/Poro-34B-GPTQ
85
+ name: Open LLM Leaderboard
86
+ - task:
87
+ type: text-generation
88
+ name: Text Generation
89
+ dataset:
90
+ name: Winogrande (5-shot)
91
+ type: winogrande
92
+ config: winogrande_xl
93
+ split: validation
94
+ args:
95
+ num_few_shot: 5
96
+ metrics:
97
+ - type: acc
98
+ value: 71.35
99
+ name: accuracy
100
+ source:
101
+ url: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=TheBloke/Poro-34B-GPTQ
102
+ name: Open LLM Leaderboard
103
+ - task:
104
+ type: text-generation
105
+ name: Text Generation
106
+ dataset:
107
+ name: GSM8k (5-shot)
108
+ type: gsm8k
109
+ config: main
110
+ split: test
111
+ args:
112
+ num_few_shot: 5
113
+ metrics:
114
+ - type: acc
115
+ value: 5.08
116
+ name: accuracy
117
+ source:
118
+ url: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard?query=TheBloke/Poro-34B-GPTQ
119
+ name: Open LLM Leaderboard
120
  ---
121
  <!-- markdownlint-disable MD041 -->
122
 
 
581
  ## License
582
 
583
  Poro is released under the Apache 2.0 license.
584
+
585
+ # [Open LLM Leaderboard Evaluation Results](https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard)
586
+ Detailed results can be found [here](https://huggingface.co/datasets/open-llm-leaderboard/details_TheBloke__Poro-34B-GPTQ)
587
+
588
+ | Metric |Value|
589
+ |---------------------------------|----:|
590
+ |Avg. |44.67|
591
+ |AI2 Reasoning Challenge (25-Shot)|47.01|
592
+ |HellaSwag (10-Shot) |73.75|
593
+ |MMLU (5-Shot) |32.47|
594
+ |TruthfulQA (0-shot) |38.37|
595
+ |Winogrande (5-shot) |71.35|
596
+ |GSM8k (5-shot) | 5.08|
597
+