procedure2012 commited on
Commit
1ff988e
·
verified ·
1 Parent(s): 10acb35

Publish Helios-Reasoner-7B (step_900) with eval results

Browse files
Files changed (6) hide show
  1. README.md +72 -0
  2. config.json +6 -0
  3. figures/fig1.png +0 -0
  4. figures/fig2.png +0 -0
  5. figures/fig3.png +0 -0
  6. pytorch_model.bin +3 -0
README.md ADDED
@@ -0,0 +1,72 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ library_name: transformers
4
+ ---
5
+ # Helios-Reasoner-7B
6
+ <!-- markdownlint-disable first-line-h1 -->
7
+ <!-- markdownlint-disable html -->
8
+ <!-- markdownlint-disable no-duplicate-header -->
9
+
10
+ <div align="center">
11
+ <img src="figures/fig1.png" width="60%" alt="Helios-Reasoner-7B" />
12
+ </div>
13
+ <hr>
14
+
15
+ <div align="center" style="line-height: 1;">
16
+ <a href="LICENSE" style="margin: 2px;">
17
+ <img alt="License" src="figures/fig2.png" style="display: inline-block; vertical-align: middle;"/>
18
+ </a>
19
+ </div>
20
+
21
+ ## 1. Introduction
22
+
23
+ Helios-Reasoner-7B is our flagship open reasoning model. Through extended RLHF and a curated math/code post-training mix, it now closes much of the gap with frontier closed models on hard multi-step reasoning.
24
+
25
+ <p align="center">
26
+ <img width="80%" src="figures/fig3.png">
27
+ </p>
28
+
29
+ ## 2. Evaluation Results
30
+
31
+ ### Comprehensive Benchmark Results
32
+
33
+ <div align="center">
34
+
35
+ | | Benchmark | Falcon-S | Mistral-Lite | Helios-v1 | Helios-Reasoner-7B |
36
+ |---|---|---|---|---|---|
37
+ | **Core Reasoning Tasks** | Math Reasoning | 0.533 | 0.515 | 0.519 | 0.537 |
38
+ | | Logical Reasoning | 0.808 | 0.781 | 0.830 | 0.801 |
39
+ | | Common Sense | 0.699 | 0.723 | 0.714 | 0.727 |
40
+ | **Language Understanding** | Reading Comprehension | 0.684 | 0.679 | 0.669 | 0.689 |
41
+ | | Question Answering | 0.606 | 0.571 | 0.584 | 0.600 |
42
+ | | Text Classification | 0.791 | 0.797 | 0.818 | 0.820 |
43
+ | | Sentiment Analysis | 0.767 | 0.774 | 0.755 | 0.786 |
44
+ | **Generation Tasks** | Code Generation | 0.641 | 0.651 | 0.668 | 0.636 |
45
+ | | Creative Writing | 0.593 | 0.624 | 0.603 | 0.595 |
46
+ | | Dialogue Generation | 0.617 | 0.613 | 0.632 | 0.634 |
47
+ | | Summarization | 0.728 | 0.736 | 0.739 | 0.759 |
48
+ | **Specialized Capabilities** | Translation | 0.780 | 0.764 | 0.783 | 0.800 |
49
+ | | Knowledge Retrieval | 0.637 | 0.677 | 0.661 | 0.670 |
50
+ | | Instruction Following | 0.755 | 0.751 | 0.719 | 0.750 |
51
+ | | Safety Evaluation | 0.693 | 0.702 | 0.709 | 0.732 |
52
+
53
+ </div>
54
+
55
+ ### Overall Performance Summary
56
+ The Helios-Reasoner-7B demonstrates strong performance across all evaluated benchmark categories, with particularly notable results in reasoning and generation tasks.
57
+
58
+ ## 3. Chat Website & API Platform
59
+ We offer a chat interface and API for you to interact with Helios-Reasoner-7B. Please check our official website for more details.
60
+
61
+ ## 4. How to Run Locally
62
+
63
+ Please refer to our code repository for more information about running Helios-Reasoner-7B locally.
64
+
65
+ ### Temperature
66
+ We recommend setting the temperature parameter to 0.6.
67
+
68
+ ## 5. License
69
+ This repository is released under the apache-2.0 license. The model supports commercial use.
70
+
71
+ ## 6. Contact
72
+ If you have any questions, please contact us at research@helios-labs.ai.
config.json ADDED
@@ -0,0 +1,6 @@
 
 
 
 
 
 
 
1
+ {
2
+ "model_type": "bert",
3
+ "architectures": [
4
+ "BertModel"
5
+ ]
6
+ }
figures/fig1.png ADDED
figures/fig2.png ADDED
figures/fig3.png ADDED
pytorch_model.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:965362299a238de576a92dfdd3e32aea7a2bacc94b2c41541c8c9258b923f587
3
+ size 23