Longin-Yu commited on
Commit
1754bfa
·
verified ·
1 Parent(s): b613f5e

Update technical report link and citation

Browse files
Files changed (2) hide show
  1. README.md +12 -1
  2. README.zh-CN.md +12 -1
README.md CHANGED
@@ -13,7 +13,7 @@ language:
13
  [![Project Page](https://img.shields.io/badge/-Project%20Page-4B5563?style=flat&logo=googlechrome&logoColor=white&labelColor=6B7280)](https://tencent.github.io/WeVisDoc)
14
  [![WeVisDoc-4B](https://img.shields.io/badge/-WeVisDoc--4B-E6A700?style=flat&logo=huggingface&logoColor=FFD21E&labelColor=6B7280)](https://huggingface.co/Tencent/WeVisDoc-4B)
15
  [![WeVisDoc-2B](https://img.shields.io/badge/-WeVisDoc--2B-E6A700?style=flat&logo=huggingface&logoColor=FFD21E&labelColor=6B7280)](https://huggingface.co/Tencent/WeVisDoc-2B)
16
- [![Technical Report](https://img.shields.io/badge/-Technical%20Report-C94C4C?style=flat&logo=arxiv&logoColor=B31B1B&labelColor=6B7280)](# "Coming soon")
17
 
18
  WeVisDoc is an end-to-end document parser for page images. Fine-tuned from [Qwen3-VL-2B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-2B-Instruct) and [Qwen3-VL-4B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-4B-Instruct), it turns a page into structured Markdown, with LaTeX formulas and HTML tables.
19
 
@@ -175,3 +175,14 @@ python -m wevisdoc.local --model Tencent/WeVisDoc-2B \
175
  ```
176
 
177
  Use `Tencent/WeVisDoc-4B` for the 4B version. `--model` can be omitted when `WEVISDOC_MODEL_PATH` is set. Local inference supports `--device-map` (default `auto`) and `--max-tokens` (8192). Render PDFs to page images first.
 
 
 
 
 
 
 
 
 
 
 
 
13
  [![Project Page](https://img.shields.io/badge/-Project%20Page-4B5563?style=flat&logo=googlechrome&logoColor=white&labelColor=6B7280)](https://tencent.github.io/WeVisDoc)
14
  [![WeVisDoc-4B](https://img.shields.io/badge/-WeVisDoc--4B-E6A700?style=flat&logo=huggingface&logoColor=FFD21E&labelColor=6B7280)](https://huggingface.co/Tencent/WeVisDoc-4B)
15
  [![WeVisDoc-2B](https://img.shields.io/badge/-WeVisDoc--2B-E6A700?style=flat&logo=huggingface&logoColor=FFD21E&labelColor=6B7280)](https://huggingface.co/Tencent/WeVisDoc-2B)
16
+ [![Technical Report](https://img.shields.io/badge/-Technical%20Report-C94C4C?style=flat&logo=arxiv&logoColor=B31B1B&labelColor=6B7280)](https://arxiv.org/abs/2609.20423)
17
 
18
  WeVisDoc is an end-to-end document parser for page images. Fine-tuned from [Qwen3-VL-2B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-2B-Instruct) and [Qwen3-VL-4B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-4B-Instruct), it turns a page into structured Markdown, with LaTeX formulas and HTML tables.
19
 
 
175
  ```
176
 
177
  Use `Tencent/WeVisDoc-4B` for the 4B version. `--model` can be omitted when `WEVISDOC_MODEL_PATH` is set. Local inference supports `--device-map` (default `auto`) and `--max-tokens` (8192). Render PDFs to page images first.
178
+
179
+ ## Citation
180
+
181
+ ```bibtex
182
+ @article{wevisdoc,
183
+ title = {WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing},
184
+ author = {Hao Yu and Kang Liu and Linnan Zhao and Jiabo Zhan and Chong Sun and Chen Li and Jing Lyu},
185
+ year = {2026},
186
+ journal = {arXiv preprint arXiv: 2609.20423}
187
+ }
188
+ ```
README.zh-CN.md CHANGED
@@ -6,7 +6,7 @@
6
  [![Project Page](https://img.shields.io/badge/-Project%20Page-4B5563?style=flat&logo=googlechrome&logoColor=white&labelColor=6B7280)](https://tencent.github.io/WeVisDoc)
7
  [![WeVisDoc-4B](https://img.shields.io/badge/-WeVisDoc--4B-E6A700?style=flat&logo=huggingface&logoColor=FFD21E&labelColor=6B7280)](https://huggingface.co/Tencent/WeVisDoc-4B)
8
  [![WeVisDoc-2B](https://img.shields.io/badge/-WeVisDoc--2B-E6A700?style=flat&logo=huggingface&logoColor=FFD21E&labelColor=6B7280)](https://huggingface.co/Tencent/WeVisDoc-2B)
9
- [![Technical Report](https://img.shields.io/badge/-Technical%20Report-C94C4C?style=flat&logo=arxiv&logoColor=B31B1B&labelColor=6B7280)](# "Coming soon")
10
 
11
  WeVisDoc 是面向文档图片的端到端解析模型,由 [Qwen3-VL-2B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-2B-Instruct) 与 [Qwen3-VL-4B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-4B-Instruct) 微调得到,将页面转为结构化 Markdown,并输出 LaTeX 公式与 HTML 表格。
12
 
@@ -168,3 +168,14 @@ python -m wevisdoc.local --model Tencent/WeVisDoc-2B \
168
  ```
169
 
170
  如需使用 4B 版本,将模型 ID 替换为 `Tencent/WeVisDoc-4B`。设置 `WEVISDOC_MODEL_PATH` 后可省略 `--model`。本地推理支持 `--device-map`(默认 `auto`)和 `--max-tokens`(默认 8192)。PDF 需先转成页面图片。
 
 
 
 
 
 
 
 
 
 
 
 
6
  [![Project Page](https://img.shields.io/badge/-Project%20Page-4B5563?style=flat&logo=googlechrome&logoColor=white&labelColor=6B7280)](https://tencent.github.io/WeVisDoc)
7
  [![WeVisDoc-4B](https://img.shields.io/badge/-WeVisDoc--4B-E6A700?style=flat&logo=huggingface&logoColor=FFD21E&labelColor=6B7280)](https://huggingface.co/Tencent/WeVisDoc-4B)
8
  [![WeVisDoc-2B](https://img.shields.io/badge/-WeVisDoc--2B-E6A700?style=flat&logo=huggingface&logoColor=FFD21E&labelColor=6B7280)](https://huggingface.co/Tencent/WeVisDoc-2B)
9
+ [![Technical Report](https://img.shields.io/badge/-Technical%20Report-C94C4C?style=flat&logo=arxiv&logoColor=B31B1B&labelColor=6B7280)](https://arxiv.org/abs/2609.20423)
10
 
11
  WeVisDoc 是面向文档图片的端到端解析模型,由 [Qwen3-VL-2B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-2B-Instruct) 与 [Qwen3-VL-4B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-4B-Instruct) 微调得到,将页面转为结构化 Markdown,并输出 LaTeX 公式与 HTML 表格。
12
 
 
168
  ```
169
 
170
  如需使用 4B 版本,将模型 ID 替换为 `Tencent/WeVisDoc-4B`。设置 `WEVISDOC_MODEL_PATH` 后可省略 `--model`。本地推理支持 `--device-map`(默认 `auto`)和 `--max-tokens`(默认 8192)。PDF 需先转成页面图片。
171
+
172
+ ## 引用
173
+
174
+ ```bibtex
175
+ @article{wevisdoc,
176
+ title = {WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing},
177
+ author = {Hao Yu and Kang Liu and Linnan Zhao and Jiabo Zhan and Chong Sun and Chen Li and Jing Lyu},
178
+ year = {2026},
179
+ journal = {arXiv preprint arXiv: 2609.20423}
180
+ }
181
+ ```