initial release

Files changed (7) hide show

README.md ADDED Viewed

+---
+language:
+- "th"
+tags:
+- "thai"
+- "token-classification"
+- "pos"
+- "dependency-parsing"
+base_model: KoichiYasuoka/camembert-thai-base
+datasets:
+- "universal_dependencies"
+license: "apache-2.0"
+pipeline_tag: "token-classification"
+widget:
+- text: "หลายหัวดีกว่าหัวเดียว"
+---
+# camembert-thai-base-upos
+## Model Description
+This is a CamemBERT model pre-trained on Thai texts for POS-tagging and dependency-parsing, derived from [camembert-thai-base](https://huggingface.co/KoichiYasuoka/camembert-thati-base). Every word is tagged by [UPOS](https://universaldependencies.org/u/pos/) (Universal Part-Of-Speech).
+## How to Use
+```py
+from transformers import pipeline
+nlp=pipeline("token-classification","KoichiYasuoka/camembert-thai-base-upos",aggregation_strategy="simple")
+print(nlp("หลายหัวดีกว่าหัวเดียว"))
+```
+or
+```
+import esupar
+nlp=esupar.load("KoichiYasuoka/camembert-thai-base-upos")
+print(nlp("หลายหัวดีกว่าหัวเดียว"))
+```
+## See Also
+[esupar](https://github.com/KoichiYasuoka/esupar): Tokenizer POS-tagger and Dependency-parser with BERT/RoBERTa/DeBERTa models

config.json ADDED Viewed

The diff for this file is too large to render. See raw diff

pytorch_model.bin ADDED Viewed

+version https://git-lfs.github.com/spec/v1
+oid sha256:176fecfd3434ae3519103a2cc1bfc557694ee69a111fe6e997ec4ab888d1a7e9
+size 1109330278

special_tokens_map.json ADDED Viewed

+{
+  "additional_special_tokens": [
+    "<s>NOTUSED",
+    "</s>NOTUSED",
+    "<_>"
+  ],
+  "bos_token": {
+    "content": "<s>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "cls_token": {
+    "content": "<s>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "eos_token": {
+    "content": "</s>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "mask_token": {
+    "content": "<mask>",
+    "lstrip": true,
+    "normalized": true,
+    "rstrip": false,
+    "single_word": false
+  },
+  "pad_token": {
+    "content": "<pad>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "sep_token": {
+    "content": "</s>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "unk_token": {
+    "content": "<unk>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  }
+}

supar.model ADDED Viewed

+version https://git-lfs.github.com/spec/v1
+oid sha256:00ce85aba1a732475a667f0602a827c0d416b1728ec65eaf9704cd1e61259e73
+size 1163935350

tokenizer.json ADDED Viewed

+version https://git-lfs.github.com/spec/v1
+oid sha256:67d187215f962d5cce64e220651641808a1afeb332d7c6e22447bfe9b2aa9138
+size 16916048

tokenizer_config.json ADDED Viewed

The diff for this file is too large to render. See raw diff