huggingface-course
/

albert-tokenizer-without-normalizer

SaulLu commited on Oct 19, 2021

Commit

50a007b

•

1 Parent(s): 67b371b

add modified albert tokenizer

Files changed (1) hide show

README.md CHANGED Viewed

@@ -11,6 +11,5 @@ print(tokenizer.convert_ids_to_tokens(tokenizer.encode(text)))
 # ['[CLS]', '▁this', '▁is', '▁a', '▁text', '▁with', '▁accent', 's', '▁and', '▁capital', '▁letters', '[SEP]']
 tokenizer = AutoTokenizer.from_pretrained("huggingface-course/albert-tokenizer-without-normalizer")
 print(tokenizer.convert_ids_to_tokens(tokenizer.encode(text)))
-#
-['[CLS]', '▁', '<unk>', 'his', '▁is', '▁a', '▁text', '▁with', '▁', '<unk>', 'cc', '<unk>', 'nts', '▁and', '▁', '<unk>', '▁', '<unk>', '[SEP]']
 ```

 # ['[CLS]', '▁this', '▁is', '▁a', '▁text', '▁with', '▁accent', 's', '▁and', '▁capital', '▁letters', '[SEP]']
 tokenizer = AutoTokenizer.from_pretrained("huggingface-course/albert-tokenizer-without-normalizer")
 print(tokenizer.convert_ids_to_tokens(tokenizer.encode(text)))
+# ['[CLS]', '▁', '<unk>', 'his', '▁is', '▁a', '▁text', '▁with', '▁', '<unk>', 'cc', '<unk>', 'nts', '▁and', '▁', '<unk>', '▁', '<unk>', '[SEP]']
 ```