HuggingFaceTB/SmolLM-360M continue pretraining on Malaysian context dataset
Continue pretraining on 50B tokens, dataset prepared at https://github.com/malaysia-ai/pretrain-text-dataset/tree/main/smollm
Wandb at https://wandb.ai/huseinzol05/finetune-HuggingFaceTB-SmolLM-360M/
- Downloads last month
- 22
Model tree for mesolitica/malaysian-HuggingFaceTB-SmolLM-360M
Base model
HuggingFaceTB/SmolLM-360M