Add Sentence Transformers integration (#2)

- Add Sentence Transformers integration with CLS pooling (fc6047009e0ee0f2ad4a36e4bae86bb10ba961fd)
- Add tags/library_name for tighter integration with HF (249dd29816720d49dc21809339445932237b866b)

Files changed (5) hide show

1_Pooling/config.json ADDED Viewed

+{
+  "word_embedding_dimension": 768,
+  "pooling_mode_cls_token": true,
+  "pooling_mode_mean_tokens": false,
+  "pooling_mode_max_tokens": false,
+  "pooling_mode_mean_sqrt_len_tokens": false,
+  "pooling_mode_weightedmean_tokens": false,
+  "pooling_mode_lasttoken": false,
+  "include_prompt": true
+}

README.md CHANGED Viewed

@@ -1,6 +1,9 @@
 ---
 tags:
 - feature-extraction
 language: en
 datasets:
 - SciDocs
@@ -28,6 +31,30 @@ PubMedNCL: Working with biomedical papers? Try [PubMedNCL](https://huggingface.c
 ## How to use the pretrained model
 ```python
 from transformers import AutoTokenizer, AutoModel
@@ -49,6 +76,12 @@ result = model(**inputs)
 # take the first token ([CLS] token) in the batch as the embedding
 embeddings = result.last_hidden_state[:, 0, :]
 ```
 ## Triplet Mining Parameters

 ---
 tags:
 - feature-extraction
+- sentence-transformers
+- transformers
+library_name: sentence-transformers
 language: en
 datasets:
 - SciDocs
 ## How to use the pretrained model
+### Sentence Transformers
+```python
+from sentence_transformers import SentenceTransformer
+# Load the model
+model = SentenceTransformer("malteos/scincl")
+# Concatenate the title and abstract with the [SEP] token
+papers = [
+    "BERT [SEP] We introduce a new language representation model called BERT",
+    "Attention is all you need [SEP] The dominant sequence transduction models are based on complex recurrent or convolutional neural networks",
+]
+# Inference
+embeddings = model.encode(papers)
+# Compute the (cosine) similarity between embeddings
+similarity = model.similarity(embeddings[0], embeddings[1])
+print(similarity.item())
+# => 0.8440517783164978
+```
+### Transformers
 ```python
 from transformers import AutoTokenizer, AutoModel
 # take the first token ([CLS] token) in the batch as the embedding
 embeddings = result.last_hidden_state[:, 0, :]
+# calculate the similarity
+embeddings = torch.nn.functional.normalize(embeddings, p=2, dim=1)
+similarity = (embeddings[0] @ embeddings[1].T)
+print(similarity.item())
+# => 0.8440518379211426
 ```
 ## Triplet Mining Parameters

config_sentence_transformers.json ADDED Viewed

+{
+  "__version__": {
+    "sentence_transformers": "3.0.0",
+    "transformers": "4.41.2",
+    "pytorch": "2.3.0+cu121"
+  },
+  "prompts": {},
+  "default_prompt_name": null,
+  "similarity_fn_name": "cosine"
+}

modules.json ADDED Viewed

+[
+  {
+    "idx": 0,
+    "name": "0",
+    "path": "",
+    "type": "sentence_transformers.models.Transformer"
+  },
+  {
+    "idx": 1,
+    "name": "1",
+    "path": "1_Pooling",
+    "type": "sentence_transformers.models.Pooling"
+  }
+]

sentence_bert_config.json ADDED Viewed

+{
+  "max_seq_length": 512,
+  "do_lower_case": false
+}