opensearch-project
/

opensearch-neural-sparse-encoding-v1

@@ -10,6 +10,12 @@ tags:
 - query-expansion
 - document-expansion
 - bag-of-words
 ---
 # opensearch-neural-sparse-encoding-v1
@@ -37,6 +43,69 @@ This model is trained on MS MARCO dataset.
 OpenSearch neural sparse feature supports learned sparse retrieval with lucene inverted index. Link: https://opensearch.org/docs/latest/query-dsl/specialized/neural-sparse/. The indexing and search can be performed with OpenSearch high-level API.
 ## Usage (HuggingFace)
 This model is supposed to run inside OpenSearch cluster. But you can also use it outside the cluster, with HuggingFace models API.

 - query-expansion
 - document-expansion
 - bag-of-words
+- sentence-transformers
+- sparse-encoder
+- sparse
+- splade
+pipeline_tag: feature-extraction
+library_name: sentence-transformers
 ---
 # opensearch-neural-sparse-encoding-v1
 OpenSearch neural sparse feature supports learned sparse retrieval with lucene inverted index. Link: https://opensearch.org/docs/latest/query-dsl/specialized/neural-sparse/. The indexing and search can be performed with OpenSearch high-level API.
+## Usage (Sentence Transformers)
+First install the Sentence Transformers library:
+```bash
+pip install -U sentence-transformers
+```
+Then you can load this model and run inference.
+```python
+from sentence_transformers.sparse_encoder import SparseEncoder
+# Download from the 🤗 Hub
+model = SparseEncoder("opensearch-project/opensearch-neural-sparse-encoding-v1")
+query = "What's the weather in ny now?"
+document = "Currently New York is rainy."
+query_embed = model.encode_query(query)
+document_embed = model.encode_document(document)
+sim = model.similarity(query_embed, document_embed)
+print(f"Similarity: {sim}")
+# Similarity: tensor([[22.3299]])
+# Visualize top tokens for each text
+top_k = 8
+print(f"\nTop tokens {top_k} for each text:")
+decoded_query = model.decode(query_embed, top_k=top_k)
+decoded_document = model.decode(document_embed)
+for i in range(top_k):
+    query_token, query_score = decoded_query[i]
+    doc_score = next((score for token, score in decoded_document if token == query_token), 0)
+    if doc_score != 0:
+        print(f"Token: {query_token}, Query score: {query_score:.4f}, Document score: {doc_score:.4f}")
+# Top tokens 30 for each text:
+# Token: ny, Query score: 2.9262, Document score: 2.1335
+# Token: weather, Query score: 2.5206, Document score: 1.5277
+# Token: york, Query score: 2.0373, Document score: 2.3489
+# Token: cool, Query score: 1.5786, Document score: 0.8752
+# Token: current, Query score: 1.4636, Document score: 1.5132
+# Token: season, Query score: 0.7761, Document score: 0.8860
+# Token: 2020, Query score: 0.7560, Document score: 0.6726
+# Token: summer, Query score: 0.7222, Document score: 0.6292
+# Token: nina, Query score: 0.6888, Document score: 0.6419
+# Token: storm, Query score: 0.6451, Document score: 0.8200
+# Token: brooklyn, Query score: 0.4698, Document score: 0.7635
+# Token: julian, Query score: 0.4562, Document score: 0.1208
+# Token: wow, Query score: 0.3484, Document score: 0.3903
+# Token: usa, Query score: 0.3439, Document score: 0.4160
+# Token: manhattan, Query score: 0.2751, Document score: 0.8260
+# Token: fog, Query score: 0.2013, Document score: 0.7735
+# Token: mood, Query score: 0.1989, Document score: 0.2961
+# Token: climate, Query score: 0.1653, Document score: 0.3437
+# Token: nature, Query score: 0.1191, Document score: 0.1533
+# Token: temperature, Query score: 0.0665, Document score: 0.0599
+# Token: windy, Query score: 0.0552, Document score: 0.3396
+```
 ## Usage (HuggingFace)
 This model is supposed to run inside OpenSearch cluster. But you can also use it outside the cluster, with HuggingFace models API.

config_sentence_transformers.json ADDED Viewed

	@@ -0,0 +1,14 @@

+{
+  "model_type": "SparseEncoder",
+  "__version__": {
+    "sentence_transformers": "5.0.0",
+    "transformers": "4.50.3",
+    "pytorch": "2.6.0+cu124"
+  },
+  "prompts": {
+    "query": "",
+    "document": ""
+  },
+  "default_prompt_name": null,
+  "similarity_fn_name": "dot"
+}

modules.json ADDED Viewed

	@@ -0,0 +1,14 @@

+[
+  {
+    "idx": 0,
+    "name": "0",
+    "path": "",
+    "type": "sentence_transformers.sparse_encoder.models.MLMTransformer"
+  },
+  {
+    "idx": 1,
+    "name": "1",
+    "path": "1_SpladePooling",
+    "type": "sentence_transformers.sparse_encoder.models.SpladePooling"
+  }
+]

sentence_bert_config.json ADDED Viewed

	@@ -0,0 +1,4 @@

+{
+    "max_seq_length": 512,
+    "do_lower_case": false
+}