kyutai
/

moshika-pytorch-bf16

Moshi

Safetensors

English

Model card Files Files and versions Community

jegou commited on Sep 18, 2024

Commit

f45037c

verified ·

1 Parent(s): 894dc28

Update README.md

Browse files

Files changed (1) hide show

README.md +11 -13

README.md CHANGED Viewed

@@ -1,4 +1,6 @@
 ---
 license: cc-by-4.0
 language:
 - en
@@ -25,16 +27,13 @@ Monologue” method significantly improves the linguistic quality of generated s
 ### Model Sources
-<!-- Provide the basic links for the model. -->
 - **Repository:** [repo](https://github.com/kyutai-labs/moshi)
-- **Paper:** [paper (soon!)](https://arxiv.org/abs/2409.XXXXX)
 - **Demo:** [demo](https://moshi.chat/)
 ## Uses
-<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
 ### Direct Use
 The model can be used as a conversational agent for casual conversations, basic facts and advice (e.g. recipes, trivia), roleplay, etc. However, the model has limited abilities for complex tasks and cannot access tools, but rather focues on natural, low-latency interactions.
@@ -54,8 +53,6 @@ This model is for research only and we do not recommend it for providing advices
 ## Bias, Risks, and Limitations
-<!-- This section is meant to convey both technical and sociotechnical limitations. -->
 The model has been trained with a few safeguards to try to limit potential toxic usages, however our toxicity analysis shows that it behaves in the middle of existing models with respect to textual generation. It has some bias towards certain domains and topics that are over-represented in the training data. Its capabilities are relatively limited so far and it is trained to produce only one voice to avoid impersonation. Yet, we need the perspective in time to establish the sociotechnical limitations.
@@ -92,16 +89,17 @@ The training was performed on 127 DGX nodes provided by Scaleway, accounting for
 ## Citation
 ```
-@article{defossez2024moshi,
-  title={Moshi: a speech-text foundation model for real-time dialogue},
-  authors={Alexandre D\'efossez, Laurent Mazar\'e, Manu Orsini, Am\'elie Royer, Patrick P\'erez, Herv\'e J\'egou, Edouard Grave, Neil Zeghidour},
-  year={2024},
-  month={September},
-  journal={arXiv preprint arXiv:2409.XXXXX}
 }
 ```
 ## Model Card Authors
-Alexandre Défossez, Laurent Mazaré, Manu Orsini, Amélie Royer, Patrick Pérez, Hervé Jégou, Edouard Grave, Neil Zeghidour

 ---
+# For reference on model card metadata, see the spec: https://github.com/huggingface/hub-docs/blob/main/modelcard.md?plain=1
+# Doc / guide: https://huggingface.co/docs/hub/model-cards
 license: cc-by-4.0
 language:
 - en
 ### Model Sources
 - **Repository:** [repo](https://github.com/kyutai-labs/moshi)
+- **Paper:** [paper](http://kyutai.org/Moshi.pdf)
 - **Demo:** [demo](https://moshi.chat/)
 ## Uses
 ### Direct Use
 The model can be used as a conversational agent for casual conversations, basic facts and advice (e.g. recipes, trivia), roleplay, etc. However, the model has limited abilities for complex tasks and cannot access tools, but rather focues on natural, low-latency interactions.
 ## Bias, Risks, and Limitations
 The model has been trained with a few safeguards to try to limit potential toxic usages, however our toxicity analysis shows that it behaves in the middle of existing models with respect to textual generation. It has some bias towards certain domains and topics that are over-represented in the training data. Its capabilities are relatively limited so far and it is trained to produce only one voice to avoid impersonation. Yet, we need the perspective in time to establish the sociotechnical limitations.
 ## Citation
 ```
+@techreport{citation-key,
+    author = {Alexandre D\'efossez, Laurent Mazar\'e, Manu Orsini, Am\'elie Royer, Patrick P\'erez, Herv\'e J\'egou, Edouard Grave, Neil Zeghidour},
+    title = {Moshi: a speech-text foundation model for real-time dialogue},
+    institution = {Kyutai},
+    year={2024},
+    month={September},
+    url={http://kyutai.org/Moshi.pdf},
 }
 ```
 ## Model Card Authors
+Alexandre Défossez, Laurent Mazaré, Manu Orsini, Amélie Royer, Patrick Pérez, Hervé Jégou, Edouard Grave, Neil Zeghidour