neuralmagic
/

Llama-2-7b-pruned70-retrained

Text Generation

text-generation-inference

Inference Endpoints

Model card Files Files and versions Community

mgoin commited on Mar 18, 2024

Commit

617c84c

·

verified ·

1 Parent(s): f26bc8f

Update README.md

Files changed (1) hide show

README.md +2 -2

README.md CHANGED Viewed

@@ -1,5 +1,5 @@
 ---
-base_model: meta-llama/Llama-2-7b-hf
 inference: true
 model_type: llama
 datasets:
@@ -10,7 +10,7 @@ tags:
 # Llama-2-7b-pruned70-retrained
-This repo contains model files for a [Llama 2 7B](https://huggingface.co/meta-llama/Llama-2-7b-hf) model that has had 70% of the parameters pruned in one-shot with [SparseGPT](https://arxiv.org/abs/2301.00774), then retrained by [Cerebras](https://huggingface.co/cerebras) with XXB [UPDATE] tokens from SlimPajama while maintaining sparsity.
 **Authors**: Neural Magic, Cerebras

 ---
+base_model: neuralmagic/Llama-2-7b-prune50-retrained
 inference: true
 model_type: llama
 datasets:
 # Llama-2-7b-pruned70-retrained
+This repo contains model files for a [Llama 2 7B](https://huggingface.co/meta-llama/Llama-2-7b-hf) model that has had 50% of the parameters pruned in one-shot with [SparseGPT](https://arxiv.org/abs/2301.00774), then retrained by [Cerebras](https://huggingface.co/cerebras) with 50B [UPDATE] from SlimPajama while maintaining sparsity. It was then one-shot pruned to 70% sparsity and trained for another 100B tokens.
 **Authors**: Neural Magic, Cerebras