Hugging Face
Models
Datasets
Spaces
Posts
Docs
Enterprise
Pricing
Log In
Sign Up
line-corporation
/
p-sacpo
like
3
Follow
LINE
34
Reinforcement Learning
Transformers
Safetensors
PKU-Alignment/PKU-SafeRLHF-30K
English
llama
text-generation
reinforcement-learning-from-human-feedback
rlhf
safety
ai-safety
alpaca
text-generation-inference
Inference Endpoints
arxiv:
2404.11049
arxiv:
2305.18290
License:
cc-by-nc-4.0
Model card
Files
Files and versions
Community
Train
Deploy
Use this model
2901c06
p-sacpo
Commit History
Update README.md
2901c06
verified
akifumiwachi
commited on
Jun 21
Update README.md
2bbc875
verified
reisato80
commited on
Jun 21
Upload LlamaForCausalLM
b36098b
verified
reisato80
commited on
Jun 19
Upload tokenizer
eff089a
verified
reisato80
commited on
Jun 19
initial commit
234d482
verified
ospo-line
commited on
Jun 19