See https://github.com/thuml/RLVR-World for examples for using this model.

Citation

@article{wu2025rlvr,
    title={RLVR-World: Training World Models with Reinforcement Learning}, 
    author={Jialong Wu and Shaofeng Yin and Ningya Feng and Mingsheng Long},
    journal={arXiv preprint arXiv:2505.13934},
    year={2025},
}

Downloads last month: 5

Safetensors

Model size

1.78B params

Tensor type

BF16

Inference Providers NEW

This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for thuml/bytesized32-world-model-rlvr-binary-reward

Base model

deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B

Finetuned

thuml/bytesized32-world-model-sft

Finetuned

(2)

this model

Dataset used to train thuml/bytesized32-world-model-rlvr-binary-reward

Collection including thuml/bytesized32-world-model-rlvr-binary-reward

RLVR-World

Collection

14 items • Updated May 26 • 1