Jason Stillerman

stillerman

stillerman

AI & ML interests

None yet

Recent Activity

liked a model about 8 hours ago

teapotai/teapotllm

reacted to merve's post with 👍 1 day ago

So many open releases at Hugging Face past week 🤯 recapping all here ⤵️ https://huggingface.co/collections/merve/march-21-releases-67dbe10e185f199e656140ae 👀 Multimodal > Mistral AI released a 24B vision LM, both base and instruction FT versions, sota 🔥 (OS) > with IBM we released SmolDocling, a sota 256M document parser with Apache 2.0 license (OS) > SpatialLM is a new vision LM that outputs 3D bounding boxes, comes with 0.5B (QwenVL based) and 1B (Llama based) variants > SkyWork released SkyWork-R1V-38B, new vision reasoning model (OS) 💬 LLMs > NVIDIA released new Nemotron models in 49B and 8B with their post-training dataset > LG released EXAONE, new reasoning models in 2.4B, 7.8B and 32B > Dataset: Glaive AI released a new reasoning dataset of 22M+ examples > Dataset: NVIDIA released new helpfulness dataset HelpSteer3 > Dataset: OpenManusRL is a new agent dataset based on ReAct framework (OS) > Open-R1 team released OlympicCoder, new competitive coder model in 7B and 32B > Dataset: GeneralThought-430K is a new reasoning dataset (OS) 🖼️ Image Generation/Computer Vision > Roboflow released RF-DETR, new real-time sota object detector (OS) 🔥 > YOLOE is a new real-time zero-shot object detector with text and visual prompts 🥹 > Stability AI released Stable Virtual Camera, a new novel view synthesis model > Tencent released Hunyuan3D-2mini, new small and fast 3D asset generation model > ByteDance released InfiniteYou, new realistic photo generation model > StarVector is a new 8B model that generates svg from images > FlexWorld is a new model that expands 3D views (OS) 🎤 Audio > Sesame released CSM-1B new speech generation model (OS) 🤖 Robotics > NVIDIA released GR00T, new robotics model for generalized reasoning and skills, along with the dataset *OS ones have Apache 2.0 or MIT license

reacted to merve's post with 🤗 1 day ago

View all activity

Organizations

stillerman's activity

liked a model about 8 hours ago

teapotai/teapotllm

Text2Text Generation • Updated 1 day ago • 2.65k • • 55

reacted to merve's post with 👍🤗 1 day ago

Post

2925

So many open releases at Hugging Face past week 🤯 recapping all here ⤵️ merve/march-21-releases-67dbe10e185f199e656140ae

👀 Multimodal
> Mistral AI released a 24B vision LM, both base and instruction FT versions, sota 🔥 (OS)
> with IBM we released SmolDocling, a sota 256M document parser with Apache 2.0 license (OS)
> SpatialLM is a new vision LM that outputs 3D bounding boxes, comes with 0.5B (QwenVL based) and 1B (Llama based) variants
> SkyWork released SkyWork-R1V-38B, new vision reasoning model (OS)

💬 LLMs
> NVIDIA released new Nemotron models in 49B and 8B with their post-training dataset
> LG released EXAONE, new reasoning models in 2.4B, 7.8B and 32B
> Dataset: Glaive AI released a new reasoning dataset of 22M+ examples
> Dataset: NVIDIA released new helpfulness dataset HelpSteer3
> Dataset: OpenManusRL is a new agent dataset based on ReAct framework (OS)
> Open-R1 team released OlympicCoder, new competitive coder model in 7B and 32B
> Dataset: GeneralThought-430K is a new reasoning dataset (OS)

🖼️ Image Generation/Computer Vision
> Roboflow released RF-DETR, new real-time sota object detector (OS) 🔥
> YOLOE is a new real-time zero-shot object detector with text and visual prompts 🥹
> Stability AI released Stable Virtual Camera, a new novel view synthesis model
> Tencent released Hunyuan3D-2mini, new small and fast 3D asset generation model
> ByteDance released InfiniteYou, new realistic photo generation model
> StarVector is a new 8B model that generates svg from images
> FlexWorld is a new model that expands 3D views (OS)

🎤 Audio
> Sesame released CSM-1B new speech generation model (OS)

🤖 Robotics
> NVIDIA released GR00T, new robotics model for generalized reasoning and skills, along with the dataset

*OS ones have Apache 2.0 or MIT license

liked a Space 5 days ago

Dataground

✍

Create an account or log in to Hugging Face

liked a Space 12 days ago

The Distill Template

🌌

Craft Beautiful Blogs

upvoted a paper 13 days ago

Towards the Law of Capacity Gap in Distilling Language Models

Paper • 2311.07052 • Published Nov 13, 2023 • 2

liked a Space 27 days ago

154

LLaDA

🚀

Large Language Diffusion Models

liked a dataset about 1 month ago

HuggingFaceFW/fineweb-edu

Viewer • Updated Jan 31 • 3.3B • 389k • 655

liked a model about 1 month ago

microsoft/wham

Updated Feb 21 • 6.31k • 248

liked 3 Spaces about 1 month ago

FarmingGame

🚀

Play KexFarm Unity game

2.34k

The Ultra-Scale Playbook

🌌

The ultimate guide to training LLM on large GPU Clusters

207

AI Podcast Generator

🎙

Generate Podcast using Kokoro-TTS!

liked a model about 1 month ago

stillerman/FineLlama-0.2

Updated Feb 16 • 15 • 1

updated a model about 1 month ago

stillerman/FineLlama-0.2

Updated Feb 16 • 15 • 1

published a model about 1 month ago

stillerman/FineLlama-0.2

Updated Feb 16 • 15 • 1

reacted to davanstrien's post with ❤️ about 1 month ago

Post

1938

How do you make 1M+ Hugging Face models & datasets more discoverable?

davanstrien/Smol-Hub-tldr!

I fine-tuned HuggingFaceTB/SmolLM2-360M to generate one-line summaries from a model or dataset README.

Its own self-description?
"A model for generating concise summaries of model & dataset cards from the Hugging Face Hub"

The goal? Make it easier to find the right models and datasets for your specific needs. It's already powering a semantic search for datasets Space.

It's still a WIP but thanks to @loubnabnl , @anton-l , @eliebak et al, for cooking such a nice base model for fine-tuning small, efficient models for specific domains and tasks. 🙏

liked 2 datasets about 1 month ago

open-r1/OpenR1-Math-Raw

Viewer • Updated 29 days ago • 516k • 1.19k • 72

open-r1/OpenR1-Math-220k

Viewer • Updated Feb 18 • 450k • 52.4k • 526

liked a model about 2 months ago

kudzueye/boreal-hl-v1

Text-to-Video • Updated Feb 10 • 118