Stefan Schweter's picture

Stefan Schweter PRO

stefan-it

·

AI & ML interests

Flair Library, NER & PoS Tagging, LM Pretraining (mostly encoder-only), Historical Language Models

Recent Activity

reacted to davanstrien's post with 🚀 1 day ago

The https://huggingface.co/datasets/data-is-better-together/fineweb-c dataset is growing! This week a few more languages have got 1,000 annotations for the educational quality of data from https://huggingface.co/datasets/HuggingFaceFW/fineweb-2. Why should you care? The quality of pre-training data can have a big impact on the performance of downstream language models trained on that data (https://huggingface.co/spaces/HuggingFaceFW/blogpost-fineweb-v1). Being able to filter by educational quality is on way of improving the quality of the data you use for training an LLM. Very importantly this approach can also reduce the amount of data needed for pertaining. Why not use an LLM? LLMs can be used to annotate educational quality for a subset of data. This data can then be used to train a smaller encoder only model to label the full dataset. However, this may not work well for languages outside of english. This is where fineweb-c (community) comes in. The community is annotating the educational quality of fineweb2 data. Currently 114 languages have some annotations. These annotations will enable a number of things: - Evaluate whether an LLM can label the educational quality for texts in that language well - Directly be used for training quality classifiers - Help discover other rules and huerisitcs for refining fineweb2 further for different languages. This week the following languages where done: Swedish thanks to: @Lauler @AntonVic @ohallstrom @bjarlestam @menbom @Ekgren @apsod Ukrainian thanks to: @hannayukhymenko @robinhad @realPivo @RabotiahovDmytro @reciprocate Assamese thanks to: @moyoor97 @Arpanjyoti @nawaf-helmi123 @pahigogoi1 @aelhence @kishorekashyap Want to learn more: https://huggingface.co/blog/davanstrien/fineweb2-community Contribute yourself here: https://huggingface.co/spaces/data-is-better-together/fineweb-c

commented a paper 1 day ago

Building Foundations for Natural Language Processing of Historical Turkish: Resources and Models

upvoted a paper 1 day ago

Building Foundations for Natural Language Processing of Historical Turkish: Resources and Models

View all activity

Articles

Fine-tune Flair Models on NER Dataset with 🤗 AutoTrain SpaceRunner

Organizations

stefan-it's activity

commented a paper 1 day ago

Building Foundations for Natural Language Processing of Historical Turkish: Resources and Models

Paper • 2501.04828 • Published 3 days ago • 3 •

New activity in lang-uk/electra-base-ukrainian-cased-discriminator 7 days ago

Adding `safetensors` variant of this model

#1 opened 7 days ago by

commented a paper 20 days ago

Fietje: An open, efficient LLM for Dutch

Paper • 2412.15450 • Published 23 days ago • 4 •

commented 2 papers 23 days ago

Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference

Paper • 2412.13663 • Published 25 days ago • 121 •

Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference

Paper • 2412.13663 • Published 25 days ago • 121 •

New activity in stefan-it/span-marker-gelectra-large-germeval14 27 days ago

Adding `safetensors` variant of this model

#1 opened 27 days ago by

commented a paper 28 days ago

OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ Languages

Paper • 2412.09587 • Published about 1 month ago • 3 •

New activity in cis-lmu/bar-fineweb 30 days ago

Librarian Bot: Add language metadata for dataset

#2 opened about 1 month ago by

New activity in dbmdz/bert-base-german-europeana-uncased about 1 month ago

Adding `safetensors` variant of this model

#1 opened about 1 month ago by

New activity in PleIAs/journaux-lm-v1 about 1 month ago

Adding `safetensors` variant of this model

#1 opened about 1 month ago by

New activity in stefan-it/zeitungs-lm-v1 about 1 month ago

Adding `safetensors` variant of this model

#1 opened about 1 month ago by

commented 2 papers about 1 month ago

GerPS-Compare: Comparing NER methods for legal norm analysis

Paper • 2412.02427 • Published Dec 3, 2024 •

Scaling BERT Models for Turkish Automatic Punctuation and Capitalization Correction

Paper • 2412.02698 • Published Dec 3, 2024 •

New activity in alpindale/two-million-bluesky-posts about 1 month ago

Are my pronouns right?

#23 opened about 1 month ago by

pronounintifada

commented a paper about 1 month ago

CamemBERT 2.0: A Smarter French Language Model Aged to Perfection

Paper • 2411.08868 • Published Nov 13, 2024 • 12 •

New activity in ConquestAce/bluesky-did about 1 month ago

Many thanks!

#1 opened about 1 month ago by

New activity in alpindale/two-million-bluesky-posts about 2 months ago

🚩 Report: Ethical issue(s)

#19 opened about 2 months ago by

🚩 Report: Ethical issue(s)

#2 opened about 2 months ago by

New activity in bluesky-community/one-million-bluesky-posts about 2 months ago

Discussion about dataset removal

#12 opened about 2 months ago by

tobiasdrundridge

New activity in dbmdz/electra-small-turkish-cased-discriminator about 2 months ago

Adding `safetensors` variant of this model

#1 opened about 2 months ago by