Curated OCR models for documents, languages, handwriting and text in images. Browse four collections with short practical notes.
Daniel van Strien PRO
AI & ML interests
Machine Learning Librarian
Recent Activity
updated a Space about 3 hours ago
davanstrien/benchmark-race updated a dataset about 8 hours ago
librarian-bots/model_cards_with_metadata updated a dataset about 8 hours ago
librarian-bots/dataset_cards_with_metadataOrganizations
OCR: Handwriting & archives
Handwriting, historical print and manuscript recognition. Notes on languages, text-line segmentation and transcription conventions.
-
Riksarkivet/trocr-base-handwritten-hist-swe-2
Image-to-Text • 0.4B • Updated • 89.9k • 17 -
BDRC/tibetan-ocr
Image-Text-to-Text • 0.8B • Updated • 227 • 3 -
isaacmg/qwen3-vl-8b-hebrew-v19a-ckpt
Image-Text-to-Text • Updated • 44 -
magistermilitum/tridis_v2_HTR_historical_manuscripts
0.6B • Updated • 308 • 7
OCR: Documents
Roughly ordered by recent releases, useful updates and current usage. Practical OCR and document parsing models. Reviewed September 2026.
Datasets Wrapped 2025: Reasoning
The reasoning datasets that defined 2025. Part 1 of Datasets Wrapped 2025. #DatasetsWrapped2025
hub-tldr
Creating a smol model for tl;dr-ing the hub
-
davanstrien/Smol-Hub-tldr
Text Generation • 0.4B • Updated • 67 • 11 - Running93
Semantic Hugging Face Hub Search
🔎93Search Hugging Face datasets and models by meaning
-
davanstrien/hub-tldr-dataset-summaries-llama
Viewer • Updated • 5k • 33 • 1 -
davanstrien/hub-tldr-model-summaries-llama
Viewer • Updated • 5k • 27 • 1
synthetic-data-generation-demos
A collection of demos for various approaches to synthetic data generation
- Runtime errorAgents8
Genstruct 7B
👀8 - Running on ZeroAgentsFeatured86
Instruction Synthesizer
🐠86Generate instruction-response pairs from text
- Running on ZeroAgentsFeatured73
Magpie
🐦73Generate and rate instruction-response pairs
- Runtime errorAgents11
Bonito
💬11Generate task-specific instructions and responses from text
Synthetic (text) Dataset Generation
Papers about synthetic dataset generation
-
Better Synthetic Data by Retrieving and Transforming Existing Datasets
Paper • 2404.14361 • Published • 2 -
Generative AI for Synthetic Data Generation: Methods, Challenges and the Future
Paper • 2403.04190 • Published • 1 -
Best Practices and Lessons Learned on Synthetic Data for Language Models
Paper • 2404.07503 • Published • 32 -
A Multi-Faceted Evaluation Framework for Assessing Synthetic Data Generated by Large Language Models
Paper • 2404.14445 • Published
Historic language modeling
This collection contains models, datasets and spaces related to historic language models i.e. language models trained on historic data
Image Preference Optimization Datasets
Datasets suitable for Image Preference Optimization based on their colum names
OCR: Text recognition & pipelines
Text-line and region recognisers, plus detection and recognition pipelines for documents, manga and text in photographs.
OCR: Languages & scripts
OCR for particular languages and writing systems, from whole-page readers to dedicated text-line recognisers.
Video-game gameplay datasets for agent training
Recent Hub datasets of video game gameplay (frames/video + actions, replays, trajectories) for imitation learning and game-playing agents.
Reasoning Required?
-
davanstrien/reasoning-required
Viewer • Updated • 5k • 106 • 20 -
davanstrien/ModernBERT-based-Reasoning-Required
Text Classification • 0.1B • Updated • 15 • 11 -
davanstrien/fineweb-with-reasoning-scores-and-topics
Viewer • Updated • 10k • 36 • 2 -
davanstrien/fine-reasoning-questions
Viewer • Updated • 244 • 66 • 19
Maths reasoning
Maths reasoning datasets found using https://huggingface.co/spaces/librarian-bots/huggingface-datasets-semantic-search
sentence-transformers-from-synthetic-data
Example of using distilabel to generate synthetic triplets data for fine-tuning a Sentence Transformer model
-
bigcode/self-oss-instruct-sc2-exec-filter-50k
Viewer • Updated • 50.7k • 33.9k • 107 -
davanstrien/similarity-dataset-sc2-8b
Viewer • Updated • 2.32k • 49 • 6 -
davanstrien/code-prompt-similarity-model
Sentence Similarity • 0.1B • Updated • 89 • 6 -
davanstrien/abstract-wiki
Viewer • Updated • 5k • 16 • 2
haiku
🌸 This is a collection of synthetic datasets built to help improve the ability of open language models to better write haikus through the use of DPO
Probably DPO datasets
A collection of datasets that probably support DPO
query-to-hub-datasets-viewer-project
OCR on the Hub
Curated OCR models for documents, languages, handwriting and text in images. Browse four collections with short practical notes.
OCR: Text recognition & pipelines
Text-line and region recognisers, plus detection and recognition pipelines for documents, manga and text in photographs.
OCR: Handwriting & archives
Handwriting, historical print and manuscript recognition. Notes on languages, text-line segmentation and transcription conventions.
-
Riksarkivet/trocr-base-handwritten-hist-swe-2
Image-to-Text • 0.4B • Updated • 89.9k • 17 -
BDRC/tibetan-ocr
Image-Text-to-Text • 0.8B • Updated • 227 • 3 -
isaacmg/qwen3-vl-8b-hebrew-v19a-ckpt
Image-Text-to-Text • Updated • 44 -
magistermilitum/tridis_v2_HTR_historical_manuscripts
0.6B • Updated • 308 • 7
OCR: Languages & scripts
OCR for particular languages and writing systems, from whole-page readers to dedicated text-line recognisers.
OCR: Documents
Roughly ordered by recent releases, useful updates and current usage. Practical OCR and document parsing models. Reviewed September 2026.
Video-game gameplay datasets for agent training
Recent Hub datasets of video game gameplay (frames/video + actions, replays, trajectories) for imitation learning and game-playing agents.
Datasets Wrapped 2025: Reasoning
The reasoning datasets that defined 2025. Part 1 of Datasets Wrapped 2025. #DatasetsWrapped2025
Reasoning Required?
-
davanstrien/reasoning-required
Viewer • Updated • 5k • 106 • 20 -
davanstrien/ModernBERT-based-Reasoning-Required
Text Classification • 0.1B • Updated • 15 • 11 -
davanstrien/fineweb-with-reasoning-scores-and-topics
Viewer • Updated • 10k • 36 • 2 -
davanstrien/fine-reasoning-questions
Viewer • Updated • 244 • 66 • 19
hub-tldr
Creating a smol model for tl;dr-ing the hub
-
davanstrien/Smol-Hub-tldr
Text Generation • 0.4B • Updated • 67 • 11 - Running93
Semantic Hugging Face Hub Search
🔎93Search Hugging Face datasets and models by meaning
-
davanstrien/hub-tldr-dataset-summaries-llama
Viewer • Updated • 5k • 33 • 1 -
davanstrien/hub-tldr-model-summaries-llama
Viewer • Updated • 5k • 27 • 1
Maths reasoning
Maths reasoning datasets found using https://huggingface.co/spaces/librarian-bots/huggingface-datasets-semantic-search
synthetic-data-generation-demos
A collection of demos for various approaches to synthetic data generation
- Runtime errorAgents8
Genstruct 7B
👀8 - Running on ZeroAgentsFeatured86
Instruction Synthesizer
🐠86Generate instruction-response pairs from text
- Running on ZeroAgentsFeatured73
Magpie
🐦73Generate and rate instruction-response pairs
- Runtime errorAgents11
Bonito
💬11Generate task-specific instructions and responses from text
sentence-transformers-from-synthetic-data
Example of using distilabel to generate synthetic triplets data for fine-tuning a Sentence Transformer model
-
bigcode/self-oss-instruct-sc2-exec-filter-50k
Viewer • Updated • 50.7k • 33.9k • 107 -
davanstrien/similarity-dataset-sc2-8b
Viewer • Updated • 2.32k • 49 • 6 -
davanstrien/code-prompt-similarity-model
Sentence Similarity • 0.1B • Updated • 89 • 6 -
davanstrien/abstract-wiki
Viewer • Updated • 5k • 16 • 2
Synthetic (text) Dataset Generation
Papers about synthetic dataset generation
-
Better Synthetic Data by Retrieving and Transforming Existing Datasets
Paper • 2404.14361 • Published • 2 -
Generative AI for Synthetic Data Generation: Methods, Challenges and the Future
Paper • 2403.04190 • Published • 1 -
Best Practices and Lessons Learned on Synthetic Data for Language Models
Paper • 2404.07503 • Published • 32 -
A Multi-Faceted Evaluation Framework for Assessing Synthetic Data Generated by Large Language Models
Paper • 2404.14445 • Published
haiku
🌸 This is a collection of synthetic datasets built to help improve the ability of open language models to better write haikus through the use of DPO
Historic language modeling
This collection contains models, datasets and spaces related to historic language models i.e. language models trained on historic data
Probably DPO datasets
A collection of datasets that probably support DPO
Image Preference Optimization Datasets
Datasets suitable for Image Preference Optimization based on their colum names
query-to-hub-datasets-viewer-project