Booru Ecosystem & Dataset Harvesting Suite
Unified dataset ingestion, smart tracking, automated tagging, and scheduling pipeline for mediaboard resources
A complete ecosystem of tools, browser extensions, and backend indexing engines designed to collect, track, tag, and schedule high-volume media resources from imageboards for generative text-to-image dataset engineering.
System Components
- Runebooru Engine: Lightweight, high-throughput Danbooru-style imageboard server tailored specifically for text-to-image training dataset indexing and search.
- Szurubooru Browser Extension: Context-menu browser extension for Chromium/Edge enabling 1-click image uploads, automated metadata scraping, and tag ingestion into Szurubooru instances.
- Smart Resource Tracking & Scheduling: Automated background scraping and rate-limited scheduling pipelines to harvest high-resolution visual assets and associated tag taxonomies.
- BooruGPT: Lightweight transformer fine-tuning pipeline for tag sequence modeling and auto-prompt completion.