Speech Sources — Unchunked Collection Long-form or minimally processed source speech datasets. • 11 items • Updated about 1 hour ago
TTS — Candidate Training Corpora Collection Prepared TTS corpora awaiting end-to-end lineage and release validation. • 15 items • Updated about 1 hour ago
Model-specific Encodings and Artifacts Collection Tokenizer-, codec-, or model-specific derived representations. • 18 items • Updated about 1 hour ago
STT — Aligned Upstream Candidates Collection Restored or aligned STT repositories that are not in the validated 14-repository upstream set. • 17 items • Updated about 1 hour ago
STT — Validated Upstream Inputs Collection The 14 upstream repositories used by the completed STT training run. These are not final immutable training releases. • 14 items • Updated about 1 hour ago
Speech Sources — Segmented and Chunked Collection Reusable segmented or chunked speech datasets. • 10 items • Updated about 1 hour ago
TTS — Raw Pair Plans (Not Train-Ready) Collection Raw clone-pair planning metadata. These repositories are inputs to consensus and packaging, not training releases. • 11 items • Updated about 1 hour ago
TTS — Packaged Clone-Pair Shards (Candidate) Collection Packaged reference/target pair shards awaiting consensus, reservation, pruning, and lineage certification. • 10 items • Updated about 1 hour ago