Public accounting
The corpus tracker
The corpus is the community's asset, so its growth is public. Every number below updates as recordings arrive. Personal details stay private — the tracker shows only what a contributor has consented to share.
Milestone progress
👂 Community-validated: 0.0h · help verify clips
Targets are hours of raw described audio per the roadmap. Transcribed-hour milestones are published with each corpus release.
Dialect balance
The balance chart appears once the first recordings arrive.
Both dialects are collected by design — quotas keep the models honest.
How the backend tracker works
Every clip lands in the community database with its metadata — speaker, village, okka, dialect, age band, content type, consent tier, and duration — ready for the AI layer: automatic pre-transcription, quality scoring, and dialect analysis feed a human review queue, so people only correct, never start from scratch. Recordings with archive-only consent never leave the vault.