Public accounting

The corpus tracker

The corpus is the community's asset, so its growth is public. Every number below updates as recordings arrive. Personal details stay private — the tracker shows only what a contributor has consented to share.

0
Voices contributed
0h 00m
Thakk recorded
0
Recordings in the archive
The corpus is open — be among the first voices.Every recording is an heirloom: one elder, one story, one blessing at a time.The corpus is open — be among the first voices.Every recording is an heirloom: one elder, one story, one blessing at a time.

Milestone progress


Phase 0 · Foundation & pilot0.0h / 100h
Phase 1 · Corpus at scale0.0h / 400h
Phase 2 · Models v10.0h / 700h
Phase 3 · The speaking companion0.0h / 1000h

👂 Community-validated: 0.0h · help verify clips

Targets are hours of raw described audio per the roadmap. Transcribed-hour milestones are published with each corpus release.

Dialect balance

The balance chart appears once the first recordings arrive.

Both dialects are collected by design — quotas keep the models honest.

How the backend tracker works

Every clip lands in the community database with its metadata — speaker, village, okka, dialect, age band, content type, consent tier, and duration — ready for the AI layer: automatic pre-transcription, quality scoring, and dialect analysis feed a human review queue, so people only correct, never start from scratch. Recordings with archive-only consent never leave the vault.

Add your voice to these numbers