Rollout
36 months, four phases
The program moves from a 90-day pilot to a full stack of Kodava language AI in daily community use. Every phase has public milestones — you can watch them move on the corpus tracker.
Phase 0 · Months 0–3
Foundation & pilot — LIVE NOW
- Guardianship licence and consent tiers drafted; language council seated
- Recording kit and metadata standard finalized; orthography convention v1 with Academy and university partners
- Pilot: 25 elders and 200 crowd contributors recorded; Project Vaani's Kodagu audio audited
- This site: Level 0 voice collection open, live corpus tracker running
100 hours raw audio · 10 hours transcribed · first Whisper fine-tune demo
Phase 1 · Months 3–9
Corpus at scale
- Field teams covering both dialects (Mendele & Kiggat)
- Crowd app public, with festival and Padayatra recording drives
- Paid transcription team of 8–10 trained and working
- Legacy media digitization running; first ELAR / Indian archive deposits
400 hours raw · 50 hours transcribed · ASR v1 in the correction loop · TTS voice donors recorded
Phase 2 · Months 9–18
Models v1 & first products
- Thakk Archive listening site live; talking dictionary launched
- ASR v2 usable for search and subtitling; TTS v1 speaks
- Parallel corpus of 100k sentence pairs
- First Kodava–Kannada–English translation model
700 hours raw · 100 hours transcribed · ASR word error rate under 25% and falling
Phase 3 · Months 18–36
The speaking companion
- Ainmane speech-to-speech companion in beta, then public
- WhatsApp bot for the diaspora; classroom pack in schools
- Annual corpus releases; sustainability plan (grants, CSR, diaspora endowment) operating
1,000+ hours raw · 150+ hours transcribed · conversational LLM adapted to Kodava
Kickoff
The first 90 days
| Days | Action | Outcome |
|---|---|---|
| Days 1–15 | Meet the founding circle with the Karnataka Kodava Sahitya Academy, Akhila Kodava Samaja, the different Kodava Samajas, linguists, diaspora technologists, and elder representatives. Draft the guardianship licence and the consent forms, and open conversations with AI4Bharat, Vaani/ARTPARK, and Karya. | Governance skeleton and partner intent in writing |
| Days 16–30 | Audit existing material — Vaani Kodagu audio, Academy recordings, radio archives, family tapes. Define the metadata standard and Kannada-script orthography convention v1. Buy 3 field recording kits. | Data inventory · recording standard · kits ready |
| Days 31–60 | Pilot elder sessions: 25 elders across both dialects, 60+ hours recorded. Stand up storage and the transcription workflow; transcribe the first 5 hours. This website's 'Record 30 seconds of Thakk' drive pushed through the Nada Kodagu network. | First real corpus — the community sees and hears the project |
| Days 61–90 | Fine-tune the first Whisper checkpoint on pilot data and demo live transcription of an elder's story at a public event. Publish the manifesto and 3-year roadmap. Submit the first grant applications (ELDP, Bhashini, CSR). | Working AI demo · public launch · funding pipeline |
A note on momentum: the Kaveri-to-Igguthappa Padayatra in late August is a ready-made corpus event — a “voices of the walk” drive where every walker records a story, a song, or a blessing in Thakk along the route.
Eyes open
Risks & how we manage them
| Risk | Mitigation |
|---|---|
| Orthography disputes stall transcription | A pragmatic Kannada-script convention v1 as project standard; audio stays primary so any future script can be layered on later |
| Elder speakers pass before capture | Phase 0 prioritizes the oldest 100 voices; depth sessions start in month 2, before any technology is built |
| Volunteer energy fades | Paid Karya-style work for the grind; festivals and leaderboards for the fun; visible annual releases |
| Data extraction by outside actors | Guardianship licence, consent tiers, trust custody, gated commercial access from day one |
| Dialect imbalance skews models | Dialect quotas in collection targets and in every evaluation set |
| Funding gaps | Diversified pipeline: ELDP, Bhashini/SPPEL, Karnataka cultural funds, CSR, diaspora crowdfunding |