CHÀNG 唱
Learn Chinese by singing. A pronunciation tool built on a habit people already have instead of a new one.

The problem
Language pronunciation only improves through practice, and keeping that practice consistent means making it enjoyable. Most apps solve this through gamification, using tools like streaks and points to build the habit from scratch.
Instead of building a new routine, the challenge was to design around something people already do.
Who it’s for
- People who want to practice pronunciation and speaking
Design requirements
- Builds on something users already do.
- Involves the user speaking out loud, repeatedly.
- Splits practice into short chunks for quick wins.
- The user chooses the material, not a fixed catalog.
Scope: Pronunciation only (no grammar, translation, or vocabulary).
Ideate
Most language apps use gamification to force a brand-new habit into the user’s routine.
What if we attached the practice to something users already do?
| Habit | 1. Existing habit | 2. Involves speaking out loud | 3. Natural small chunks | 4. User’s own material |
|---|---|---|---|---|
| Podcasts | Yes | No | No | Yes |
| Shows and movies | Yes | No | Yes | Yes |
| Reading | Yes | Rarely | Yes | Yes |
| Talking with family | Yes | Yes | No | No |
| Singing | Yes | Yes | Yes | Yes |
Most habits fall short of the design requirements. Either the user isn’t speaking out loud, or the format doesn’t fit short practice sessions. Singing is the exception: it naturally involves speaking out loud, breaks into small chunks, and uses songs people already love to repeat.
The trade-off: It relies on having favorite songs in the language, which complete beginners might not have yet.
Can singing be scored?
Testing sung input on a prototype Mandarin speech checker proved unreliable. Mandarin tones depend on pitch, but singing forces pitch to follow melody, erasing the data needed for scoring.
The fix: split the loop
- Speak for accurate scoring.
- Sing to lock in the habit once sounds are solid.
Low fidelity wireframes
How much does the user say at once?

Verse — no
- Full-verse practice makes chunks too long for quick wins.
- Trade off: Line-to-line flow across verses.
Word — no
- Isolated word practice ignores natural singing habits.
- Trade off: The fastest possible wins.
Line — yes
- Smallest unit that still feels like real singing.
- Short enough to repeat many times.
- A win every few seconds.
How does the user see what was wrong?
Requirements defined the practice unit; feedback came down to a single question: does it guide the fix?

- Explored two other versions that only flagged wrong words and error distance.
- Only this version displays what was actually said, so the contrast itself acts as the instruction.
How does the user practice it?

- Explored line repeats and end reviews. One wastes reps on correct sounds, the other delays feedback.
- Targeted drilling isolates the error, then returns to the full line.
Prototype
Here is the complete experience in motion, taking the user from song selection through speaking and targeted practice.
Testing
What I tested
- Do users accept speaking first?
- Do drills stick in the full line?
- Do users return without reminders?
Method
- Day 1: Unmoderated test with 3 users doing 1 full loop.
- Week 2: Follow-up on organic usage and drop-off reasons.
Findings
- Speaking first: Accepted, but needed context on why and when to sing. Hidden info buttons went unopened.
- Drilling: Fixes carried directly back into the full line.
- Retention: Goal-driven users stayed (7 visits in 2 weeks); casual users churned after two sessions.
Takeaway
Singing sustains habits for motivated users, but doesn’t create motivation from scratch.
What I changed
Problem: Unclear transition between speaking and singing.
Fix: Labeled the speaking step directly, unlocking the singing step only after a line passes.

What I’d do next
- Retention: Design a lighter return trigger for users without explicit goals.
- Song import: Streamline lyric setup while navigating copyright constraints.
- Audio experience: Integrate background music without adding legal or setup complexity.