Hi Memo Team,
I'm a regular user of the Podcast feature and I'd like to request additional English voice options, specifically a younger, more energetic female voice for educational podcasts aimed at children and teenagers.
The Problem:
Currently, when I generate podcasts in Cantonese, the female host voice (Sarah) sounds young, vibrant, and energetic β perfect for engaging younger learners. However, when I generate the same podcast in English, the female voice defaults to a much older, more mature tone that sounds like a teacher or a parent. This creates a disconnect for students, especially:
Year 4-6 students (ages 8-11) who need youthful, relatable voices
Auditory learners who rely heavily on voice engagement for retention
Hong Kong immigrant children in the UK who respond better to peer-like voices than teacher-like voices
No matter how I adjust the custom instructions (e.g., specifying "teenage voice," "NOT a teacher," "hyperactive, high-pitched"), the underlying English TTS voice model remains the same mature female voice. This appears to be a TTS engine limitation rather than a prompt issue.
The Request:
Please add multiple English voice options to the Podcast feature, similar to how other TTS platforms (e.g., ElevenLabs) offer a voice library. At minimum, please include:
A young female voice (teenage/early 20s, energetic, suitable for children's content)
A young male voice (teenage/early 20s, casual, suitable for peer-like dialogue)
Ideally, allow users to select voices per host (Sarah and Alex) from a dropdown menu before generating the podcast.
Why This Matters:
Memo's podcast feature is uniquely powerful for auditory learners
The current mature English voice undermines the "fun, peer-to-peer banter" format that makes the Cantonese version so effective
Adding voice options would significantly improve engagement and learning outcomes for English-language content
Competitor platforms already offer voice selection β this would keep Memo competitive
Additional Note:
A related issue: when Cantonese phrases are written in English romanisation (e.g., "m goi", "hou faan"), the English TTS engine reads them out in a very unnatural way that breaks immersion. This further highlights the need for either native Cantonese TTS support within English podcasts, or clearer guidance to avoid romanised Cantonese.
Thank you for considering this request. I'm happy to provide more details or test any beta voice options if available.
Best regards
MC
Share update with 0 linked conversations as well
In Review
π₯ Feedback
About 15 hours ago

MC
Get notified by email when there are changes.
In Review
π₯ Feedback
About 15 hours ago

MC
Get notified by email when there are changes.