
We have trained an AI model called Orator, that speaks Urdu with human-like realism. To the best of our knowledge, it is the first model that can speak Urdu in an authentic way: the way everyday people speak Urdu in everyday life.
We have voices from “the whole mohalla“ (entire neighborhood). From angry molvis and nosey aunties, to thailay walas, and burger boys.
You can experience the voices at UpliftAI.org. The website will allow you to experience dozens of voices in seconds by just clicking/hovering the circles.
Making the model
We have put our heart into this work, and enjoying it more than any work we have ever done in the past.
We started with the map of Pakistan and asked: how do people sound in every district. That gave us the accent list.
Then, we asked ourselves: what are all the unique characters we would interact with in a typical month as we go about our normal life. This gave us our character/persona list.
We then hired a leading voice producer and worked with him to hire voice actors and script writers, and then spent months doing voice recordings in studios in Karachi, Lahore, Quetta and Islamabad.
Once the voices were recorded, then we hired 50 contractors to label data.
The final and the simplest step was the model training. We trained a very small 32M param VITS2 inspired model.
The results are very encouraging.
- Here is someone creating a story book using our voices: https://shehryar-stories.up.railway.app/frontend/index.html
- Here is me just having fun with it: https://www.linkedin.com/posts/hammad2_is-your-mom-complaining-about-massi-leaving-ugcPost-7468570160233107457-sbQ5/
Ask
Hammad left Apple and took a year off to figure out what he wanted to do with the next decade of his career.
First month: finished netflix :)
Then he spent two months reviewing 20 years of voice tech research papers in chronological order… and somewhere in there realized that there is an opportunity to make amazing voice models for smaller languages — faster than big tech.
At the core, we are solving the problem of universal access to digital knowledge and digital services.
People who would previously never use certain digital services (like online ordering, or online health advice), actually use the service when they are given an intuitive Voice based interface — especially in emerging markets.
Voice Interfaces make it more natural and intuitive to use technology; and we believe it will onboard 100s of millions of new users onto digital services. This will make people more productive, and increase GDP per capita.
When we truly succeed, the world will access technology through our Voice Tech.
And the GDP per capita would increase around the world, as more people will be able to take full advantage of technology using intuitive voice interfaces.
Currently a large part of the world is unable to fully utilize technology coz text-interfaces are very confusing for many people.