
Emotive foundation voice models
At Miso Labs, we're training full-duplex speech-to-speech models (i.e. models that can interrupt and talk over you). In order to do this, we need to collect large amounts of speech data (our current data set is over 100 million hours, or 11,000 years of audio), as well as human feedback data.
In order to do this, we have a 7000 square feet recording studio in North Hollywood, and a growing roster of voice actors recording from home studios. As we scale this data collection process, we are hiring for a founding data operations lead to manage the increasingly complex logistics of this process.
What you'll do
What we're looking for
You may be a great fit if
The interview consists of an initial zoom screening, and then a second onsite with the whole team.
At Miso Labs, our goal is to pass the voice Turing test by training full-duplex speech-to-speech models. We've raised a large seed round from top investors, and are hiring our founding research team.