
Emotive foundation voice models
At Miso Labs, our goal is to pass the voice Turing test. In order to do so, we must move beyond the current STT->LLM->TTS paradigm, and train full-duplex speech-to-speech models. This requires exploring new architectural ideas, and scaling them rapidly.
We’ve raised a large seed round from top investors and are hiring our founding research team. We’re a small team, so you can expect to work on every part of the model training stack from small-scale architecture experiments, to large pre-training runs on up to 1000 H100s. We are an in-person culture in San Francisco, and you will be expected to relocate to the SF area (we will cover the moving costs).
It is important to us that team members get credit for their work. Outside of our core IP, researchers are encouraged to publish their work at Miso Labs in conferences and on open source, and we are happy to sponsor travel to top conferences in AI/ML.
Preferred Qualifications:
If you’re interested in this role, please apply even if you don’t fit all the qualifications. We’re more interested in evidence of exceptional ability than prior experience in deep learning/AI.
This role comes with full benefits including:
Our hiring process consists of two rounds of interviews:
If you are not currently located in the Bay Area, we will cover the cost of your flight and hotel to attend the in-person interview at our office.
At Miso Labs, our goal is to pass the voice Turing test by training full-duplex speech-to-speech models. We've raised a large seed round from top investors, and are hiring our founding research team.