WE ARE HIRING · 採用情報
AI Researcher
San Francisco
Kotoba builds cutting-edge voice generation AI. We’re hiring an AI Researcher to advance our real-time speech systems across full-duplex speech-to-speech, speech-to-text, and text-to-speech, and to turn research breakthroughs into products used around the world.
ABOUT KOTOBA
At the frontier of voice generation AI
Kotoba is a generative AI company on a mission to “become the defacto for voice AI in East Asia,” developing state-of-the-art voice generation AI models. At our core is a low-latency, high-accuracy speech translation model that connects conversations so naturally it feels as though both speakers share the same language. It supports major language pairs — including Japanese, English, Korean, Chinese, and Spanish — breaking down the language barrier. We also develop our own ultra-low-latency speech-to-text and text-to-speech models that run everywhere from the data center to edge devices, and we license this foundational technology, which powers the next generation of voice agents, to Fortune 50 companies and major U.S. tech firms. We create world-class voice AI from our two hubs in Tokyo and San Francisco.
Kotoba brings this technology to the world through its own product, the Kotoba app, available on iOS and Android. Since launch it has won explosive support across countries and regions including Japan and Korea, growing to a steady 2,000–3,000 new downloads per day. In the App Store and Google Play category rankings, it has reached No. 1 in its category, ahead of the likes of Google Translate and Audible. Loved by individual users, it is also seeing accelerating adoption among major enterprises — centered on Japan — and has supported nearly 100 events on the ground, including SusHi Tech Tokyo. It has drawn media attention as well, featured on numerous programs such as TV Asahi, NHK, TV Tokyo, Abema Prime, and PIVOT.
Kotoba was founded in 2023 by two Japanese generative AI researchers who earned their PhDs at top U.S. universities. To date it has raised a cumulative total of more than ¥3 billion (roughly US$23 million) from prominent VCs in Japan and the U.S. (including Kindred Ventures and Globis Capital Partners) and from the corporate venture arms of leading U.S. and Japanese enterprises, and it also receives strong government support in Japan for AI model training. World-class research paired with a product that keeps winning in the market — with both, Kotoba is one of the most exciting generative AI startups today.
Media
— TV Tokyo:https://www.youtube.com/watch?v=17CG6kSv2zs
— TV Asahi:https://tver.jp/episodes/epy340m62i
— PIVOT:https://www.youtube.com/watch?v=N30YYfIqGEg
Kotoba was founded in 2023 by two Japanese generative AI researchers who earned their PhDs at top U.S. universities. To date it has raised a cumulative total of more than ¥3 billion (roughly US$23 million) from prominent VCs in Japan and the U.S. (including Kindred Ventures and Globis Capital Partners) and from the corporate venture arms of leading U.S. and Japanese enterprises, and it also receives strong government support in Japan for AI model training. World-class research paired with a product that keeps winning in the market — with both, Kotoba is one of the most exciting generative AI startups today.
ROLE DESCRIPTION
THE ROLE
As an AI Researcher at Kotoba, you will advance the next generation of real-time, interactive voice AI across full-duplex speech-to-speech, speech-to-text, and text-to-speech systems. You will build models that not only understand and generate high-quality speech, but also manage the flow of conversation naturally—including turn-taking, interruptions, overlapping speech, backchannels, response timing, prosody, and latency. Your research may also extend to AI model orchestration: coordinating speech models, language models, reasoning systems, retrieval components, and tools behind a responsive voice interface. You will work across the full research lifecycle, from identifying research questions and designing experiments to distributed training, evaluation, and production deployment. Given Kotoba’s focus, research involving Japanese, Korean, Chinese, and other East Asian languages will be especially relevant.
RESPONSIBILITIES
What you'll do
Define and execute research projects for next-generation voice AI across speech-to-speech, speech-to-text, and text-to-speech systems.
Develop full-duplex conversational models that can listen and speak simultaneously while handling turn-taking, interruptions, overlapping speech, backchannels, and end-of-turn prediction.
Improve the accuracy, naturalness, expressiveness, multilingual robustness, and streaming latency ofspeech recognition and speech generation models.
Conduct multilingual and cross-lingual research, particularly for Japanese, Korean, Chinese, English,and other languages important to Kotoba’s products.
Explore architectures that orchestrate speech, language, reasoning, retrieval, and tool-use modelsbehind a unified real-time voice interface.
Build and scale model training and inference pipelines using distributed GPU infrastructure, while optimizing models for low-latency deployment.
Collaborate with research, product, and infrastructure engineers to transfer promising research into Kotoba’s applications, APIs, SDKs, and customer projects.
QUALIFICATIONS
Who we're looking for
Required
Ph.D. or equivalent research experience in machine learning, speech processing, natural language processing, multimodal AI, human-computer interaction, or a closely related field.
A strong research track record demonstrated through publications at leading conferences or journals in machine learning, speech, NLP, or related areas.
Deep expertise in at least one relevant area, such as speech-to-speech modeling, speech translation, spoken dialogue systems, speech recognition, speech generation, multimodal
foundation models, large language models, or AI model orchestration.
Hands-on experience designing, implementing, training, and evaluating modern neural models using frameworks such as PyTorch or JAX.
Strong knowledge of modern speech and language architectures, including transformers, streaming models, autoregressive and non-autoregressive models, and foundation-model training.
The ability to formulate original research questions, design rigorous experiments, critically analyze results, and turn promising ideas into working systems.
Familiarity with large-scale model training, inference, data pipelines, distributed computing, and GPU-based experimentation.
Strong written and verbal communication skills, including professional proficiency in English.
Comfort working in an ambitious, fast-moving environment where research is closely connected to products and real-world deployment.
Preferred
Research experience in full-duplex speech-to-speech, speech recognition, or speech generation, particularly involving turn-taking, interruptions, backchannels, dialogue timing, or conversational fluency.
Research experience involving Japanese, Korean, Chinese, or other East Asian languages, including multilingual or cross-lingual modeling.
Knowledge of audio tokenization, neural audio codecs, streaming speech recognition, streaming speech generation, or low-latency speech architectures.
Experience with distributed training and efficient inference for large speech, language, or multimodal models.
Research experience with systems that orchestrate multiple models, agents, retrieval components, reasoning modules, or external tools.
Previous experience at an industrial research lab, major AI organization, technology company, or research-driven startup, particularly transferring research into production.
A record of open-source contributions
HOW TO APPLY
If you’re interested, please apply Kotoba’s hiring team directly
at hiring@kotoba.tech with your resume attached.