WE ARE HIRING · 採用情報

AI Researcher

San Francisco

Kotoba builds cutting-edge voice generation AI. We’re hiring an AI Researcher to advance our real-time speech systems across full-duplex speech-to-speech, speech-to-text, and text-to-speech, and to turn research breakthroughs into products used around the world.

ABOUT KOTOBA

⾳声⽣成 AI の最前線

Kotoba は、“東アジアの⾳声 AI のスタンダードになる”ことをミッションに掲げ、最先端の⾳声⽣ 成 AI モデルを開発する⽣成 AI 企業です。中核となるのは、まるで同じ⾔語を話しているかのよう な⾃然さで会話を繋ぐ、低遅延・⾼精度な⾳声翻訳モデル。⽇本語・英語・韓国語・中国語・スペ イン語をはじめとする主要⾔語ペアに対応し⾔語の壁を崩しています。さらに、データセンターか らエッジデバイスまで動作する超低遅延の speech-to-text/text-to-speech モデルも⾃社開発して おり、次世代の⾳声エージェントを⽀える基盤技術として、Fortune 50 企業や⽶国のビッグテック 企業へのライセンス提供を多数⼿がけています。世界トップ⽔準の⾳声 AI を、私たちは東京・サン フランシスコの2拠点から⽣み出しています。

この技術を、Kotoba は⾃社プロダクト「Kotoba app」(⽇本語名: 同時通訳)として iOS / Android で世に送り出しています。リリース以来、⽇本・韓国をはじめとする国・地域で爆発的な⽀持を獲 得し、1 ⽇あたり約 2,000〜3,000 件の新規ダウンロードをコンスタントに維持するまで普及して います。App Store および Google Play のカテゴリーランキングでは、Google 翻訳や Audible など を抑えてカテゴリー内 1 位を獲得した実績もあります。個⼈ユーザーに愛⽤されるだけでなく、⽇ 本を中⼼に⼤⼿エンタープライズへの導⼊も加速しており、SushiTech Tokyo をはじめ 100 近くの イベントの現場を⽀えてきました。メディアからも注⽬を集め、テレビ朝⽇、NHK、テレビ東京、 Abema Prime、PIVOT など数多くの番組で特集されています。

メディア掲載
— テレビ東京:https://www.youtube.com/watch?v=17CG6kSv2zs
— テレビ朝⽇:https://tver.jp/episodes/epy340m62i
— PIVOT:https://www.youtube.com/watch?v=N30YYfIqGEg

Kotoba は 2023 年、⽶国トップ⼤学で PhD を取得した 2 名の⽇本⼈の⽣成 AI 研究者によって創業されました。これまでに⽇⽶の著名 VC (Kindred Ventures, Globis Capital Partners など)や⽶国・⽇本の最⼤⼿企業の CVC から累計約 30 億円以上を調達し、さらに⽇本政府からも AI モデル学習に関する⼒強い⽀援を受けています。世界最⾼峰の研究⼒と、実際に市場で勝ち続けるプロダクト⼒。その両輪を併せ持つ、いまもっとも勢いのある⽣成 AI スタートアップのひとつです。

ROLE DESCRIPTION

役割

As an AI Researcher at Kotoba, you will advance the next generation of real-time, interactive voice AI across full-duplex speech-to-speech, speech-to-text, and text-to-speech systems. You will build models that not only understand and generate high-quality speech, but also manage the flow of conversation naturally—including turn-taking, interruptions, overlapping speech, backchannels, response timing, prosody, and latency. Your research may also extend to AI model orchestration: coordinating speech models, language models, reasoning systems, retrieval components, and tools behind a responsive voice interface. You will work across the full research lifecycle, from identifying research questions and designing experiments to distributed training, evaluation, and production deployment. Given Kotoba’s focus, research involving Japanese, Korean, Chinese, and other East Asian languages will be especially relevant.

RESPONSIBILITIES

担当いただくこと

  • Define and execute research projects for next-generation voice AI across speech-to-speech, speech-to-text, and text-to-speech systems.

  • Develop full-duplex conversational models that can listen and speak simultaneously while handling turn-taking, interruptions, overlapping speech, backchannels, and end-of-turn prediction.

  • Improve the accuracy, naturalness, expressiveness, multilingual robustness, and streaming latency ofspeech recognition and speech generation models.

  • Conduct multilingual and cross-lingual research, particularly for Japanese, Korean, Chinese, English,and other languages important to Kotoba’s products.

  • Explore architectures that orchestrate speech, language, reasoning, retrieval, and tool-use modelsbehind a unified real-time voice interface.

  • Build and scale model training and inference pipelines using distributed GPU infrastructure, while optimizing models for low-latency deployment.

  • Collaborate with research, product, and infrastructure engineers to transfer promising research into Kotoba’s applications, APIs, SDKs, and customer projects.

QUALIFICATIONS

Who we're looking for

Required

  • Ph.D. or equivalent research experience in machine learning, speech processing, natural language processing, multimodal AI, human-computer interaction, or a closely related field.

  • A strong research track record demonstrated through publications at leading conferences or journals in machine learning, speech, NLP, or related areas.

  • Deep expertise in at least one relevant area, such as speech-to-speech modeling, speech translation, spoken dialogue systems, speech recognition, speech generation, multimodal

  • foundation models, large language models, or AI model orchestration.

  • Hands-on experience designing, implementing, training, and evaluating modern neural models using frameworks such as PyTorch or JAX.

  • Strong knowledge of modern speech and language architectures, including transformers, streaming models, autoregressive and non-autoregressive models, and foundation-model training.

  • The ability to formulate original research questions, design rigorous experiments, critically analyze results, and turn promising ideas into working systems.

  • Familiarity with large-scale model training, inference, data pipelines, distributed computing, and GPU-based experimentation.

  • Strong written and verbal communication skills, including professional proficiency in English.

  • Comfort working in an ambitious, fast-moving environment where research is closely connected to products and real-world deployment.

Preferred

  • Research experience in full-duplex speech-to-speech, speech recognition, or speech generation, particularly involving turn-taking, interruptions, backchannels, dialogue timing, or conversational fluency.

  • Research experience involving Japanese, Korean, Chinese, or other East Asian languages, including multilingual or cross-lingual modeling.

  • Knowledge of audio tokenization, neural audio codecs, streaming speech recognition, streaming speech generation, or low-latency speech architectures.

  • Experience with distributed training and efficient inference for large speech, language, or multimodal models.

  • Research experience with systems that orchestrate multiple models, agents, retrieval components, reasoning modules, or external tools.

  • Previous experience at an industrial research lab, major AI organization, technology company, or research-driven startup, particularly transferring research into production.

  • A record of open-source contributions

HOW TO APPLY

ご関⼼をお持ちの⽅は、Kotoba の採⽤窓⼝(hiring@kotoba.tech)まで、

resume を添付して直接ご応募ください。

Copyright © Kotoba Technologies 2026

Japanese
Japanese
Japanese