WE ARE HIRING · 採用情報

Speech Edge AI Researcher

Tokyo & San Francisco

Kotoba builds cutting-edge voice AI. We’re hiring a Speech Edge AI Researcher to make our speech-to-speech, speech-to-text, and text-to-speech models fast, compact, and production-ready on smartphones, wearables, and other edge devices.

ABOUT KOTOBA

At the frontier of voice generation AI

Kotoba is a generative AI company on a mission to “become the defacto for voice AI in East Asia,” developing state-of-the-art voice generation AI models. At our core is a low-latency, high-accuracy speech translation model that connects conversations so naturally it feels as though both speakers share the same language. It supports major language pairs — including Japanese, English, Korean, Chinese, and Spanish — breaking down the language barrier. We also develop our own ultra-low-latency speech-to-text and text-to-speech models that run everywhere from the data center to edge devices, and we license this foundational technology, which powers the next generation of voice agents, to Fortune 50 companies and major U.S. tech firms. We create world-class voice AI from our two hubs in Tokyo and San Francisco.

Kotoba brings this technology to the world through its own product, the Kotoba app, available on iOS and Android. Since launch it has won explosive support across countries and regions including Japan and Korea, growing to a steady 2,000–3,000 new downloads per day. In the App Store and Google Play category rankings, it has reached No. 1 in its category, ahead of the likes of Google Translate and Audible. Loved by individual users, it is also seeing accelerating adoption among major enterprises — centered on Japan — and has supported nearly 100 events on the ground, including SusHi Tech Tokyo. It has drawn media attention as well, featured on numerous programs such as TV Asahi, NHK, TV Tokyo, Abema Prime, and PIVOT.

Kotoba was founded in 2023 by two Japanese generative AI researchers who earned their PhDs at top U.S. universities. To date it has raised a cumulative total of more than ¥3 billion (roughly US$23 million) from prominent VCs in Japan and the U.S. (including Kindred Ventures and Globis Capital Partners) and from the corporate venture arms of leading U.S. and Japanese enterprises, and it also receives strong government support in Japan for AI model training. World-class research paired with a product that keeps winning in the market — with both, Kotoba is one of the most exciting generative AI startups today.

Media

— TV Tokyo:https://www.youtube.com/watch?v=17CG6kSv2zs

— TV Asahi:https://tver.jp/episodes/epy340m62i

— PIVOT:https://www.youtube.com/watch?v=N30YYfIqGEg

Kotoba was founded in 2023 by two Japanese generative AI researchers who earned their PhDs at top U.S. universities. To date it has raised a cumulative total of more than ¥3 billion (roughly US$23 million) from prominent VCs in Japan and the U.S. (including Kindred Ventures and Globis Capital Partners) and from the corporate venture arms of leading U.S. and Japanese enterprises, and it also receives strong government support in Japan for AI model training. World-class research paired with a product that keeps winning in the market — with both, Kotoba is one of the most exciting generative AI startups today.

ROLE DESCRIPTION

THE ROLE

As a Speech Edge AI Researcher at Kotoba, you will develop the algorithms and model architectures that bring state-of-the-art voice AI onto resource-constrained devices. Your work will span speech-to-speech, automatic speech recognition (ASR), text-to-speech (TTS), speech translation, and neural audio codecs. You will use knowledge distillation, structured and unstructured pruning, quantization, low-rank methods, and architecture redesign to reduce model size, memory use, latency, and power consumption while preserving accuracy and naturalness. This is a research-and-engineering role across the full on-device stack. You will build and optimize streaming inference pipelines for mobile CPUs, GPUs, and NPUs; profile real devices; improve kernels, operators, memory movement, scheduling, and caching; and validate performance across smartphones, wearables, and embedded platforms. Working with model, mobile, and systems engineers, you will turn research prototypes into reliable implementations for Kotoba’s apps, SDKs, and customer products.

RESPONSIBILITIES

What you'll do

▪ Define and execute research on model compression and efficient architectures for on-device speech AI.

▪ Develop compact, streaming models for ASR, TTS, speech translation, speech-to-speech, and neural audio codecs.

▪ Advance knowledge distillation, structured and unstructured pruning, post-training and quantization-aware training, mixed precision, low-rank methods, and architecture redesign.

▪ Build and optimize end-to-end inference pipelines across CPUs, GPUs, and NPUs, including operators, kernel fusion, memory planning, caching, and execution scheduling.

▪ Benchmark models on real devices and establish reproducible metrics for quality, model size, real-time factor, end-to-end latency, time to first audio, peak memory, power, and thermal behavior.

▪ Integrate models with portable runtimes and hardware backends while maintaining performance, reliability, and consistent behavior across devices.

▪ Work with research, mobile, and systems engineers to productionize successful ideas and communicate findings through technical documentation, publications, or open-source releases when appropriate.

QUALIFICATIONS

Who we're looking for

Required

▪ A Ph.D., M.S., or equivalent research and engineering experience in machine learning, speech processing, efficient deep learning, computer systems, computer architecture, or a closely related field.

▪ A strong track record in model compression, efficient inference, or edge AI, demonstrated through research publications and/or production systems.

▪ Hands-on expertise in one or more of knowledge distillation, pruning, quantization, low-rank methods, efficient architecture design, or hardware-aware training.

▪ Strong implementation skills in PyTorch or JAX, together with C, C++, or comparable systems programming experience.

▪ Experience profiling and optimizing neural-network inference on resource-constrained hardware, with a practical understanding of CPUs, GPUs, NPUs, memory hierarchy, and numerical precision.

▪ The ability to design rigorous experiments and make informed trade-offs among quality, latency, memory, model size, power consumption, and portability.

▪ Strong written and verbal communication skills, including professional proficiency in English, and comfort working in a fast-moving research-driven startup.

Preferred

▪ Experience with ASR, TTS, speech translation, speech-to-speech systems, audio-language models, or neural audio codecs.

▪ Experience with on-device runtimes such as Core ML, ExecuTorch, ONNX Runtime, LiteRT/TFLite, TensorRT, or llama.cpp.

▪ Knowledge of acceleration backends and libraries such as Metal, Vulkan, CUDA, QNN, XNNPACK, or custom kernels and operator fusion.

▪ Experience with INT8 or INT4 quantization, mixed-precision execution, quantization-aware training, or hardware-aware model design.

▪ Experience building streaming, low-latency audio systems and deploying models across mobile, wearable, or embedded platforms.

▪ Experience training or distilling large teacher models using distributed GPU infrastructure, and evaluating compressed students at scale.

▪ Previous experience at an industrial research lab, AI company, technology company, or research-driven startup, including publications or open-source contributions.

HOW TO APPLY

If you’re interested, please apply Kotoba’s hiring team directly

at hiring@kotoba.tech with your resume attached.

Copyright © Kotoba Technologies 2026

English
English
English