China Voice Translator: RU/EN ↔ Chinese
China Voice Translator is a mobile web prototype for live spoken conversations during trips to China. The product is built for everyday situations where a user needs to talk with a Chinese-speaking person without repeatedly tapping the phone before every phrase: restaurants, shops, hotels, taxis, exhibitions, and street-level errands.
Product Scope
The MVP supports Russian, English, and Simplified Chinese. In automatic mode, Russian and English speech are translated into Chinese, while Chinese speech is translated back into Russian by default. Manual mode can lock fixed directions such as Russian to Chinese, Chinese to Russian, English to Chinese, or Chinese to English.
Conversation UX
The interface is designed around one continuous hands-free session. The user starts listening once, grants microphone access, and the app keeps listening, detecting speech boundaries, showing the original transcript, streaming the translation, finalizing the phrase, and optionally speaking the translated audio. New speech can interrupt playback, so a real dialog can continue naturally instead of waiting for the previous audio to finish.
Translation Quality
The translator is explicitly instructed to translate the current phrase only. It must not answer questions, add advice, explain the topic, or act as an assistant. The system preserves intent, negation, numbers, prices, names, dates, and politeness. Chinese output uses Simplified Chinese and Mandarin pronunciation. The product plan includes pinyin, back-translation, uncertainty warnings, and alternative text variants when recognition is unstable.
Realtime Architecture
The pilot uses a lightweight mobile PWA and Gemini Live through Vercel AI Gateway. The server keeps the permanent gateway key, issues short-lived realtime client tokens, normalizes language settings, rate-limits session starts, and records events. The browser captures microphone audio with echo cancellation, noise suppression, and automatic gain control, sends PCM audio frames, receives input and output transcripts, and plays translated audio.
Data And Observability
The prototype records normalized lifecycle events, final turns, source and translated audio chunks, detected languages, model/provider metadata, and latency metrics such as time to preliminary translation. This creates a reviewable dataset for improving language detection, measuring partial/final translation quality, and building regression cases before a wider release.
Engineering Focus
The hard parts are mobile latency, China network conditions, barge-in, unstable audio, language auto-detection, and keeping translation faithful without turning the model into a conversational agent. The technical design keeps provider details behind an adapter, plans a Hong Kong relay fallback if direct Vercel WebSocket access is unstable from mainland China, and leaves room for Alibaba/Qwen as a reserve realtime provider.
Product Value
The project turns a phone into a practical travel interpreter for China. It reduces dependence on human intermediaries, makes common everyday conversations faster, and gives the team measurable translation data instead of treating every field use as an unlogged one-off experiment.
Media Gallery
