← Back to Showcase
🧠
AI Project

VoiceVerse — Multilingual AI Translator

👤 by Tehzeeb Shah 📍 Karachi 📅 Jul 22, 2026
❤️ 0
Votes
85
🤖 AI Score
2
💬 Reviews
217
👁 Views

📦 Deliverables

🌐
Live Demo
https://voice.s4s.tehzeeb.to.frosty-c…
↗
🐙
Source Code
https://github.com/tehzeebshahs4s-hub…
↗

🖼️ Screenshots

screenshot

🎯 Problem Statement

When people travel to another country, the language barrier prevents even basic communication — asking for a hospital, reading a street sign, understanding a menu, or making sense of a voice note. Existing tools are fragmented: one app for speech, another for images, another for documents. There's no single lightweight tool that translates any input — voice, image, audio, video, or document — into any language and speaks it back. What I Built A universal, browser-based AI translator deployed live on a Plesk server. It translates between 35+ languages (English, Urdu, Arabic, Chinese, Japanese, Hindi, French, Spanish…) through four input methods: 🎤 Live speech — push-to-talk mic with 10s auto-stop, transcribed by Groq Whisper. ⌨️ Typed text — instant translation. 📎 File upload — handles any file: Images → in-browser OCR (Tesseract.js) reads signs/menus/documents. Audio/Video → decoded and auto-chunked, transcribed by Groq Whisper. PDF/Word/Text → text extracted in-browser (pdf.js / mammoth.js). 🔊 Spoken output — every translation is spoken aloud, with cloud TTS for languages browsers can't voice natively (Urdu, Arabic, Chinese…). Architecture is fully client-side extraction + thin PHP backend for translation/TTS/STT proxies — no heavy server dependencies. Translation runs through Google Translate (~0.6s) with an LLM fallback and a persistent cache for instant repeats. Includes a live Diagnostics panel for transparency.

🛠️ What I Built

A universal, browser-based AI translator deployed live on a Plesk server. It translates between 35+ languages (English, Urdu, Arabic, Chinese, Japanese, Hindi, French, Spanish…) through four input methods: 🎤 Live speech — push-to-talk mic with 10s auto-stop, transcribed by Groq Whisper. ⌨️ Typed text — instant translation. 📎 File upload — handles any file: Images → in-browser OCR (Tesseract.js) reads signs/menus/documents. Audio/Video → decoded and auto-chunked, transcribed by Groq Whisper. PDF/Word/Text → text extracted in-browser (pdf.js / mammoth.js). 🔊 Spoken output — every translation is spoken aloud, with cloud TTS for languages browsers can't voice natively (Urdu, Arabic, Chinese…). Architecture is fully client-side extraction + thin PHP backend for translation/TTS/STT proxies — no heavy server dependencies. Translation runs through Google Translate (~0.6s) with an LLM fallback and a persistent cache for instant repeats. Includes a live Diagnostics panel for transparency.

🔥 Challenges I Faced

No FTP credentials on Plesk — reverse-engineered Plesk's internal File-Manager REST API (session login + forgery token + domain-scoped upload endpoints) to deploy entirely via the panel. Web Speech API unsupported on mobile/Firefox/Safari — built a multi-engine fallback: native → Groq cloud → in-browser Whisper (transformers.js). No Urdu/Arabic TTS voice in any browser — added a server-side Google-TTS proxy so 100+ languages are spoken. LLM rate limits (429) made translation slow/unreliable — switched primary engine to Google Translate with on-disk caching (3–5× faster). Audio sample-rate bug — browser decoded at 48kHz but I labeled WAV as 16kHz → garbled transcription; fixed to use the real decoded rate with byte-safe chunking.

💡 What I Learned

Client-side ML libraries (Tesseract.js, transformers.js, pdf.js, mammoth.js) make powerful, key-less features possible entirely in the browser. A graceful fallback chain (native → cloud → WASM) is the key to an app that works on every device. Hosted control panels like Plesk expose scriptable internal APIs if you trace the session/forgery-token flow. Caching + choosing the right primary engine (Google vs LLM) matters more than micro-optimization for perceived speed.

🚀 Future Improvements

Conversation mode — auto-detect each speaker's language for back-and-forth dialogue. Segmented subtitles for long audio/video (timestamped, line-by-line). Scanned-PDF OCR — render pages to images and OCR them. Offline mode and a PWA / mobile wrapper for travel use. Conversation history export and a saved phrasebook.

🧰 Tech Stack & AI Tools

JavaScript (vanilla)PHPHTML/CSSGroq Whisper APIGoogle TranslateOpenAI-compatible LLM (jugaar.ai)Tesseract.jspdf.jsmammoth.jstransformers.jsPlesk
🤖 AI Tools Used:
opencode z.ai GLM 5.2

🤖 AI Reviewer Feedback

Great job building a comprehensive, browser‑based translator that handles speech, text, images, audio/video, and documents with clever client‑side processing and fallback chains. To push it further, polish the mobile experience and add more robust error handling/documentation for the various APIs you integrate.

💬 Leave a Quick Review

Log in to rate and review this project.

⭐ Community Reviews (2)

T
Tehzeeb Shah 🤖 AI
★★★★★
Sep 3, 2026

Great job building a comprehensive, browser‑based translator that handles speech, text, images, audio/video, and documents with clever client‑side processing and fallback chains. To push it further, polish the mobile experience and add more robust error handling/documentation for the various APIs you integrate.

T
Tehzeeb Shah 🤖 AI
★★★★★
Jul 22, 2026

Great job pulling together a wide range of client‑side ML tools and APIs into a smooth, multilingual translator that works across many input types. To push this further, consider adding clearer UI cues for error states and offline fallback, and document the API keys and deployment steps so others can reproduce the setup more easily.