VoxPilot AI — Multilingual Autonomous Conversational Voice Copilot
Link to open source: https://github.com/SHIa12l0000000/VoxPilot-AI
Link to Live Project: https://voxpilot-ai.surge.sh/
🌟 Overview & Problem Statement
Traditional AI chat interfaces demand continuous screen engagement, manual typing, and high reading bandwidth. In fast-paced learning and mobile environments, learners and developers need a hands-free, low-latency, and personalized voice assistant that can converse naturally, switch languages fluidly (English, Hindi, and Hinglish), and never require complex server setups.
VoxPilot AI is an autonomous, real-time conversational voice copilot designed for Android and modern web platforms. Built around an intelligent dual-engine architecture, it delivers contextual tutoring, interactive Q&A, and live translation with sub-second voice interactions and acoustic echo isolation.
🚀 Key Features
1. Autonomous Dual-Engine Voice Architecture:
-> Cloud Streaming (Agora RTC + MiniMax / Deepgram): Low-latency conversational audio pipeline for cloud voice streaming.
-> On-Device Autonomous Fallback: Runs locally on every Android device and browser without requiring an external server or fixed IP.
2. Multilingual Intelligence (Powered by Gemini):
->Dynamically matches the user's spoken language: speak in English for fluent English explanations, or speak in Hindi / Hinglish for natural native tutoring.
3 specialized modes:
🎓 Learn Mode: Interactive Socratic tutor explaining core concepts with real-world analogies.
💬 Ask Mode: Real-time conversational knowledge assistant across engineering, science, and math.
🌐 Translate Mode: Live bidirectional English-Hindi voice translation.
3. Acoustic Echo Isolation & Tap-To-Interrupt:
-> Intelligent audio gating prevents microphone feedback loops from speaker output.
-> Tap-to-interrupt allows instant barge-in whenever the user wants to pivot or ask a follow-up.
4. Contextual Smart Notes:
-> Say "Save this as a note" or "Notes banao", and key takeaways are automatically extracted and stored locally.
🛠️ Tech Stack & Architecture
-> Mobile (Android): Kotlin, Jetpack Compose, Material 3, Android Text-To-Speech (TTS), Android SpeechRecognizer, Room DB / SharedPreferences.
-> AI Core: Google Gemini REST API (`gemini-3.1-flash-lite` with fallback to `gemini-3.5-flash`) for real-time natural language reasoning.
-> Voice & Web Client: Web Speech Recognition API, Web SpeechSynthesis API, HTML5 / CSS Glassmorphism Canvas Orb visualizer.
->Communication Pipelines:*Agora RTC/RTM engine support with seamless offline autonomous fallback.
💡 Impact & Vision
->VoxPilot AI bridges the digital divide for millions of multilingual students and developers by making personalized, interactive AI learning as natural and effortless as talking to a friend.
This build was uploaded as a hackathon project





