Welcome to Ruthvik Nath Bandari's portfolio

All projects
Agentic AI · Speech

Language Mirror

AI Partner Catalyst hackathon · project

A voice-conversation prototype I orchestrated across hosted speech, translation and language models, offering guided speaking practice in five languages.

  • FastAPI
  • Gemini
  • ElevenLabs
  • Next.js
  • React 19
  • SSE
Abstract pair of mirrored speech bubbles, on a green fieldIllustration

Highlights

  • 15 dialect personas across 5 languages (Italian, Spanish, French, German, Japanese), voiced through 3 ElevenLabs voices
  • Google Speech-to-Text, Google Translate, and Gemini on a FastAPI backend: a six-stage SSE pipeline with a seven-model fallback chain
  • ElevenLabs Multilingual v2 text-to-speech for spoken replies
  • Next.js 16 / React 19 / MUI frontend with GSAP animations, live subtitles, and language-mismatch detection

How it was built

For the AI Partner Catalyst hackathon I orchestrated a voice tutor from hosted services. A learner's speech goes through a six-stage streaming pipeline on a FastAPI backend: Google speech-to-text, language detection through Google Translate, a Gemini reply in a fixed JSON format, and ElevenLabs speech synthesis. If the learner uses the wrong language, the pipeline stops before calling Gemini to save quota, and a seven-model chain handles quota failures. A Next.js frontend shows live subtitles. Limits: every model is hosted and nothing was trained; there are no tests or measurements.