My Projects

Building innovative AI solutions that push the boundaries of what's possible

Speech AI

DIS-Vector

Zero-shot voice cloning framework via advanced speaker embedding extraction from limited audio samples.

PyTorchWhisperXTTS-v2HuBERT+2
accuracy: 97.3%type: Research Framework
Speech AI

EMOD: Emotional Speech Synthesis

Standalone emotion embedding extraction framework for neural TTS with continuous intensity control and zero-shot emotion transfer.

PyTorchVITSHuBERTPyWorld+2
stars: 18forks: 3languages: 9
LLM & Agents

Real-time Conversational AI Agent

End-to-end conversational AI with streaming ASR, LLM reasoning, and neural TTS for real-time natural conversations.

PythonFastAPIWebSocketsLLaMA-3+3
latency: <500msrequests: 10K+
RAG & Search

Multimodal Vision RAG

CLIP-powered multimodal embedding pipeline with LLaVA reasoning for enterprise document intelligence.

CLIPLLaVAQdrantYOLO+1
vectors: 2.1Maccuracy: 96%
RAG & Search

Enterprise RAG Platform

Hybrid search RAG with knowledge graph traversal for 40% improved answer relevance over naive RAG.

LangChainWeaviateBGE-M3GPT-4o+2
relevance: +40%documents: 50K+
Multimodal AI

Real-Time Object Tracker

ONNX-quantized YOLOv9 + ByteTrack deployed on edge hardware at 30fps for retail analytics.

YOLOv9ONNXByteTrackTensorRT+2
fps: 30stores: 50+
Speech AI

Multilingual TTS Engine

Production TTS engine supporting 12 languages with sub-200ms first-token latency and custom prosody model.

StyleTTS2BarkCoquiTriton+2
latency: <200mslanguages: 12
Speech AI

Voice Cloning Studio

Professional voice cloning and synthesis platform supporting 20+ languages with emotion control.

ReactFastAPIXTTS-v2VITS+2
voices: 50+languages: 20
ML Core

ML Optimization Suite

Advanced optimization algorithms with convergence analysis and interactive visualizations.

PythonNumPyMatplotlibJupyter+1
stars: 45forks: 23