Publications & Research
Research papers, technical reports, and open-source contributions
DIS-Vector: A Framework for Voice Intelligence and Speaker Embedding Extraction
A comprehensive framework for extracting speaker embeddings, emotional features, and prosodic patterns from limited audio samples, enabling zero-shot voice conversion and few-shot speaker adaptation.
EMOD: An Efficient Approach for Low-Resource Controllable Emotional Speech Synthesis
A standalone emotion embedding extraction framework for neural TTS systems that converts emotional speech signals into fixed-dimensional latent embeddings. Supports categorical emotion representation and continuous intensity control through a scalar parameter α for zero-shot emotional conditioning.
Real-time Conversational AI: Integrating Streaming ASR, LLM Reasoning, and Neural TTS
Architecture and implementation of end-to-end conversational AI systems with multilingual support, context persistence, and emotion-aware response generation.
Advanced RAG Systems: Hybrid Search and Agentic Reasoning
Novel approaches to retrieval-augmented generation combining dense and sparse retrieval with multi-document reasoning for enhanced accuracy.