Publications & Research

Research papers, technical reports, and open-source contributions

DIS-Vector: A Framework for Voice Intelligence and Speaker Embedding Extraction

A comprehensive framework for extracting speaker embeddings, emotional features, and prosodic patterns from limited audio samples, enabling zero-shot voice conversion and few-shot speaker adaptation.

EMOD: An Efficient Approach for Low-Resource Controllable Emotional Speech Synthesis

Research Project 2025Demo: nn-project-2.github.io → Emotion-TTS-web/ 18 3Python

A standalone emotion embedding extraction framework for neural TTS systems that converts emotional speech signals into fixed-dimensional latent embeddings. Supports categorical emotion representation and continuous intensity control through a scalar parameter α for zero-shot emotional conditioning.

Code🎤 Zero-Shot Emotion Transfer

Real-time Conversational AI: Integrating Streaming ASR, LLM Reasoning, and Neural TTS

Technical Report 2024

Architecture and implementation of end-to-end conversational AI systems with multilingual support, context persistence, and emotion-aware response generation.

Advanced RAG Systems: Hybrid Search and Agentic Reasoning

Technical Report 2024

Novel approaches to retrieval-augmented generation combining dense and sparse retrieval with multi-document reasoning for enhanced accuracy.