Engineering Case Study

Dhvani

Real-Time Audio Analysis & Speech Emotion Recognition Platform

Interactive Visual Telemetry
Dhvani

Project Overview

Dhvani is a high-performance web audio processing application that performs real-time acoustic signal spectrum visualization, speech transcription, and sentiment classification.

The Problem & Motivation

Traditional speech analysis required heavy desktop software, lacking web-accessible real-time FFT spectrum visualizations and instant sentiment feedback.

System Architecture Diagram

CLIENT UIReact & Web Audio API
➔
NEXT.JS ROUTERNext.js Client Application
➔
BACKEND SERVICEPython FastAPI & PyTorch
➔
DATABASE & CLOUDMongoDB & AWS S3

Key Engineering Features

01

60 FPS Spectrum Analyzer

Low-latency Web Audio API FFT visualizer.

Renders frequency spectrum bars and waveform oscillograms smoothly without UI thread lag.

02

Speech Emotion Classification

Machine learning acoustic tone detector.

Classifies vocal emotions (Joy, Calm, Surprise, Anger) with high precision confidence metrics.

Full Technology Stack

Frontend
ReactNext.jsWeb Audio APIHTML5 Canvas
Backend & ML
PythonFastAPIPyTorchLibrosa
Deployment
DockerAWS EC2Nginx

Technical Challenges & Engineering Solutions

Canvas 60 FPS Frame Rate Stability

Challenge: Continuous Web Audio API buffer sampling was triggering excessive React re-renders.

Solution: Decoupled canvas drawing loop into requestAnimationFrame ref handles, bypassing React component state updates entirely.

Measurable Results & Outcomes

115msInference LatencyReal-time
95%Model AccuracyTested benchmark
60 FPSRender FPSZero frame drops

What I Learned

  • Deep understanding of Web Audio API, Fast Fourier Transforms (FFT), and raw binary sound buffers.
  • Gained hands-on experience deploying PyTorch models inside containerized Python microservices.