Skip links

Voice Lab

A lab specialized in speech processing, transforming spoken interactions and multilingual communications, into actionable information that enables intelligent automation, customer engagement, and data analytics.

Where Speech BecomesIntelligence

The Voice Lab builds on the technological heritage of PerVoice, founded in 2007 as a spin-off of Fondazione Bruno Kessler and specialized in Automatic Speech Recognition.

It develops advanced technologies for visual perception and multimodal AI, helping organizations extract value from visual data and transform it into scalable operational knowledge. Combining computer vision, image processing, document understanding, and generative AI, the lab delivers reliable solutions that support inspection, monitoring, automation, and information access across complex enterprise environments.

Distinctive Approach

Reliability

Enterprise-grade speech technologies designed for traceable, auditable, and measurable performance across mission-critical communication and analytics workflows.

Operations Ready

Solutions validated on real-world deployments including broadcast monitoring, institutional meetings, customer care, and large-scale conversational analytics.

Flexibility

Modular architectures that integrate with multiple ASR, speaker analysis, and language technologies while adapting to diverse domains and deployment scenarios.

Focus Areas

Speech Recognition & Multilingual Transcription
Development of advanced automatic speech recognition solutions, domain-adapted transcription systems, and multilingual processing technologies for meetings, broadcast media, institutional environments, and customer interactions.
01
Multimodal Speech Foundation Models
VelvetSpeech is a multimodal speech-language model specialized in automatic speech recognition, speech translation, speech understanding, spoken content summarization, listening comprehension, and speech emotion recognition. It is designed to perform these tasks consistently and reliably across a wide range of prompts, acoustic environments and speaking styles.
02
Speaker Analytics & Language Intelligence
Research and deployment of speaker diarization, speaker identification, and spoken language identification technologies that enable accurate transcriptions, searches, and analytics across multilingual and multi-speaker communications.
03
Speech Emotion Recognition
Development of AI systems capable of identifying emotions and interaction dynamics from speech, supporting customer experience analysis, service quality monitoring, and advanced conversational intelligence applications.
04