Frequently Asked Questions

If you don’t see what you are looking for, please reach out to our team at support@canaryspeech.com.

How does Canary technology work?

Voice contains inherent qualities that can be used as vocal biomarkers to reveal emotional, physiological, and cognitive states. Canary Speech uses these biomarkers to create models that evaluate disease from conversational speech in seconds.

What parts of speech are used to create vocal models?

Using a proprietary system, speech features for the deep learning model are extracted from the audio signal, or the vocal sample, and include analysis of three elements of speech: acoustic, prosodic, and linguistic.

How is data acquired when creating vocal models?

We have built highly accurate predictive models through machine learning. Alongside our respected healthcare partners, we have collected data samples of patient voices to train our models. There are several types of data used to build a robust machine learning model:

  • First, in order to learn which biomarkers are indicative of a specific condition, the algorithms are given samples of patients already diagnosed with the condition and compare them against healthy controls.
  • Second, samples from a variety of people are used to incorporate diverse demographics. 
  • Third, the model incorporates data from a range of acoustic environments and speech input devices to control for the different settings that the technology may be used.
  • Finally, we use a technique called data augmentation to improve the performance and outcomes of machine learning models by changing audio characteristics to form new examples to train datasets.

What does the vocal stress score indicate?

Stress triggers measurable physiological changes — and many of them surface in your voice. Canary Speech detects acoustic markers like elevated pitch, irregular speech cadence, and vocal tension to assess how much autonomic arousal may be present when you speak.

 

Scores fall into three categories:

  • Low – Your vocal patterns reflect a calm, regulated state with effective stress response.
  • Medium – There are some stress indicators present, but nothing that appears to be significantly disrupting daily function. Simple interventions — movement, breathing, social support — can help prevent escalation.
  • High – Your voice is showing markers consistent with significant stress load. If this persists, it’s worth speaking with a doctor or counselor.

Everyone has stressful stretches. Tracking this score over time helps you spot when stress is becoming more than just a rough week.

What does the vocal mood score indicate?

Your Mood Score is based on acoustic features in your voice — things like intonation, pitch variability, and speech rhythm — that shift in measurable ways depending on your emotional state. Research has shown these patterns correlate with indicators used in clinical mood assessments like the PHQ-9. It’s not about what you say, but how you sound when you say it.

 

Scores fall into three categories:

  • Excellent – Your vocal patterns align with strong emotional wellbeing.
  • Good – Your voice reflects generally positive emotional health, with some mild fluctuation. Light wellness habits can help keep you on track.
  • Low – Your vocal indicators are consistent with more significant mood disruption. If this score persists, it’s worth having a conversation with someone you trust — or a healthcare professional.

One low score on a bad day isn’t something to worry about. A pattern of low scores over time is worth taking seriously.

Glossary of acoustic features

  • Mel-frequency cepstral coefficients (MFCC): Coefficients collectively make up an MFC, where MFC is a representation of the short-term power spectrum of sound
  • Perceptual Linear Predictive: An alternative to MFCC, a combination of spectral analysis and linear prediction analysis
  • Pitch: The fundamental period of the speech signal
  • Spectral flux: A measure of how quickly the power spectrum of a signal is changing
  • Spectral centroid: A measure where the center of mass of the spectrum is located
  • Spectral bandwidth: A bandwidth of signal spectrum
  • Spectral contrast: Decibel difference between spectral peaks and valleys
  • Spectral flatness: A measure of how much a sound resembles a pure tone
  • Spectral roll-off: The frequency below which a specified percentage of the total spectral energy
  • Harmonics-to-noise ratio (HNR): The ratio between periodic and non-periodic components of a speech sound
  • F0: Fundamental frequency of a speech signal, approximate frequency of the (quasi-)periodic structure of the voiced speech signal
  • Jitter: Variations in signal frequency
  • Shimmer: Variations in signal amplitude
  • WER: Word error rate. Usually, it is a measure to check the ASR (automatic speech recognition) accuracy. We use this rate for measuring the articulation of speech compared to the reading script.
  • Word_prob: A probability of a spoken word’s appearance in a big corpus. Common words such as happy, thank, etc. will get high probability and uncommon words such as extraordinary, canary, etc. will get low probability.
  • Filler ratio: Ratio of filler word usage (hmm, uh, oh, eh, …)
  • SYN: Syntactic part of speech ratio. for ex. ADJ (ratio of adjective word usage).
  • Lexical difficulty: Smog grade, age of acquisition of words, concreteness, ambiguity, familiarity
  • OOV: Out of vocabulary (unrecognizable word from ASR)

What does the vocal energy score indicate?

Your Energy Score measures vocal indicators of vitality and engagement, drawn from three acoustic components:

 

  • Dynamics – The degree of pitch variation in your speech. A wider range reflects expressive, engaged communication; a flatter pattern can indicate low affect or fatigue.
  • Speed – Your rate of speech in words per minute. Both significantly elevated and reduced speeds can be clinically relevant markers of mental and physical state.
  • Power – The strength and respiratory effort behind your voice. Lower power scores reflect reduced vocal projection, which can correlate with disengagement or low energy.

A single low score may simply reflect a tired day — that’s normal. What’s more meaningful is how your score trends over weeks and months.

What is a biomarker, and more specifically, a vocal biomarker?

The FDA defines a digital biomarker to be a characteristic or set of characteristics, collected from digital health technologies, that is measured as an indicator of normal biological processes, pathogenic processes, or responses to an exposure or intervention. A vocal biomarker is a feature, or a combination of features from the audio signal of the voice that is associated with a clinical outcome.

What technologies were used to create Canary Speech?

Machine learning, a form of artificial intelligence, is at the core of Canary Speech’s technology.  We use state-of-the-art machine learning methods, including deep neural networks, to automatically learn to make predictions based on features extracted from labeled data.

What is the ideal audio recording size for analysis?

An audio length between 20-40 seconds is ideal.

How large are the data sets?

Data sets by disease range in the hundreds to tens of thousands and differ based on the quantity required for science based machine learning techniques.

Is my information secure?

Canary is backed by a team of experts working tirelessly to maintain data security–from our development lifecycle, to continuous monitoring of our infrastructure and applications. Our technology is HIPAA-compliant and offers fully anonymous solutions. Our technology is CIS Control Audited, vulnerability scanned, and has cleared all API penetration testing. Canary Speech’s technology and services are HITRUST e1 certified, and ISO/IEC 27001:2022 ISO/IEC 42001:2023 certified, demonstrating our continued commitment to data security, privacy, and responsible AI governance.