Free tool · No signup · Audio stays on your device

Free AI Voice Analyzer: Check Your Pitch Range, Volume and Pauses

An AI voice analyzer measures how you sound rather than what you say. This free tool listens through your microphone for up to 60 seconds and reports three things: your pitch, as a median and a range in semitones with a monotone flag; your loudness consistency, meaning how much your volume moves between phrases and how often it drops away; and your pauses, counting every silence longer than 250 milliseconds with its average and longest length and the share of time you spend speaking. As rules of thumb, a pitch range under about 4 semitones often sounds flat, phrase-to-phrase volume that varies by 2 to 6 decibels sounds natural, and most confident presenters spend 70 to 90% of their time talking. It runs in your browser with no signup or upload, so your audio never leaves your device. It measures sound, not emotion.

Runs entirely in your browser. Your audio is never recorded to a file, uploaded or sent to any server.

Press Start and talk for 30 to 60 seconds as if you were presenting: explain your job, pitch an idea or read a paragraph aloud. Results appear when you press Stop or at 60 seconds.

How the analyzer measures your voice

Frames. Every 50 milliseconds the tool reads the last 2,048 samples from your microphone (about 43 ms of sound) through a Web Audio AnalyserNode. Each frame becomes three numbers: its loudness in dBFS, a pitch estimate, and whether it counts as speech. The raw sound is then discarded.

Speech or silence. The threshold adapts to your room: a frame counts as speech when it is about 12 dB louder than the quietest tenth of the recording (your background noise), capped so that a recording with no silence still works.

Pitch.For voiced speech frames, the fundamental frequency (F0) is found by autocorrelation: the frame is compared with delayed copies of itself, and the delay that matches best is one vocal-fold cycle. The search covers 70 to 400 Hz. Quiet frames and frames without a clear repeating pattern (breaths, “s” and “f” sounds) are skipped. Your range is the distance in semitones between the 10th and 90th percentile of those estimates, which ignores the odd stray reading at either end.

Loudness.Speech frames are grouped into half-second phrases so that syllable-to-syllable flicker is smoothed out. Variability is the standard deviation of those phrase levels in decibels; a phrase is “too quiet” when it is more than 8 dB below your median phrase.

Pauses. A pause is any silence of 250 ms or more between two stretches of speech; silence before you start and after you finish is ignored. Speaking time is the share of the span from your first word to your last that is not inside a pause.

How to read your results

MeasureRule-of-thumb bandsWhat to do
Pitch range (10th-90th percentile)Under 4 st: flat. 4-8 st: some variety. 8-12 st: lively. Over 12 st: very wide.If flat, stress one key word per sentence and let pitch fall at full stops. If very wide, check for statements rising at the end.
Monotone flagShown when the range is under 4 semitones.Read a paragraph aloud as if to a child, then again as normal. The second take usually gains a few semitones.
Loudness variability (0.5 s phrases)Under 2 dB: very even. 2-6 dB: healthy. Over 6 dB: uneven.If uneven, find a steady base volume first, then add emphasis on purpose.
Too-quiet phrasesUnder 10%: fine. 10-25%: some trailing off. Over 25%: frequent drops.Keep enough breath for the last word of each sentence; that is where volume usually goes.
Speaking timeOver 90%: few pauses, can sound rushed. 70-90%: balanced. Under 70%: lots of silence.Add a deliberate half-second pause after each key point; rehearse openings and transitions if gaps were unplanned.
Pauses of 0.5 s or more per speaking minuteFamous speeches in the EchoPitch study: median about 15.Fewer than 5 usually means you are not giving the audience time to absorb anything.

st = semitones. These bands are practical rules of thumb for presentations, pitches and interviews, not clinical or diagnostic thresholds. Typical speaking pitch differs a lot between people, so the median pitch is shown for your information and is not graded.

Your audio never leaves your device

The analysis happens in JavaScript on your own computer or phone. No recording is saved, nothing is uploaded, and there is no account. The microphone is released as soon as you press Stop or reach 60 seconds. This is different from the full EchoPitch coach, which sends audio to a transcription service so it can also judge your words; that is explained in the FAQ below.

Voice analyzer FAQ

What does an AI voice analyzer measure?

A voice analyzer measures the sound of your voice: pitch (how high or low it is and how much it moves), loudness (how loud you are and how steady), and timing (pauses and how much of the time you are speaking). This tool reports a median pitch, a pitch range in semitones with a monotone flag, loudness variability, the share of phrases that drop away, and pause counts and lengths. Tools that also transcribe your words can add pace and filler words, which is what the speaking rate calculator and filler word counter on this site do.

Is my audio uploaded or stored?

No. The analyzer runs entirely in your browser using the Web Audio API. The microphone signal is processed in memory, a few dozen milliseconds at a time, and turned into numbers: a level, a pitch estimate, speech or silence. No audio file is created, nothing is sent to EchoPitch or any other server, and closing the tab discards everything. You can check this by opening your browser's network panel while you record. The only thing sent is an anonymous analytics event saying the tool was used.

What is a good pitch range when speaking?

As a rule of thumb, speech that moves across about 4 to 12 semitones between its lower and upper pitch (the 10th and 90th percentiles) tends to sound engaged, while a range under about 4 semitones often sounds flat, especially over several minutes. A semitone is one step on a piano keyboard, so 12 semitones is an octave. These are practical guides for presenting, not norms: a calm, authoritative delivery can sit at the lower end, and storytelling naturally runs wider. Compare your own takes rather than chasing a single number.

Can a voice analyzer detect emotions or confidence?

This one does not try. It measures acoustic signals, pitch, loudness and silence, and reports them with plain thresholds. Listeners do use those signals when judging whether someone sounds confident or flat, but the same numbers can come from very different states, so turning them into an emotion label would be guesswork. Treat the results as a mirror for delivery habits you can change, such as a narrow pitch range or sentences that trail off, not as a reading of how you feel.

How accurate is a browser voice analyzer?

Pitch is estimated by autocorrelation, which is accurate to within a few hertz on clean voiced speech. Accuracy drops with background noise, music, a second voice, or a microphone that is far away, and very short recordings give unstable ranges. Loudness figures are relative: the absolute dBFS level depends on your microphone and its distance, so the variability and too-quiet share matter more than the level itself. For the most useful comparison, record in the same place, at the same distance, for 30 to 60 seconds each time.

How is this different from EchoPitch's full analysis?

This free tool measures pitch, loudness and pauses from sound alone, on your device. The full EchoPitch coach records a practice session of up to three minutes, transcribes it, and scores pace, filler words, pauses, vocal variety, facial expression signals and the content and clarity of what you said. That deeper analysis needs a transcript, so the audio is sent to OpenAI Whisper for transcription and the transcript to Anthropic for feedback. If you want a scored report without paying, the free Pitch Scorer analyses a 60-second pitch.

Related free tools and reading

Want feedback on your words too?

EchoPitch scores full practice sessions on pace, filler words, pauses, vocal variety, facial expression signals and the clarity of what you said: 5 sessions for $12.99, paid once, no subscription. Or score a 60-second pitch free first.