Glossary

Speech-to-text

The technology that turns spoken words into written text.

Speech-to-text (STT, also called automatic speech recognition) is technology that converts spoken audio into written text. It listens to sound, models the patterns of speech, and produces the words you said — the engine underneath dictation, voice notes, captions, and transcription.

Speech-to-text can run in the cloud, where audio is uploaded to a server, or on-device, where it is processed locally so the audio never leaves your phone. The on-device approach is the more private one, and it can work offline.

In Notabe, on-device speech-to-text is how voice notes become searchable text — the audio stays on your iPhone or iPad, and only the resulting words are titled, tagged, and made findable. See also our guide on on-device vs cloud transcription.