User Guide
dicto displays what you say or type as large, clear text. This guide describes how the app works.
Table of contents
Getting Started
The first time dicto is launched users are taken through a brief onboarding phase where a default spoken language and color scheme are chosen. Both can be changed in Settings at any time. These preference screens are followed by three screens of basic instructions on the apps features, which can be skipped.
The first time the Microphone button is pressed, iOS will request permission to access Speech Recognition and the Microphone on your device by presenting the Access Permissions screens. These are both Apple services and dicto needs them to display your speech as text. Tap Allow, or to enable access later open the iOS Settings app, go to Settings ▸ dicto, and turn on Microphone and Speech Recognition.
The Main Screen
The main screen displays the current message and the controls for creating, clearing, and revisiting messages. Portrait orientation presents slightly larger text (good for shorter messages), and landscape orientation presents longer text (good for slightly longer or contextual messages).
- Message display area: Shows on-screen instructions and displays the current message, scaled as large as the available space allows and updating in real time while you speak. Double-tap the area to start or stop dictation. Long-press to edit the current message. Long messages scroll automatically; switch to manual scrolling in Settings.
- Microphone button: Starts or stops a spoken message. See Speaking & Dictation. Press and hold to begin enrollment in Solo.
- Keyboard button: Opens the keyboard to type or paste a message instead of speaking. See Typing.
- Clear button: Clears the current message. The cleared message is not saved to Earlier. A long press on the Clear button brings up a "Start a fresh session?" dialog which resets the app.
- Earlier: A strip of completed messages along the bottom of the screen. An uncleared message is added to the Earlier section when a new message is started. Tap a message to display it again. Earlier is cleared when the app is closed, and is inactive while Whisper Mode is on or a Solo session is running.
- Read aloud button: Reads the current message aloud, using the voice selected in Settings. See Read Aloud.
- Settings button: Opens the Settings screen, where preferences such as language, color scheme, speaking pace, and Whisper Mode are set. See Settings.
Speaking & Dictation
Start dictation by tapping the microphone button or by double-tapping the message display area. Speak normally; words appear as you say them. If the dictation is not working well (sentences not ending properly), there is a Speaking pace setting for faster or slower speakers in the Settings menu that may help.
By default, a pause of about one second ends a sentence, then dicto adds a period or question mark and keeps listening. A pause of about two seconds ends the current speaking turn and displays the final text, scrolling automatically for messages that extend below the main screen. Messages can also be manually scrolled, and auto-scrolling can be turned off in the Settings menu. Adjust how long dicto waits at pauses with Speaking pace in Settings. A Solo session stops automatically after about a minute and a half, so the microphone does not stay open in a noisy room. Solo sentence timings are also set using the Speaking pace setting.
Dictation is provided by Apple's speech recognition and is not perfect. Accents, background noise, uncommon names, very short phrases, and people speaking over each other can be misheard. Speaking clearly at a natural pace, and using the Speaking pace setting that matches the speaker, gives the best results.
To clear the text on the screen press the Clear text at the bottom of the screen. Pressing and holding the Clear text will bring up the Start a fresh session? dialog. Press Start Fresh to reset dicto, and clear the EARLIER section.
Solo
Solo uses voice-isolation technology to focus dicto on one enrolled voice - designed to show that speaker's words and ignore other voices nearby. Solo allows for hands-free operation. Speech from other voices is not displayed and is never transcribed. Solo works best with one person speaking at a time in a quiet to moderate room; when dicto is not sure a voice is the enrolled one, it shows nothing. Phrases under three words that are spoken quickly may be too short to be identified (e.g. "Hello", "That's great"), and may not appear on screen.
How Solo works
While Solo is active, dicto listens to the room and turns each stretch of speech into a small set of numbers that describe the sound of the voice, not the words. Those numbers are compared with the numbers captured at enrollment. Speech that matches the enrolled voice is transcribed and displayed; speech that does not match is discarded without ever being turned into text. The comparison runs on small machine-learning models that ship inside the app and run entirely on the device; no cloud service, account, or connection is involved. No speech is kept or recorded, ever.
Enrolling a voice
- Press and hold the microphone button and speak naturally. What you say during the hold is displayed as a normal message.
- A panel shows enrollment progress. The bar is red until enough voice has been captured for basic sorting.
- At about three seconds of actual speech the bar turns yellow and a tone plays: basic enrollment has succeeded, and the button can be released at any time.
- Continuing to hold the button improves the enrollment, up to ten seconds of actual speech for full enrollment. The bar shifts from yellow toward green as more voice is captured, reaching green at about ten seconds. Full enrollment is indicated by the Solo Glow, a full-screen glow effect.
- Release the button. Solo begins listening for the enrolled voice. Successful enrollment is indicated by the microphone icon changing to the Solo icon: an icon of a person with speech bars with a solid teal circle outline.
If too little voice was captured when the button is released, dicto listens for about two more seconds, then cancels the current enrollment if enrollment still cannot complete.
Enrollment top-up and the Solo Glow
A basic enrollment is enough to start, and dicto quietly finishes the job on its own. During a session, moments of speech that unmistakably match the enrolled voice are added to the enrollment, improving accuracy as the conversation goes on. This is called enrollment top-up; it uses only the enrolled voice, never other speech. When the enrollment reaches its full size - whether by holding the button to the green bar or through top-up during conversation - the Solo Glow appears: a brief, silent, full-screen glow that means dicto has fully enrolled the voice. Nothing needs to be done when it appears.
During a session
The enrolled voice is displayed one message at a time, using the Speaking pace timings. Other voices produce nothing on screen. Earlier is hidden while a Solo session is active. Tap the microphone button to pause or resume listening; the session also pauses on its own after about thirty seconds of silence in the room. A paused Solo session is indicated by the solid teal line becoming a dashed line. To re-activate the current session tap the Solo button.
Ending a session
Press and hold the Clear button and confirm Start Fresh to erase the current text and the enrolled voice; a brief "Session cleared" notice confirms the erasure. A session also ends by itself after about three minutes without the enrolled voice, or when the app is moved to the background. Press and hold the microphone button during a session to enroll a different voice in place of the enrolled one.
The voice profile
Voice matching runs on small machine-learning models that ship inside the app and run entirely on your device. There is no cloud service, no account, and no connection involved; the models compare voices, and nothing more. The enrollment itself is a small set of numbers used to classify a voice, held in memory for the length of the session: never saved, never sent anywhere, and gone when the session ends.
The voice-matching models are the pyannote community-1 speaker models (WeSpeaker family), converted to Core ML by FluidInference and used under the Creative Commons Attribution 4.0 license.
Typing
Tap the Keyboard button to open Type text. Type or paste a message, then tap Done (or press the blue, check button on the keyboard) to display it on the main screen, or Cancel to close without changing the current message.
Whisper Mode
Whisper mode toggles the disappearing messages function. Turn on Whisper mode in Settings to send Wisps: messages that appear on screen, stay for a few seconds, then fade away permanently. Wisps are never saved to Earlier, which is hidden while Whisper mode is on. Auto-scrolling is also turned off while Whisper mode is on.
Read Aloud
Tap the Read aloud button in the top-right corner to have the currently displayed message read out loud. The button is unavailable when no message is on screen. The reading voice is set under Voice in Settings; additional voices can be added in the iOS Settings app under Accessibility ▸ Spoken Content ▸ Voices (not Siri).
Settings
Tap the Settings button (the gear, top-left) to open the Settings menu.
- Whisper mode: Toggles disappearing messages. See Whisper Mode.
- Speaking pace: Fast, Average, or Slow. Adjusts how long dicto waits at pauses to end sentences or messages. If sentences don't separate, set to Fast; if sentences get cut off, set to Slow.
- Solo voice match: Strict, Balanced, or Relaxed. Adjusts how closely a voice must match the enrolled voice to be shown in Solo. Applies from the next enrollment. Choose Strict when voices in the room sound similar.
- Conversation distance (iPad only): Adjusts font size for Near (smaller font) and Far (larger font) conversations.
- Auto-scroll long messages: When on, long messages automatically scroll to the end and back at a readable pace; turn off to stop automatically scrolling messages. Messages can always be scrolled manually.
- Input on open: Spoken opens dicto in listening mode for the first message as a quick conversation partner; Manual opens dicto with the normal interface, allowing spoken or text-based message input for the first message.
- Voice: The voice used for Read aloud.
- Color scheme: Select preferred display colors. System, Light, and Dark options, as well as high-contrast pairs: Black on white, White on black, Yellow on blue, Green on black, Orange on black, Crimson on white, Neon pink on black, and White on green. Every color combination meets the WCAG AA accessibility standard for contrast, and at the large sizes dicto displays, every combination meets AAA.
- Language: The language you speak for dictation (not the app's language). dicto supports 50+ spoken languages, drawn from Apple's on-device dictation. Any language can be chosen and used as long as it is supported by the device, regardless of the user interface language (e.g. Chinese dictation can be used in an English language app interface). See the dictation languages list. The dicto interface itself is available in 14 languages; see interface languages.
- About dicto: App version, a short description, and a link to this site.
Privacy & Data
dicto processes everything on your device. Nothing you say or type is sent to us, and nothing is stored once you close the app. For full details, see the Privacy Policy.
Support
Questions, problems, or feedback? Email support@speechdisplay.com - we're happy to help.