User Guide

dicto displays what you say or type as large, clear text. This guide describes how the app works.

Table of contents

Getting Started

The first time dicto is launched users are taken through a brief onboarding phase where a default spoken language and color scheme are chosen. Both can be changed in Settings at any time. These preference screens are followed by three screens of basic instructions on the apps features, which can be skipped.

The first time the Microphone button is pressed, iOS will request permission to access Speech Recognition and the Microphone on your device by presenting the Access Permissions screens. These are both Apple services and dicto needs them to display your speech as text. Tap Allow, or to enable access later open the iOS Settings app, go to Settings ▸ dicto, and turn on Microphone and Speech Recognition.

The Main Screen

The main screen displays the current message and the controls for creating, clearing, and revisiting messages. Portrait orientation presents slightly larger text (good for shorter messages), and landscape orientation presents longer text (good for slightly longer or contextual messages).

Speaking & Dictation

Start dictation by tapping the microphone button or by double-tapping the message display area. Speak normally; words appear as you say them. If the dictation is not working well (sentences not ending properly), there is a Speaking pace setting for faster or slower speakers in the Settings menu that may help.

By default, a pause of about one second ends a sentence, then dicto adds a period or question mark and keeps listening. A pause of about two seconds ends the current speaking turn and displays the final text, scrolling automatically for messages that extend below the main screen. Messages can also be manually scrolled, and auto-scrolling can be turned off in the Settings menu. Adjust how long dicto waits at pauses with Speaking pace in Settings. A Solo session stops automatically after about a minute and a half, so the microphone does not stay open in a noisy room. Solo sentence timings are also set using the Speaking pace setting.

Dictation is provided by Apple's speech recognition and is not perfect. Accents, background noise, uncommon names, very short phrases, and people speaking over each other can be misheard. Speaking clearly at a natural pace, and using the Speaking pace setting that matches the speaker, gives the best results.

To clear the text on the screen press the Clear text at the bottom of the screen. Pressing and holding the Clear text will bring up the Start a fresh session? dialog. Press Start Fresh to reset dicto, and clear the EARLIER section.

Solo

Solo uses voice-isolation technology to focus dicto on one enrolled voice - designed to show that speaker's words and ignore other voices nearby. Solo allows for hands-free operation. Speech from other voices is not displayed and is never transcribed. Solo works best with one person speaking at a time in a quiet to moderate room; when dicto is not sure a voice is the enrolled one, it shows nothing. Phrases under three words that are spoken quickly may be too short to be identified (e.g. "Hello", "That's great"), and may not appear on screen.

How Solo works

While Solo is active, dicto listens to the room and turns each stretch of speech into a small set of numbers that describe the sound of the voice, not the words. Those numbers are compared with the numbers captured at enrollment. Speech that matches the enrolled voice is transcribed and displayed; speech that does not match is discarded without ever being turned into text. The comparison runs on small machine-learning models that ship inside the app and run entirely on the device; no cloud service, account, or connection is involved. No speech is kept or recorded, ever.

Enrolling a voice

  1. Press and hold the microphone button and speak naturally. What you say during the hold is displayed as a normal message.
  2. A panel shows enrollment progress. The bar is red until enough voice has been captured for basic sorting.
  3. At about three seconds of actual speech the bar turns yellow and a tone plays: basic enrollment has succeeded, and the button can be released at any time.
  4. Continuing to hold the button improves the enrollment, up to ten seconds of actual speech for full enrollment. The bar shifts from yellow toward green as more voice is captured, reaching green at about ten seconds. Full enrollment is indicated by the Solo Glow, a full-screen glow effect.
  5. Release the button. Solo begins listening for the enrolled voice. Successful enrollment is indicated by the microphone icon changing to the Solo icon: an icon of a person with speech bars with a solid teal circle outline.

If too little voice was captured when the button is released, dicto listens for about two more seconds, then cancels the current enrollment if enrollment still cannot complete.

Enrollment top-up and the Solo Glow

A basic enrollment is enough to start, and dicto quietly finishes the job on its own. During a session, moments of speech that unmistakably match the enrolled voice are added to the enrollment, improving accuracy as the conversation goes on. This is called enrollment top-up; it uses only the enrolled voice, never other speech. When the enrollment reaches its full size - whether by holding the button to the green bar or through top-up during conversation - the Solo Glow appears: a brief, silent, full-screen glow that means dicto has fully enrolled the voice. Nothing needs to be done when it appears.

During a session

The enrolled voice is displayed one message at a time, using the Speaking pace timings. Other voices produce nothing on screen. Earlier is hidden while a Solo session is active. Tap the microphone button to pause or resume listening; the session also pauses on its own after about thirty seconds of silence in the room. A paused Solo session is indicated by the solid teal line becoming a dashed line. To re-activate the current session tap the Solo button.

Ending a session

Press and hold the Clear button and confirm Start Fresh to erase the current text and the enrolled voice; a brief "Session cleared" notice confirms the erasure. A session also ends by itself after about three minutes without the enrolled voice, or when the app is moved to the background. Press and hold the microphone button during a session to enroll a different voice in place of the enrolled one.

The voice profile

Voice matching runs on small machine-learning models that ship inside the app and run entirely on your device. There is no cloud service, no account, and no connection involved; the models compare voices, and nothing more. The enrollment itself is a small set of numbers used to classify a voice, held in memory for the length of the session: never saved, never sent anywhere, and gone when the session ends.

The voice-matching models are the pyannote community-1 speaker models (WeSpeaker family), converted to Core ML by FluidInference and used under the Creative Commons Attribution 4.0 license.

Typing

Tap the Keyboard button to open Type text. Type or paste a message, then tap Done (or press the blue, check button on the keyboard) to display it on the main screen, or Cancel to close without changing the current message.

Whisper Mode

Whisper mode toggles the disappearing messages function. Turn on Whisper mode in Settings to send Wisps: messages that appear on screen, stay for a few seconds, then fade away permanently. Wisps are never saved to Earlier, which is hidden while Whisper mode is on. Auto-scrolling is also turned off while Whisper mode is on.

Read Aloud

Tap the Read aloud button in the top-right corner to have the currently displayed message read out loud. The button is unavailable when no message is on screen. The reading voice is set under Voice in Settings; additional voices can be added in the iOS Settings app under Accessibility ▸ Spoken Content ▸ Voices (not Siri).

Settings

Tap the Settings button (the gear, top-left) to open the Settings menu.

Privacy & Data

dicto processes everything on your device. Nothing you say or type is sent to us, and nothing is stored once you close the app. For full details, see the Privacy Policy.

Support

Questions, problems, or feedback? Email support@speechdisplay.com - we're happy to help.

Back to home