Shaberu for Mac and iPhone
Talk, the way you already talk.
Chinese with English in it, written down as the sentence you said, straight into whatever you were already typing in.
What the recogniser heard
呃 那個 我們這個component等一下要re factor一下
記得把文件更新一下
What Shaberu sent
我們這個 component 等一下要 refactor 一下, 記得把文件更新一下。
macOS 14 or later, and a network connection. 5,000 words a week free, no card.
Anywhere there is a cursor
Nothing to copy, nothing to paste, no window to switch to. The finished text goes into whatever app is in front of you. If you can type there, you can talk there.
- Terminal
- Gmail
- Slack
- Notion
- ChatGPT
- LINE
- Messages
Built for the way people here actually talk
Half of what gets said in a Taiwanese office is Chinese with English nouns dropped into it, like 這個 component 要 refactor. Most dictation decides that sentence is one language and translates the other half away. Shaberu keeps it as the sentence you said, and writes the Chinese the way people here write it: the recognition, the wording, the spacing between the two scripts and a final polish all run on our server before the text comes back, which is why it keeps getting better without you updating anything. The recording is transcribed and dropped in the same request, and never written to disk on our side.
Compared
What dictation usually does, and what this does instead
| Shaberu | The usual dictation app | |
|---|---|---|
| Two languages in one sentence | Kept as the sentence you said, with the English left in English | Decided to be one language, and the other translated away |
| Punctuation and spacing | Punctuated, with a space either side of the English | One long transcript with the "um"s left in |
| Your own vocabulary | Handed to the recogniser before it listens | A find-and-replace afterwards, if anything |
| What you are charged for | Words of finished text, per character in Chinese and per word in English | Minutes of audio, including the ones you spent thinking |
| The recording | Transcribed and dropped inside the same request | Usually kept on a server for a while |
| Two devices | One account, one vocabulary, one set of settings | Usually set up once per device |
The right-hand column is what this kind of app commonly does, not what any one site does. If a line of it is wrong, write to us and we will change it.
Say what you want done with it, mid-sentence
Ask for a new line, a new paragraph or a list while you are still talking and it is carried out rather than typed. The instruction never reaches the page, and the sentence around it does. It is the difference between dictation that transcribes you and dictation that takes notes for you.
分點列出來 先確認預算 再排時程 最後找人
- 先確認預算
- 再排時程
- 最後找人
Tap to start, tap to send
There is no hold-to-talk. Tap right Shift and a capsule appears at the bottom of the screen: cancel on the left, the live waveform and a running clock in the middle, send on the right. Tap again to send, Esc to throw it away. A quiet room draws a flat waveform, so the number that keeps moving is what tells you it is really listening.
Your names and your jargon
Add them once and they reach the recogniser before it listens, rather than being swapped in afterwards by a rule that cannot hear anything. Both devices read the same list.
- Shaberu
- walkccc
- OpenCC
- Cloudflare
- 台積電
- appkit
On iPhone it is a keyboard
Switch to Shaberu with the globe key, talk, and the finished text is typed into the field you were in. The same account and the same settings as the Mac.
Coming to the App Store.
Everything else
- The shortcut can be either Shift, Option, Command, Control or Fn. While it is held, touching any other key, the mouse or the scroll wheel is read as a chord and the dictation is abandoned, so choosing Shift does not break capital letters.
- Say "new line", "new paragraph" or "make that a list" mid-sentence and it is obeyed rather than typed.
- Chinese and English are detected as you speak. A one or two second phrase is where detection is worst, so the language can be pinned in Settings at the cost of mixing on that dictation.
- Text arrives by paste, with the clipboard borrowed and put back exactly, images and rich text included, or keystroke by keystroke where an app refuses a synthetic paste.
- Polishing can be turned off, and takes extra instructions: "I write technical docs, keep it formal, no exclamation marks." Off, the wording, the punctuation and the spacing between scripts still run, because those are rules rather than a model.
- A microphone that is unplugged, muted or held by another app is reported within a second, rather than after you have finished talking to nothing.
- The last 200 transcripts are kept in plain text on the machine itself, so text that landed in the wrong window can be copied back out.
- Optionally starts at login, and adds a line to the menu when there is a new version. Updates are announced, never installed behind you.
- One account across the Mac and the iPhone. Sign in with Apple, and that is the whole of the sign-up.
The recording is not kept anywhere
Your audio goes to our server, is transcribed, converted and polished, and the text comes back. It exists only in memory for the length of that one request. It is never written to disk, and there is no storage bucket it could be written to.
- Transcripts are kept on your account by default so both devices show the same history. Turning off Sync history stops the server writing them at all, rather than writing then deleting.
- Recognition is performed by OpenAI and polishing by Anthropic, as processors on our behalf. Neither trains on data submitted through their APIs.
- Deleting your account removes the transcripts, the vocabulary, the settings and the device list immediately and irreversibly.
Plans
The free column is deliberate
Free
US$0
- 5,000 words of finished text a week, resetting every Monday
- gpt-4o-mini-transcribe, polished by Claude Haiku 4.5
- Your vocabulary, both devices and every feature, all of it
- No card, and no expiry
Pro
US$7.99/ month
or US$59.99 a year, which is US$5.00 a month
- Unlimited words, under a fair-use ceiling set far above human speech
- gpt-4o-transcribe, polished by Claude Sonnet 5
- Worth paying for on long sentences, a strong accent, and switching scripts several times in one breath
- Cancel any time. Pro stays until the period you paid for ends
Upgrading happens in the app: a button on the capsule the moment a dictation runs out of words, or Settings, General, Account.
FAQ
Fair questions
Does it work offline?
No. Recognition, the wording and the polish all run on the server, and there is no degraded path on a plane. That is the trade the design makes on purpose: a word it gets wrong is fixed by a deploy the same day, rather than by an app update everyone has to install.
Is my audio stored?
No. It lives in memory for one request and is gone when the text comes back, and there is no bucket for it to be written to. Transcripts are a separate thing: kept on your account so both devices agree, switchable off, and deletable in one action.
Which apps does it work in?
All of them. Shaberu types into whatever is in front of you, so Gmail, Slack, Notion, Messages, the terminal, ChatGPT and LINE are the same case. If there is a cursor, there is somewhere for the text to land.
Why does it need Accessibility permission?
It is what lets Shaberu see a shortcut key pressed in another app and put text into that app. macOS offers no other way. You grant it once, under System Settings, Privacy and Security, Accessibility.
How is the free allowance counted?
In words of finished text rather than minutes of audio, so a pause, a false start, or a recording your microphone never picked up costs nothing. Chinese counts per character and English per word. A minute of ordinary speech is about 200 characters or 150 words, which makes 5,000 words roughly half an hour a week.
What does the Chinese come out like?
Like something you would have typed yourself: punctuated, with a space either side of any English, and in the wording people in Taiwan use. There is no setting for it and there does not need to be.
Does it run on an Intel Mac?
Yes. Nothing runs on your machine, so there is no Apple Silicon requirement. macOS 14 or later is the whole of it, plus a network connection.
What about the iPhone?
It is a dictation keyboard with the same account, the same vocabulary and the same pipeline, reached with the globe key instead of a modifier key. It is coming to the App Store.