Dear AppleVis community,
ToBe SAID, a new local TTS app for macOS is officially here!
It integrates Pocket TTS directly into macOS, allowing you to use custom voices system-wide with accessibility tools like VoiceOver, Spoken Content, and Speak Selection.
ToBe SAID comes with two flexible ways to create personalized voices: Voice Cloning and Voice Design.
With Voice Cloning, you can record your own voice and turn it into a custom TTS engine for reading back text. However, since you can legally only clone your own voice—and many people don't particularly enjoy listening to the sound of their own recorded voice—we also built Voice Design.
Voice Design lets you create completely custom voices simply by describing how you want them to sound in plain text. For example, you can enter a prompt like:
"A gentle, soft-spoken 30s American female voice in an ASMR style. Medium-high pitch spoken in an unhurried, delicate murmur. Smooth and velvety timbre with subtle breathing pauses, clear articulation, and a comforting, attentive tone."
Once generated, you can preview the audio, adjust your description, and dial in a voice that feels comfortable for long reading sessions.
Under the hood, the voice design is powered by Qwen3-TTS. If you need help crafting detailed descriptions, you can even ask ChatGPT or Gemini to generate tailored Voice Design prompts for you.
The free version includes 4 standard built-in voices, with both Voice Cloning and Voice Design available as a preview to test out. Upgrading to the Pro version removes these limits so you can generate and save as many custom voices as you need.
We’d love for you to give it a try.
Free: https://apps.apple.com/app/tobe-said/id6801981584
Pro: https://apps.apple.com/app/tobe-said-pro/id6803422049
Comments
Hungarian
Please add Hungarian as a TTS language!
More language support
Since pocket tts already release it's training program so adding underrepresented language is possible. But it's not cheap and I'm still trying to get funding for this.
Can we eventually get something like this for iPhones?
Can we get something like this for iPhones eventually? I’m asking because I don’t have a Mac. I think it would be fun to try something like this though. I’ve never had any kind of TTS like a customize, but that concept sounds pretty cool. I just use the regular built-in voice right now. Those are still great though.
Hope to have this on iOS
Tobesaid TTS works so well on Android, so I really hope it comes to iOS soon too.
iOS Support
for iOS version, I'm still trying to made it possible. Apple has very strict RAM limitation on iOS tts engine that I cannot solve yet. And if app is possible it's will be more like android version where you can clone but cannot do voice design. Voice design is exclusive to mac.
question
Hi, I have a question: How many voice designs can you do with the free version versus the pro version?
voice design
in free version you can do voice design as many as you can, but it single slot, so it will override older one, and you cannot use it as system tts. so it only preview of it capabilities. while on pro you will have unlimited slot for voice design and you can use it as system tts. it's better you try free version first before deciding to buy pro.
Turkish language support should be added.
It's a very nice synthesizer.However, if Turkish voice is added, users in TĂĽrkiye can also purchase the pro version.
just bought pro
just bought pro. currently using default as s system TTS. ot seems to work well, but inconsistent audio levels when the voices says surtain thibngs i feel is a bug. so far, loving it! did not try free first, as to install pro later would be trickey, as there are parts of the mac app store, which are different in macos 27 beta.
How are you justifying charging for a free project?
Pocket TTS and Quenn are free and open source. Why are you charging for free code? Isn't that illegal or at least immoral? Everybody cares about money. Greedy people.
RE: How are you justifying charging for a free project?
I think that you don't understand several things.
Let's start with that if you want to publish something on the App Store you need a fairly recent Mac and to pay a membership fee of $100/year. There are many projects that end up without any user-distribution (only code that someone compiles) as paying this is non-trivial and may come cumbersome over the years.
Further, the developer has invested an enormous amount of his own work. I have tried to make something like this and failed to do in a short time - not because I plan to earn anything but because no one seems to be interested and my app needs this infrastructure. So I can tell you that you can't do it in an hour, a day and likely not even in a week to do just the basics of this, not a complete project. Mostly because there isn't too much data from Apple's side on how it should be done and you need to try a million times before it works.
But what makes it even worse is that the developer is giving away most of his work for free. His first app, that contains voices is completely free. He is charging only a bit for additional work that he has made and it is completely optional.
And yes I think that moral people would consider paying for this as a donation even if they don't need additional functions to support the project to keep on going. If we talk about morale and ethics.
Re:
And one thing that I forgot - there is the open-source project sherpa-onnx that deals with text-to-speech and speech-to-text.
Based on that there is a Google Play app VoxSherpa made by another developer that uses sherpa-onnx to provide system voices on Android. VoxSherpa shows ads (even full screen) and has in-app purchases for similar things like this app has. Further, Google Play distribution is much cheaper - $20/lifetime and you can have any PC to work with. But besides that, VoxSherpa is officially endorsed by sherpa-onnx with the link on the very top of their home page.
Just to note that I made an official complaint when they collected the location as I think that is too much, but it was fixed. And I could see generally positive comments on VoxSherpa and its commercial model, most complaints were that it needs to double check that some voices may have non-commercial licenses and as such couldn't be distributed under those terms.
We can’t just expect something for nothing
We can’t just expect something for nothing. People work really hard to develop apps and they deserve to be paid for that. I don’t know a thing about app development or programming or anything like that, but if I were going to get something like this, I would be willing to pay for it. You can’t just expect everybody’s work to be for free. Anything service that they are providing an app to use, I don’t know why so many people come out here expecting things to be free. Sure there’s things that I don’t pay for for apps but then I won’t use them. If there’s something that I’m not, I’m going to use I’ll be willing to pay for it. That’s why there’s a free version of this available if people can’t afford to pay for it. And even if it’s free to ask, you know somebody’s paying with their time and effort and as a previous comment, I membership everyone being a developer. So somebody’s paying either way.
iOS vs. Android
- They charge too much on grounds that their app targets blind and visually-impaired users and they therefore have to make it paid in order to continue maintaining it.
- They think we're ignorant fools and they just publish a paid, or even free, app that basically does what other free apps, and often numerous free apps, can do, and then promote their stuff as if it's a truly unique thing that fulfils an essential function. This is problematic even if the app is a so-called mainstream app used by sighted users as well.
- They really have to deal with Apple's greedy policies and make their apps more expensive compared to their counterparts or apps with similar functionality published on Google Play.
I would say this app falls under category 3, even if the models themselves are free, but there are two key points to keep in mind here:Another suggestion
Let me inform everyone in advance that Dennis might just do what he always does and dismiss my proposal before others even notice it, but since this developer is already working on such an app, I would like to request him/her to consider developing a TTS runtime that can run anything and everything from Qwen3-TTS to Supertonic-3, MOSS-TTS (Nano), Kitten TTS (Nano for VoiceOver, or larger variants for audio export), Kokoro TTS, etc. but the app should let us use the voices as the system voice, and be responsive enough to let us use the voices with VoiceOver while also exporting any synthesized speech as audio files, including lossless ones. Speech should be streamed for real-time synthesis, while it can be synthesized in chunks for exporting to WAV or other formats once done. The app should also support voice cloning, and support multiple languages. Start with an app that supports multiple TTS models that are already multilingual, instead of training your own multilingual model, which I wouldn't mind by the way. But it takes more time and effort, and we already have several multilingual models that can generate high-quality speech but are small enough to run on CPU. You can use the Apple Neural Engine for even faster and more power-efficient synthesis.
Re: iOS vs. Android
I would agree with most of this, but payment here is completely optional (correct me if I am wrong, but as far as I have tried that is the case). I don't see the point of discussing whether the developer can ask for an optional payment.
So the app No. 1 that does everything that the Piper app does and some more is completely free. You can't pay even if you want to.
If it were limited by character numbers, or in any other way, one could discuss something though even that wouldn't deserve immediate red flag but I would agree it could be controversial. I don't see how you can discuss whether a completely free app is ethical, exploitative or anything like that...
Yes, but there's a subtle distinction:
There's no such thing as "one free app" and "another separate paid app"; the two apps are essentially the same thing. The paid app offers more functionality, but it's actually a premium version of the free app. So you definitely can pay for the same app if you think of the two apps as practically two versions of one single app. The free app can perform tasks A and B, while the paid app can also perform tasks C and D in addition. If they were two distinct apps rather than basic and upgraded versions of one single app, the free app would perform tasks A and B, while the paid app would only perform tasks C and D, or even if a and B also did overlap, the two apps would do each of these tasks in their own way. And I didn't say that the app being paid was unfair in the first place, and that the developer never had the right to charge us anything for bundling free models within his/her app. My point is that developers can charge us a reasonable price for the effort they put into their apps, and unfortunately also to make enough revenue to cover their own costs thanks to Apple, but they shouldn't exploit our needs and charge us more than they should, or more than their app deserves. PS: Being unable to use the voices as the system voice is a big limitation for me, because this is what enables us to use the voices with VoiceOver or third-party apps like Speech Central. So the free app has an important limitation that really lowers its value for me, and for this reason, I wouldn't regard the paid app merely as an optional purchase that I don't have to make. Another thing is, we shouldn't pay the same price for using the same app on iOS and MacOS, because they won't have the same functionality. Even the paid app on iOS won't let me clone and use my own voice or other voices I record, so it will already have limited functionality compared to the paid MacOS app.
this app is amazing, and i don/t mind the npminal charge
this app is amazing! the AI model for voice design is 6.5 GB. if you want to download it, make sure you have a deesent internet connection. I do, and I have, downloaded the required model for voice design, and prompted it to get a voice I like. it is werth every nicle. would've said penny, but in Canada, pennies are no longer made.
Free vs Paid
Ok, let's do some clarification. Free version you will have 4 voice for system voices and that's it. Other feature like clone and voice design is is preview only, meaning you can try but you cannot add this to system voice. Paid doesn't have limitation. Ethical thing, I put attribution in license page since that is required, I also put many open source project for pocket tts in my github for free, and many of them are related with this project, and I already see some developer use it but didn't give any attribution. So basically if any developer interested, they can make it if they try, the recipe is all there.
my thoughts
Hi,
OK, after playing with the app and the voices, here are my thoughts. First, the voices are amazing. Sadly it does not include any extra pauses such as when speaking the time. Also, it cuts off the year the custom voice design is cool. overall, amazing work..
@Enes Deniz - No Personal Attacks, Please
Enes,
You wrote:
Personal attacks are not permitted on AppleVis. And yes, naming and shaming someone definitely falls under this category. Any further conduct of this nature will be evaluated as trolling and may lead to revocation of your posting privileges.
Model explanation
Just some explanation about the model. There is different license for each model even though all free to use. Piper are free but you must open source the app. That's why you get all those free app using Piper, because it's hard to make money with it, if everyone know the code. Pocket TTS also free you do not need to open source the app but attribution must be written inside the app which I already does. And the code for running pocket tts are everywhere, but most of them packed inside their own app not system wide. There is code for running system wide, but you will be surprised on how slow and unstable it is. Why I know this? I tried those, but original developer has different path and priority, so probability you will get totally free app using those free library is exist if someone pick that up. So how about this app? It has specialized codes that crafted for pocket tts to be very fast and system wide implementation. Due to this other model will not be implemented. Why again? Those model are heavy and has multilingual only for showing, you need to generated 3-4 times to get correct sentence. Their preview will not showing this but after long hour if continuos generation the hallucinations will show up quite often. Pocket TTS already has train code and someone in the community already built those model. But again it work but not for production grade, most of those community project will stop at this "it work" phase. Even official production grade model has problem and I have many report of this. So training by myself, even though expensive and I'm still trying to get funding, will benefit in long term.
Thanks!
@lookbe thanks for the explanation! So the free version also lets us use the voices system-wide. I would definitely love to have this thing on iOS but then again, the pro version should still be cheaper than the macOS one because it won’t have voice design capability. Have you considered quantizing the model by the way?
@Michael Hansen I didn’t know merely mentioning someone else’s attitude and personal attacks directed at me constituted a personal attack itself. Sorry, I will be more careful next time and look for another way to circumvent your collaborative censorship.
model quant
Currently I have pocket tts model that only 57mb in size and run with 110mb RAM, but I'm still unsure if this satisfied RAM requirement on iOS, since after reaching certain RAM usage, it process will just be killed by iOS and replaced by default Samantha voice. I don't have phone to test anymore, and don't have enough app sales to buy iPhone. My initial test using rented phone and rented mac was few months ago using piper model and I know this issue immidiately. Whether this RAM limitation will be vary by devices not tested yet. So my plan will be public beta release for free iOS version, if that was possible. And the price will be the same, android version also have clone only, the reason I don't bring voice design to phone, because it will make your phone unusable for hour, draining battery and heating it up, which is not good.
that's cool!
hi,
oh wow! that's cool! however, when using the voice with VoiceOver, instead of saying the time it says it like this: 650 p.m. September 1, 2026 and then the speech gets cut off. I also noticed that a voice read super fast and not adding extra pauses. For example. If you are in the Messages app, it says something like this: messages messages window two has keyboard focused instead of adding an extra pause.
time and pause
For the time and pauses, it's problem on text parsing side. I'll try to fix this.
@lookbe
model quant
1. The 57mb model that I have now for ios testing is already int4 the lowest quant I can get for onnx.
2. I know this but maybe someday for this implementation, I'm focusing on multilingual support for now.
3. You can find training code here https://github.com/kyutai-labs/pocket-tts/tree/main/training
It already has hindi, czech, korean, and russian model from community, I already tested this but none are production grade for now.
Okay, let me clarify further.
I was talking about the 6.5 GB model that lets users design new voices. Can't you quantize that one?
As for training new models, I'm not a professional coder myself, nor do I have the required dependencies installed on my own computer, but I can probably help with finding datasets as a native speaker of Turkish, as well as other languages and dialects like Arabic, Azerbaijani, Uzbek, Kazakh, etc.
@Enes Deniz - Circumventing the Rules
Hi Enes,
You wrote,
There were no personal attacks directed at you in this thread. You are the one who brought the other member into it by mentioning their name and your unsolicited negative characterization of their behavior.
Everyone is expected to follow our Posting Guidelines in their contributions on AppleVis. You are free to disagree with our rules and openly criticize our moderation, but the rules are the rules.
If you agree to follow our Posting Guidelines going forward, you are absolutely welcome to continue contributing and participating on AppleVis. If you are unwilling to follow the rules of our platform (this includes attempts to circumvent them), please let me know and I will assist with closure of your account.
Interface a tad confusing
The buttons don't seem to be labelled very well. For example "volume high" means play or preview I think. The about button at the top of the form is also weirdly labelled. It took me a while to figure out how to get anything to play my prompt.
It's nice to have more choice, though. I presume I need to download the design thing to use my own custom prompts which sounds fun. Was quite a big download though so will have to wait for now.
Michael
- You're responsible for enforcing the rules, not setting them. We set the rules together, and we should try to stop anyone breaking them together.
- There's no such rule that legitimizes oppression or censorship. I want exact quotes referring to that alleged rule if it does exist, so you'd better give me those rather than quotes from my own posts that I can always read them myself and I'm the one who wrote them in the first place.
If you do manage to base your reaction on any of the existing guidelines, you still have to deal with two more points.- You can't just do something arbitrary and attempt to ground it on the rules just because they exist and aim to legitimize your justifiable actions as a moderator.
- You can't just command me to follow the rules while totally ignoring them when certain other members explicitly violate them and you don't impose them fairly and impartially yourself.
I promise, remember to tell Dennis to follow the very same rules you've been reminding me and I'll comply with the guidelines at least as carefully as you and Dennis. I currently doubt whether they even actually exist in practice, if not at all.Enes Deniz
Hi Enes,
The AppleVis Editorial Team reviews posts to the site 365 days per year, and we do our best to thoughtfully and appropriately address moderation issues as they arise. As outside observers reviewing a large amount of content, it is possible that we may sometimes miss the nuance of an individual comment on first reading, and we cannot know how you perceived an interaction unless you tell us.
If you feel personally attacked or otherwise wronged by another member on AppleVis, rather than trying to handle the situation yourself, I ask that you please contact me or another member of the AppleVis Editorial Team. My email address is [email protected], and you can also report concerns about harassment to [email protected].
Responsibility for general content moderation rests solely with the AppleVis Editorial Team. Members are welcome to bring potential guideline violations to our attention through the proper channels, but we ask that community members not attempt to take moderation action themselves. Determining whether and how the Posting Guidelines apply in a particular situation is the responsibility of the AppleVis Editorial Team. If you have concerns about the conduct of another member, please contact me or another member of the AppleVis Editorial Team. I cannot promise that we will handle every situation exactly as you would if you were in our position, but we will do our best to address each situation thoughtfully and equitably.
Regarding your comment in this thread specifically, as requested, the relevant guidelines are:
Lastly, compliance with our rules is not contingent upon how well you believe those rules are enforced with respect to another member. AppleVis' Posting Guidelines apply equally to everyone. You wrote at the beginning of your comment, “Thanks for the kind offer, but I'll have to decline politely.” I am not entirely clear whether you mean by this that you are unwilling to follow AppleVis' rules going forward. If that is what you mean, the next step would be closure of your account. Please let me know your intentions.
more thoughts on pro version
Hi,
I have more feedback on pro version. first, the custom voice design voice is awesome! It is a bit sluggish when navigating with the VoiceOver modifier keys plus the arrow keys but it's not too bad also, the pitch doesn't change when inputting a capital letter nor deleting a character even though say capital letter and deleting is set to change pitch. Other speech synthesizers do this, such as Eloquence and Vocalizer. other than that, I'm loving it!.