This Speech-to-Text App Is 16.9 MB. Your Keyboard Is Bigger. Here's What I Found.
In this article
Bottom line: Cactus Compute's Whistle is a speech-to-text model that fits in 16.9 MB and runs on a CPU, with no GPU and no cloud call. Per the Hugging Face model card, it is Apache-2.0 licensed.
Cactus reports a 4.31% word error rate on LibriSpeech test-clean and 11.1 ms to first token on an Apple M4 Pro, against 73.2 ms for Whisper base.
Those figures are self-reported. The tradeoffs are real: seven languages and 30 seconds of audio per pass.
Still, the "speech recognition needs a data center" assumption looks weaker than it did a month ago.
Stop calling AI models "big." Most of them aren't big. They're fat, and the industry has spent three years mistaking fat for capability.
I'm saying this because of a 16.9 MB file that landed on Hacker News this week. That's smaller than a single burst of photos from your phone.
It turns speech into text, and it does it on a CPU, offline, with no subscription attached.
I haven't run my own benchmarks, so everything below comes from the published model card, third-party listings, and the usual caveats.
But the numbers are strange enough that the caveats don't kill the story.
The Sacred Cow: Good AI Needs Big Compute
Here's the belief. Speech recognition is hard, so the good stuff lives on a server somewhere.
You send your audio up, a GPU chews on it, and text comes back, with your voice stored somewhere along the way.
I get why people believe this. For years it was true.
The models that made voice dictation feel like magic arrived alongside the "scale is everything" era, and the lesson everyone absorbed was that more parameters means better results.
Every launch deck since has been a bigger number.
The other half of the belief is that on-device means worse. Offline dictation has historically meant slow, clumsy, and wrong in funny ways.
If you've ever shouted "CALL MOM" at an old phone and gotten a text to your landlord, you know.
That was a fair read of the past. It's a lazy read of the present.
The Evidence
The size
The Hugging Face card describes Whistle as a 16.9 MB model running on the same CPU-first engine Cactus built for its Needle model. No GPU is required.
Mixpeek's listing puts the release in September 2026 under Apache-2.0, while one comparison write-up dates it to October 2. The sources disagree on the date, so treat the exact day loosely.
Apache-2.0 matters more than it sounds. You can ship it inside a commercial product without a lawyer-shaped headache.
The accuracy
Cactus reports a 4.31% word error rate on LibriSpeech test-clean and 10.49% on test-other. In plain English, on clean read-aloud audio it gets roughly 96 of every 100 words right.
On the messier test-other set it drops to about 90 of 100.
Two caveats. LibriSpeech is audiobook-style speech, which is much friendlier than a crowded café or a car with the windows down.
And these are the vendor's own numbers, so wait for independent testing before you bet a product on them.
The speed
This is the number that made me sit up. Cactus reports 11.1 ms to first token on a 10-second clip on an Apple M4 Pro. It lists Whisper base at 73.2 ms on the same task.
That's about 6.6 times faster, from a model small enough to email. Once latency drops that low, dictation stops feeling like a request you submit and starts feeling like a keyboard.
The footprint
There's already a Railway template that wraps Whistle in an OpenAI-compatible transcription API. It describes the service as running in around 120 MB of RAM.
That's a transcription endpoint that fits on the cheapest server tier you can rent.
The Real Problem Nobody Talks About
Whistle isn't the story. The story is how much of the size we've been tolerating was never about quality.
Think about what "big" buys a company. A model that lives on a server is a model you have to keep paying for. Every minute of audio you transcribe becomes a line item on a bill you never see.
It's also a model that sees everything you say.
A tiny local model breaks that business model. There's no per-minute fee because there's no server. There's no data pipeline because the audio never leaves the phone.
The incentive to keep models heavy was never purely technical, and size is also how you keep customers renting instead of owning.
I'll be fair here. Large models exist for good reasons. Accents, jargon, 99 languages, and noisy rooms are hard, and scale helps with all of them.
But "needed for the hard cases" quietly turned into "needed for everything," and almost nobody checked.
I made this mistake myself. I assumed for years that dictation quality was a cloud feature.
I paid for transcription services to turn meeting recordings into text and never once asked whether the job needed a data center.
It's the same pattern behind a lot of subscription fatigue, which I picked at in Why Are You Still Paying For This. We keep renting things we could carry.
Let's Be Honest About the Limits
I promised receipts, and the receipts cut both ways.
According to the comparison write-up, Whistle supports seven languages: English, German, French, Spanish, Italian, Dutch, and Polish. Whisper base covers 99.
If you need Vietnamese, Swahili, or Hindi, this isn't your model.
The model card also lists a 30-second limit per pass. A 90-minute meeting means chopping the audio into chunks, handling the seams, and hoping nobody's sentence gets cut in half.
That's engineering work, not a dealbreaker.
And "4.31% on LibriSpeech" is a lab number. Real life is a toddler screaming over your voice memo. Until independent people throw ugly audio at it, the right posture is interested, not convinced.
None of that changes the shape of the argument. A model this small doesn't have to beat the giants everywhere.
It only has to be good enough for the jobs most people do most days: dictating a text, capturing a note, transcribing a short clip.
What You Should Do Instead
You don't need to rip out your stack tonight. But you can stop assuming the cloud is the default.
1. Audit what you pay to transcribe. If you're sending short clips to a paid API, price out a local model. For voice notes under 30 seconds, you might not need the service at all.
2. Try it where privacy matters. Medical notes, legal dictation, journal entries, anything you'd hesitate to read aloud to a stranger.
A model that never touches the network removes a whole category of worry.
3. Test on your own audio, not a leaderboard. Record your real mess: your accent, your microphone, your noisy kitchen. Run it through Whistle and whatever you use now, then compare the output yourself.
4. Check your own phone's storage screen. Look at what's eating space. The title of this piece is a provocation, but the exercise is real.
Find the largest things on your device and ask what each one actually does for you.
If you build software, the takeaway is sharper. Ask "can this run on the device?" before you ask "which API should we call?" The answer used to be no. For narrow tasks like this, it's increasingly yes.
The Uncomfortable Truth
We've spent years being told that intelligence is something you have to rent.
Speech recognition was supposed to be one of the clearest examples, and now a 16.9 MB file is making that look like a pricing decision rather than a law of physics.
Not every cloud service is a scam, and I'm not telling you to delete your accounts.
But the next time a product tells you it needs your data, your connection, and your monthly fee to do something simple, it's fair to ask what that complexity is actually for.
So here's what I'm left wondering. What else are you paying to send away that could have stayed in your pocket the whole time?