Performance & Memory
MimicScribe runs speech recognition, speaker diarization and echo cancellation on your Mac. Here is what that costs, with the measurement behind each number.
Every table below says when it was measured and on what. A figure I can’t trace to a real measurement isn’t on this page.
How much memory does it use when idle?
Release build, M1 Max (10-core, 32 GB), macOS 26.6, footprint -p at 90 seconds idle against a real database. Measured 2026-09-14.
| Build | Physical footprint |
|---|---|
| Current | 145 MB |
| Before on-demand model loading | 210 MB |
The drop is real: models aren’t loaded until something needs them. The speaker-detection model loads when a meeting starts; the text model for search loads the first time something is embedded. Sitting in your menu bar, MimicScribe is holding very little.
What about the ML models — don’t they take 500 MB?
They occupy about 583 MB, and it is not charged against the app the way ordinary memory is.
The speech model runs on the Neural Engine, Apple’s dedicated ML accelerator. macOS maps its weights as a clean, file-backed region: it counts toward RSS, but the system can reclaim it at any time without asking, and it is not what macOS looks at when deciding which app is under memory pressure. It also does not compete with your other apps for RAM the way a heap allocation does.
If you run footprint yourself you will see this region listed separately. The phys_footprint line at the bottom is the number macOS actually charges the app.
How much memory does an active meeting use?
I don’t have a current measurement to publish. The figure that used to sit here was taken before models moved to on-demand loading, so it no longer means anything.
What I can tell you is how the heaviest moment is bounded. It comes just after you stop a long meeting, when speaker diarization runs over the whole recording. That work is deliberately capped: audio is processed in streamed, memory-mapped chunks instead of being held in memory, the two channels are handled one at a time, and under system memory pressure the app sheds caches and shrinks its batch sizes mid-run. A three-hour meeting finishes within a bounded footprint rather than scaling with its length.
How much CPU does a meeting use?
Release build, M1 Max, idle desktop, offline replay of a 432-second field recording — on-device speech recognition, diarization and echo cancellation running together, no network calls. Sampled over 45 seconds of steady state. Measured 2026-08-23.
About 34% of one core, of which echo cancellation is 17 points. On a 10-core machine that is roughly 3% of the total CPU.
The heavy ML inference is on the Neural Engine, which is why the CPU number is a third of a core rather than several. Echo cancellation is the largest CPU consumer because it runs continuously on every audio frame, unlike speech recognition which works in bursts.
What about battery?
The Neural Engine is built for low-power inference and draws from a different power budget than the CPU and GPU, so the ML work is cheaper than its size suggests. The sustained cost is the audio pipeline — capture and echo cancellation — at the CPU figure above.
I haven’t measured battery life directly. When I do, the number will be here with its method.
Does it do anything when I’m not recording?
Yes, three things, all small and all network:
| What | How often |
|---|---|
| Checking for app updates | Every 24 hours |
| Validating your license | Every 6 hours |
| Reporting usage counts against your plan | Every 15 minutes while the app is running |
No audio is captured and no transcription runs when you are not recording. See Network Activity for everything the app connects to and why.
How long does the first launch take?
The first time the speech model runs, macOS compiles it for the Neural Engine. That compile takes about 26 seconds and happens once — the result is cached, and every launch after that loads the model in 135–235 ms.
Measured 2026-09-14, M1 Max, empty compile cache.
How much disk does it use?
| Data | Size |
|---|---|
| Models, downloaded on first run | ~1.4 GB |
| Database — transcripts, meetings, speaker profiles | Grows with use; tens of MB is typical |
| Local backup snapshots | Roughly 10× the database, rotated |
| Meeting audio | 14.4 MB per hour, deleted after 30 days by default |
The model download is one time. See Privacy & Data for what each model is, and Audio Playback & Retention for changing how long recordings are kept.
Can I measure this myself?
Yes. With MimicScribe running:
footprint $(pgrep mimicscribe) That prints the full breakdown, including the Neural Engine region listed separately from the app’s own memory. phys_footprint at the bottom is the figure macOS charges the app, and the one comparable to Activity Monitor’s Memory column.