Glossary · How it hears you

Neural EngineANE

The Neural Engine is the part of an Apple silicon chip built to run machine-learning models efficiently, alongside the CPU and GPU. It is why a laptop can run speech recognition continuously without heat, fan noise, or a dead battery.

Every M-series Mac has one. It is a set of cores specialized for the matrix arithmetic that neural networks are made of, and it runs that arithmetic at a fraction of the power a CPU or GPU would need. Speech recognition is the kind of steady, always-on workload the Neural Engine was designed for: a model that has to keep up with audio in real time, for as long as you are talking, without making the rest of the machine sluggish.

Not every model fits. The Neural Engine has its own constraints on model shape and precision, and a model has to be converted for it. That engineering is one of the real differences between local dictation tools that feel effortless and ones that run the same model on the GPU and warm up your lap.

Neural Engine, GPU, and CPU

The CPU can run anything but is the slowest and hungriest at neural-network math. The GPU is fast and flexible and is where most local language models run, through frameworks like MLX or Metal. The Neural Engine is the most efficient of the three for models that fit its constraints, and speech recognition models fit well. A well-built local dictation tool puts each model where it runs best and leaves the CPU alone for the app you are typing into.

Why it is a requirement

Local dictation on an Intel Mac means running the model on a CPU that was not built for it. It works, slowly, and it is a big part of why on-device tools set Apple silicon as the floor. On an M1 or later the same models run in real time with the machine barely noticing, which is the difference between a tool you leave on all day and one you turn on for special occasions.

In Flit

Flit’s models are built for Apple silicon, which is why text lands in about 0.7 seconds and why an M1 or later is the requirement.

Requirements, and the reason for each →

Questions

Fair questions.

What does the Neural Engine do for dictation?

It runs the speech recognition model in real time at very low power, so a local dictation tool can transcribe continuously without heating the machine or draining the battery. That is why local dictation feels effortless on Apple silicon and sluggish on older Macs.

Do I need Apple silicon for local dictation?

For it to feel good, yes. The models can technically run on an Intel CPU, but slowly. Most on-device Mac dictation tools require an M1 or later for that reason.

Your voice was always faster.

7 days free, then $19 USD once. No account, no subscription, nothing uploaded.

macOS 14 or later · Apple silicon