New Research
Every claim comes with numbers.
We publish our benchmark results and engineering deep-dives openly. On-device inference - fast, private, hardware-native.
Benchmarks · Apple M4 Max
LLM Decode
higher is betterTime to First Token
lower is betterSpeech-to-Text
lower is betterSpeech-to-Speech
higher is betterQHexRT
NPU inference for all Qualcomm Hexagon devices. First benchmarks on Snapdragon silicon.
MetalRT
Custom kernel inference engine for Apple Silicon. Record-setting LLM, speech, vision, and speech-to-speech performance.
MetalRT Now Does Speech-to-Speech. 1.52x Faster Than mlx-audio.
Read the benchmarks123 tok/s
S2S throughput
MetalRT Now Runs Vision Language Models. Fastest on Apple Silicon.
Read the benchmarks287 tok/s
vision decode
The First Complete AI Inference Engine for Apple Silicon. Now with Speech.
Read the benchmarks101ms
STT latency
We Built the Fastest LLM Decode Engine for Apple Silicon.
Read the benchmarks658 tok/s
LLM decode
FastVoice
End-to-end on-device voice AI. Co-scheduled inference for sub-100ms first-audio latency.