Private by default
Prompts, photos and voice are processed on the device. Nothing is uploaded unless your app chooses to send it.
Run language and vision models on phones, laptops and edge hardware. Faster responses, zero cloud bills, data that stays put.
Runs on the chips your users already carry
Ship the same model to iOS, Android, desktop and edge boxes, tuned for the hardware each one has.
Prompts, photos and voice are processed on the device. Nothing is uploaded unless your app chooses to send it.
On a plane, in a basement or on a factory floor. No signal, same answers.
Picks the fastest unit on each device at runtime.
Quantization shrinks model weights. Pick a size to see what lands on the device.
Push a better model to every install over the air, with staged rollouts and one-click rollback.
Enter your own numbers. On-device inference has no per-request server cost.
Estimated cloud inference spend per month
$6,480
Open-weight language, vision and speech models, including the Llama, Gemma, Qwen and Whisper families. You can also bring your own fine-tuned weights.
iOS, Android, macOS, Windows, Linux and modern browsers through WebGPU. Edge boxes such as NVIDIA Jetson and Raspberry Pi work too.
The runtime is small. Models download after install, so the binary you submit to the app stores stays light.
No prompts and no outputs. If you opt in, we collect anonymous performance numbers such as load time and tokens per second.
Free during early access. After that, a flat price per app. Never per request.
Early access is open for teams building on mobile, desktop and edge hardware.
We'll send your SDK key to .