On-device AI SDK

AI that never leaves the device.

Run language and vision models on phones, laptops and edge hardware. Faster responses, zero cloud bills, data that stays put.

Try it: private redaction Runs in this tab
Hi, it's NAME. Move my 3pm with the design team to Thursday and send the notes to EMAIL or call PHONE.
Detected intent: Reschedule a meeting
Processed in<1 ms
Sent to a server0 bytes
Items hidden3

Runs on the chips your users already carry

Every cloud call costs you a round trip, a bill and a little trust. Edgewize runs the model where the data already is.

One SDK. Every device.

Ship the same model to iOS, Android, desktop and edge boxes, tuned for the hardware each one has.

Private by default

Prompts, photos and voice are processed on the device. Nothing is uploaded unless your app chooses to send it.

Works offline

On a plane, in a basement or on a factory floor. No signal, same answers.

Uses the right chip

Picks the fastest unit on each device at runtime.

NPUGPUCPU

Small enough to ship

Quantization shrinks model weights. Pick a size to see what lands on the device.

Model parameters
Precision
1.5 GB Weights only, before runtime memory.

Update models without an app release

Push a better model to every install over the air, with staged rollouts and one-click rollback.

Read the FAQ

From model to app in an afternoon


        
Cost calculator

What the cloud is charging you

Enter your own numbers. On-device inference has no per-request server cost.

Estimated cloud inference spend per month

$6,480

$77,760Per year at today's usage
12MServer round trips per month
$0Per-request cost with Edgewize
0Round trips with Edgewize

Questions

Which models can I run?

Open-weight language, vision and speech models, including the Llama, Gemma, Qwen and Whisper families. You can also bring your own fine-tuned weights.

Which devices are supported?

iOS, Android, macOS, Windows, Linux and modern browsers through WebGPU. Edge boxes such as NVIDIA Jetson and Raspberry Pi work too.

How much does it add to my app?

The runtime is small. Models download after install, so the binary you submit to the app stores stays light.

Does any data reach your servers?

No prompts and no outputs. If you opt in, we collect anonymous performance numbers such as load time and tokens per second.

How is it priced?

Free during early access. After that, a flat price per app. Never per request.

Ship AI that works without the cloud.

Early access is open for teams building on mobile, desktop and edge hardware.