On-Device & Open-Weight AI
AI That Runs on Your Device
Until 2024, running a capable AI model required cloud infrastructure — a GPU cluster serving inference requests over an API. That assumption has broken. Small models run with near-instant latency on consumer laptops and smartphones, and they now handle a growing share of routine tasks — summarisation, drafting, transcription, classification — that used to require a cloud call. Apple Intelligence, Google Gemini Nano, and the small open-weight model families are production deployments, not experiments. The implications for privacy, latency, cost, and data sovereignty are significant.
- What "on-device" means: the model weights and inference run locally — no network request, no data leaving the device
- Scale: small Llama 4 variants, Gemma 3, the Phi small-model line, Gemini Nano — models small enough for a phone but capable enough for real tasks
- Apple Intelligence: on-device models for writing assistance, summarisation, and image generation — runs entirely on the A-series chip
- Privacy benefit: sensitive data (medical records, legal documents, personal messages) never leaves the device for processing
- Latency benefit: no network round-trip — responses in milliseconds, not hundreds of milliseconds
Open-Weight Models: AI You Can Run Yourself
Open-weight models — where the trained model weights are publicly released — have transformed the AI landscape. The Llama 4 family, the DeepSeek V3/R1 line, Qwen 3, Gemma 3, and the Phi small-model line can be downloaded, run, and fine-tuned without API costs or vendor lock-in. The honest framing of where they stand: open-weight models now trail the closed frontier by months rather than years — and for a wide range of production tasks, that gap no longer matters.
- Open-weight ≠ open-source: weights are published but training data and code may not be — "open" refers to the weights specifically
- Llama 4 family (Meta): open-weight models spanning on-device sizes to datacentre scale, with a permissive commercial licence
- DeepSeek V3/R1 line: open-weight models — including RL-trained reasoning — that dramatically narrowed the gap with the closed frontier
- Qwen 3 (Alibaba) and Gemma 3 (Google): strong multilingual open-weight families; Phi (Microsoft) continues the small-model line for devices and private deployment
- Fine-tuning economics: organisations can customise open-weight models on their own data using modest GPU resources — no full pre-training required
When On-Device or Open-Weight Is the Right Choice
Cloud API models are not always the right answer. For privacy-sensitive data, air-gapped environments, high-volume inference economics, or regulatory reasons, on-device or self-hosted open-weight models may be the better architecture. Understanding the tradeoffs prevents both under- and over-engineering.
- Choose on-device when: data is too sensitive to send to the cloud, latency must be sub-50ms, the device has intermittent connectivity
- Choose open-weight self-hosted when: regulatory requirements prohibit third-party data processing, volume makes API costs prohibitive, customisation requires fine-tuning
- Choose cloud API when: you need frontier capability (reasoning models, largest context windows), fast iteration matters more than cost at current scale
- Hybrid pattern: on-device or open-weight for sensitive or routine tasks; cloud frontier models for complex tasks that need maximum capability
- Cost signal: at high token volume (billions per month), self-hosting an open-weight model often reaches payback within 6 months vs. API pricing
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.