The AI Learning Hub Journal

On-Device & Open-Weight AI

On-Device & Open-Weight AI: Three Architectures, One Decision Diagram comparing three AI deployment architectures — on-device, self-hosted open-weight, and cloud frontier API — with their tradeoffs and the decision framework for choosing between them. On-Device & Open-Weight AI: Three Architectures, One Decision On-Device Runs locally on your hardware WHO'S DOING IT Apple Intelligence Gemini Nano · Llama 3.2 Phi-4-mini on edge devices WINS ON Privacy — data never leaves device Latency — millisecond responses Offline — no network required USE WHEN Data sensitivity is highest Sub-50ms latency required Intermittent connectivity Trade-off: capped capability Self-Hosted Open-Weight Your infrastructure, public weights WHO'S DOING IT Llama 3 (8B / 70B) Mistral · Mixtral · Phi-4 Qwen — frontier-competitive WINS ON No vendor lock-in or licence fees Fine-tune freely on your own data Data residency & sovereignty solved USE WHEN Regulation forbids third-party data API cost > self-hosted break-even Fine-tuning is core to the product Trade-off: operate the GPU stack Cloud Frontier API Vendor-hosted, max capability WHO'S DOING IT OpenAI · Anthropic Google · xAI Frontier model APIs WINS ON Frontier capability & reasoning No infrastructure to run Fast iteration via API change USE WHEN You need top-of-frontier quality Volume below break-even Data sensitivity is manageable Trade-off: data leaves your boundary The 2026 default: hybrid On-device or open-weight for sensitive & routine work · cloud frontier for the hard cases that need maximum capability Route by data sensitivity, task complexity, and cost — not by vendor relationship Cloud-only is no longer the default — data sovereignty objections now have technical answers

AI That Runs on Your Device

Until 2024, running a capable AI model required cloud infrastructure — a GPU cluster serving inference requests over an API. That assumption has broken. Small models run with near-instant latency on consumer laptops and smartphones, and they now handle a growing share of routine tasks — summarisation, drafting, transcription, classification — that used to require a cloud call. Apple Intelligence, Google Gemini Nano, and the small open-weight model families are production deployments, not experiments. The implications for privacy, latency, cost, and data sovereignty are significant.

  • What "on-device" means: the model weights and inference run locally — no network request, no data leaving the device
  • Scale: small Llama 4 variants, Gemma 3, the Phi small-model line, Gemini Nano — models small enough for a phone but capable enough for real tasks
  • Apple Intelligence: on-device models for writing assistance, summarisation, and image generation — runs entirely on the A-series chip
  • Privacy benefit: sensitive data (medical records, legal documents, personal messages) never leaves the device for processing
  • Latency benefit: no network round-trip — responses in milliseconds, not hundreds of milliseconds

Open-Weight Models: AI You Can Run Yourself

Open-weight models — where the trained model weights are publicly released — have transformed the AI landscape. The Llama 4 family, the DeepSeek V3/R1 line, Qwen 3, Gemma 3, and the Phi small-model line can be downloaded, run, and fine-tuned without API costs or vendor lock-in. The honest framing of where they stand: open-weight models now trail the closed frontier by months rather than years — and for a wide range of production tasks, that gap no longer matters.

  • Open-weight ≠ open-source: weights are published but training data and code may not be — "open" refers to the weights specifically
  • Llama 4 family (Meta): open-weight models spanning on-device sizes to datacentre scale, with a permissive commercial licence
  • DeepSeek V3/R1 line: open-weight models — including RL-trained reasoning — that dramatically narrowed the gap with the closed frontier
  • Qwen 3 (Alibaba) and Gemma 3 (Google): strong multilingual open-weight families; Phi (Microsoft) continues the small-model line for devices and private deployment
  • Fine-tuning economics: organisations can customise open-weight models on their own data using modest GPU resources — no full pre-training required

When On-Device or Open-Weight Is the Right Choice

Cloud API models are not always the right answer. For privacy-sensitive data, air-gapped environments, high-volume inference economics, or regulatory reasons, on-device or self-hosted open-weight models may be the better architecture. Understanding the tradeoffs prevents both under- and over-engineering.

  • Choose on-device when: data is too sensitive to send to the cloud, latency must be sub-50ms, the device has intermittent connectivity
  • Choose open-weight self-hosted when: regulatory requirements prohibit third-party data processing, volume makes API costs prohibitive, customisation requires fine-tuning
  • Choose cloud API when: you need frontier capability (reasoning models, largest context windows), fast iteration matters more than cost at current scale
  • Hybrid pattern: on-device or open-weight for sensitive or routine tasks; cloud frontier models for complex tasks that need maximum capability
  • Cost signal: at high token volume (billions per month), self-hosting an open-weight model often reaches payback within 6 months vs. API pricing

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.