← back
State of the Union: Why Local, Why Now — NVIDIA, Osmantic, Roboflow, EXO Labs, @matthew_berman
Takeaway
Local inference becomes practical when capable open models meet sustained workloads, privacy needs, and suitable hardware.
Summary
- The panel connects local AI adoption to simultaneous improvements in open models, agent harnesses, and desktop hardware.
- Always-on agents generate sustained token demand, making predictable local costs and control over sensitive data increasingly attractive.
- Early Llama experiments evolved into practical inference on phones, Macs, and NVIDIA systems; participants cite DeepSeek as a major performance milestone.
- Roboflow highlights vision workloads as a natural local use case because processing often needs low latency near the camera.
local-aiopen-modelsedge-inference
Original description
Alex Cheema's team spent three weeks inside a conference room at NVIDIA headquarters and left with a 10x inference speedup on the DGX Spark, no new computer science required. The update they emailed to Jensen said the wins came from assembling optimizations NVIDIA experts had already solved, pulled together by teams swarming the room all day. The hardware math frames the whole panel: the Spark shares its Grace Blackwell architecture with the data center, Nemotron 3 Ultra runs at 30 tokens per second on four Sparks in the demo room next door, and a four billion parameter Qwen 3.5 on an iPhone now matches what GPT 4o once needed a data center to serve. Joseph Nelson remembers a passenger on his flight whose phone described the seat back in front of them as a printer while a freshly released LLaVA identified it correctly, his proof that no trillion dollar company holds a monopoly on frontier intelligence. Ahmad Osman places the ecosystem in the 1990s of Linux, where the missing piece is point and click onboarding rather than capability, and Matthew Berman sets the bar for mainstream adoption at nothing harder than opening Cursor. The market is already voting: route planning to frontier models and execution to small specialized ones, which is how Coinbase reports exploding token counts on flat costs. Speaker info: Nader Khalil, moderator (NVIDIA): https://x.com/naderlikeladder / naderlikeladder Joseph Nelson (Roboflow): https://x.com/josephofiowa https://roboflow.com Alex Cheema (EXO Labs): https://x.com/alexocheema https://exolabs.net Ahmad Osman (Osmantic): https://x.com/TheAhmadOsman / theahmadosman Matthew Berman (Forward Future): - / @matthew_berman https://forwardfuture.com Timestamps: 0:00 - Welcome to the Local AI Summit 0:40 - Karpathy twice right on keeping up 1:16 - Reasoning models and always on agents 2:32 - Panelist introductions 4:41 - When the inflection point hit 6:36 - GPT 4o quality in your pocket 7:14 - Llama 405B to DeepSeek to GLM 5.2 8:47 - The airplane accessibility story 10:25 - Harnesses give models the real world 11:27 - What language learns from vision 13:19 - A multimodel world in practice 13:57 - Coinbase: tokens up, costs flat 15:03 - Control, sovereignty, no rug pulls 17:50 - Small specialized models and data flywheels 19:45 - A second headquarters inside NVIDIA 21:50 - 10x on the DGX Spark by swarming 24:32 - Desk and data center share an architecture 26:11 - ODS and point and click onboarding 27:14 - Where local still falls short 32:33 - Why finetuning as a service stalled 35:34 - Distillation down to a submarine 39:42 - The biggest open problems in local 42:01 - Open source advocacy and closing