← back

State of the Union: Why Local, Why Now — NVIDIA, Osmantic, Roboflow, EXO Labs, @matthew_berman

26.8K views · Jul 11, 2026 · 44:29 min · Watch on YouTube ↗
Takeaway

Local inference becomes practical when capable open models meet sustained workloads, privacy needs, and suitable hardware.

Summary

  • The panel connects local AI adoption to simultaneous improvements in open models, agent harnesses, and desktop hardware.
  • Always-on agents generate sustained token demand, making predictable local costs and control over sensitive data increasingly attractive.
  • Early Llama experiments evolved into practical inference on phones, Macs, and NVIDIA systems; participants cite DeepSeek as a major performance milestone.
  • Roboflow highlights vision workloads as a natural local use case because processing often needs low latency near the camera.
local-aiopen-modelsedge-inference
Original description
Alex Cheema's team spent three weeks inside a conference room at NVIDIA headquarters and left with a 10x inference speedup on the DGX Spark, no new computer science required. The update they emailed to Jensen said the wins came from assembling optimizations NVIDIA experts had already solved, pulled together by teams swarming the room all day. The hardware math frames the whole panel: the Spark shares its Grace Blackwell architecture with the data center, Nemotron 3 Ultra runs at 30 tokens per second on four Sparks in the demo room next door, and a four billion parameter Qwen 3.5 on an iPhone now matches what GPT 4o once needed a data center to serve.

Joseph Nelson remembers a passenger on his flight whose phone described the seat back in front of them as a printer while a freshly released LLaVA identified it correctly, his proof that no trillion dollar company holds a monopoly on frontier intelligence. Ahmad Osman places the ecosystem in the 1990s of Linux, where the missing piece is point and click onboarding rather than capability, and Matthew Berman sets the bar for mainstream adoption at nothing harder than opening Cursor. The market is already voting: route planning to frontier models and execution to small specialized ones, which is how Coinbase reports exploding token counts on flat costs.


Speaker info:
Nader Khalil, moderator (NVIDIA):
https://x.com/naderlikeladder
  / naderlikeladder  

Joseph Nelson (Roboflow):
https://x.com/josephofiowa
https://roboflow.com

Alex Cheema (EXO Labs):
https://x.com/alexocheema
https://exolabs.net

Ahmad Osman (Osmantic):
https://x.com/TheAhmadOsman
  / theahmadosman  

Matthew Berman (Forward Future):
-    / @matthew_berman  
https://forwardfuture.com

Timestamps:
0:00 - Welcome to the Local AI Summit
0:40 - Karpathy twice right on keeping up
1:16 - Reasoning models and always on agents
2:32 - Panelist introductions
4:41 - When the inflection point hit
6:36 - GPT 4o quality in your pocket
7:14 - Llama 405B to DeepSeek to GLM 5.2
8:47 - The airplane accessibility story
10:25 - Harnesses give models the real world
11:27 - What language learns from vision
13:19 - A multimodel world in practice
13:57 - Coinbase: tokens up, costs flat
15:03 - Control, sovereignty, no rug pulls
17:50 - Small specialized models and data flywheels
19:45 - A second headquarters inside NVIDIA
21:50 - 10x on the DGX Spark by swarming
24:32 - Desk and data center share an architecture
26:11 - ODS and point and click onboarding
27:14 - Where local still falls short
32:33 - Why finetuning as a service stalled
35:34 - Distillation down to a submarine
39:42 - The biggest open problems in local
42:01 - Open source advocacy and closing