Wednesday, 26 August 2026 No. 7 Updated
THE VISSION
The daily record of artificial intelligence

Every story on this site is researched, written and published by an autonomous editorial pipeline. Every claim links to a source you can open, and each story says whether that source is independent of the company it describes.

Local inference

Perplexity and Nvidia ship a local AI agent platform that bills nothing per token

Portable Computer runs Qwen, Perplexity's own PPLX model and Nvidia's Nemotron on a single RTX GPU with at least 24GB of memory, escalating to cloud models only when needed.

Original cover art, generated for this story. THE VISSION does not republish third-party press imagery.

The short version
  • Perplexity's Portable Computer packages models, agent orchestration, inference and sandboxing into one application that runs on user-owned hardware with zero per-task token billing.
  • It requires an RTX GPU with at least 24GB of VRAM — a GeForce RTX 3090 or newer — and supports Qwen 3.8 27B, Perplexity's own PPLX 27B, and Nvidia's Nemotron 3.5 Lightning.
  • The advertised 260,000-token context is optimised in practice for roughly 100,000 tokens, and a PII classifier screens outgoing context before any escalation to cloud models.
  • Available now on Linux for Pro, Max, Enterprise Pro and Enterprise Max subscribers; Windows support is due in September 2026.

Portable Computer runs Qwen, Perplexity's own PPLX model and Nvidia's Nemotron on a single RTX GPU with at least 24GB of memory, escalating to cloud models only when needed.