Local inference
Perplexity and Nvidia ship a local AI agent platform that bills nothing per token
Portable Computer runs Qwen, Perplexity's own PPLX model and Nvidia's Nemotron on a single RTX GPU with at least 24GB of memory, escalating to cloud models only when needed.
Original cover art, generated for this story. THE VISSION does not republish third-party press imagery.
The short version
- Perplexity's Portable Computer packages models, agent orchestration, inference and sandboxing into one application that runs on user-owned hardware with zero per-task token billing.
- It requires an RTX GPU with at least 24GB of VRAM — a GeForce RTX 3090 or newer — and supports Qwen 3.8 27B, Perplexity's own PPLX 27B, and Nvidia's Nemotron 3.5 Lightning.
- The advertised 260,000-token context is optimised in practice for roughly 100,000 tokens, and a PII classifier screens outgoing context before any escalation to cloud models.
- Available now on Linux for Pro, Max, Enterprise Pro and Enterprise Max subscribers; Windows support is due in September 2026.
Portable Computer runs Qwen, Perplexity's own PPLX model and Nvidia's Nemotron on a single RTX GPU with at least 24GB of memory, escalating to cloud models only when needed.