Perplexity's Portable Computer Puts the Whole Agent Stack on Your Desk: Local-First AI on Nvidia's DGX Spark
Perplexity's new Portable Computer runs the full agent harness, orchestrator, and sandbox entirely on an Nvidia DGX Spark — zero per-token costs for local steps, with a PII-checked escalation gate to 15+ cloud models only when you approve it.
When Perplexity launched its cloud-based “Computer” in February 2026, the pitch was a single AI system that researches, codes, deploys, and manages projects end to end. Six months later, the company has flipped the architecture inside out. Portable Computer, launched this week in partnership with Nvidia, moves the entire agent stack — the harness, the orchestrator, the planner, the tool router, the sandbox, and the post-trained models — onto a machine that sits next to your monitor: the NVIDIA DGX Spark.
The result is a “local-first” agent, and the economics are as unusual as the design. Every task begins on the device, and work handled by local models carries no per-token charge at all. For anyone who has watched an autonomous agent burn through dollars of API credit on a single repo-scale migration or a long verification loop, that single line item changes what becomes economically rational on hardware you already own.
Not a local chat app — a packaged agent system
The critical distinction is what actually ships on the device. Portable Computer is not a local LLM with a file picker bolted on. Perplexity packages the local model, the inference engine, the agent harness, an OS-enforced tool sandbox, and app connectors as one system — eliminating the usual weekend project of standing up an inference server and wiring tools by hand.
Users choose between Qwen 3.8 27B and PPLX 27B, Perplexity’s own post-trained variant tuned specifically for its harness. NVIDIA’s Nemotron 3.5 Lightning, an open 30B mixture-of-experts model, is listed as coming soon, and bring-your-own-model setups with a custom inference server are also supported.
Code and tool calls execute inside an OS-enforced sandbox that restricts processes, filesystem paths, and network access. Notably, if the sandbox is unavailable, tool execution is disabled outright rather than silently downgraded — a defensible security posture that treats the sandbox as a hard prerequisite, not an optional feature. Gmail, Outlook, Slack, and GitHub connectors all route through the local orchestrator.
The escalation gate is the real design decision
Local-first, importantly, is not local-only. When a step genuinely needs the live web or frontier-grade reasoning, the orchestrator stops and asks. Before any cloud call goes out, the harness selects the relevant context, runs a PII classifier over it, and shows the user exactly what would leave the machine. Only then does the approved step route to one of 15+ cloud models — and the remote adviser returns text guidance only, never receiving direct access to local files, tools, or the ongoing conversation.
This per-step consent model is the most consequential design choice in the product. It splits the difference between the privacy purists (who want nothing to ever leave the machine) and the pragmatists (who know a 27B model won’t match frontier reasoning on hard steps). The data that leaves is minimal, inspected, and approved — a structure aimed squarely at finance, legal, healthcare, government, defense, and IP-heavy engineering, anywhere data residency rules or contractual confidentiality effectively block cloud inference.
Perplexity also did honest engineering around small-model limits. Qwen 3.8 27B advertises a 260K-token context window but degrades past roughly 100K, so the harness keeps the system prompt and toolset small, loads specialized skills on demand, exposes connectors as compact CLI tools instead of full MCP definitions, and compacts stale context mid-run. It’s a catalogue of context-engineering tricks that any local-agent builder can learn from.
The benchmarks — including the honest one
On Perplexity’s 53-task Local Knowledge Work Bench (spanning deep research, financial analysis, and document creation, which the company says it plans to open-source), Computer running Qwen 3.8 27B on a DGX Spark scored 82.6%, against 77.6% for the open-source Pi harness and 74.0% for Hermes on the identical model. PPLX 27B raised the score to 85.4% — evidence that post-training tuned to your own harness is worth real points.
On BrowseComp, Computer hit 66.7% versus 50.2% for Pi and 43.9% for Hermes, while using 51% less wall time and 70% fewer tokens than Pi. On ParseBench-100 for visual document understanding, it scored 65.1% against 34.6% and 13.9%.
The most informative number is the hybrid result. On Terminal Bench 2.1, the fully local run scored 59.6% at effectively zero marginal cost. Escalating to the cloud adviser lifted that to 73.0% at roughly $0.415 per rollout — compared with 82.4% at about $0.65 for Claude Opus 5 running alone. Escalation narrows the gap to frontier models without closing it. That framing — local as the floor, cloud as an approved, priced, opt-in boost — may be the honest template for hybrid agents in 2026.
Hardware reality check
None of this is free. DGX Spark installs require the GB10 superchip, 128 GB of unified memory, and at least 1 TB of storage — a machine that launched at $3,999, with supply-driven price adjustments reported since. The Qwen 3.8 27B orchestrator ships at 3-bit quantization as a 17.4 GB download requiring 32 GB of RAM; Nemotron 3.5 Lightning is 4-bit, 19 GB, needing 36 GB. Alternatively, DGX OS or Ubuntu on ARM or x64 with any RTX GPU carrying 24 GB+ of VRAM works, and installation is a standard apt repository add.
Availability is Linux-first, for Pro, Max, Enterprise Pro, and Enterprise Max subscribers — with Windows following in September and, pointedly, macOS not on the roadmap. Only a single DGX Spark is supported at launch; clustering is roadmap, not shipped.
Why this matters
Portable Computer is the third act in Perplexity’s 2026 arc: cloud Computer (February), Computer for Builders (August 12), and now a local-first build. But it’s also part of a bigger industry shift. Nvidia gets to position DGX Spark as the personal AI supercomputer for the agent era — a box that makes you a repeat customer for silicon rather than tokens. Perplexity gets an answer to the enterprise privacy objection that has slowed agent adoption in regulated industries. And users get something genuinely new: an agent whose default is silence, whose cloud calls are audited, and whose marginal cost for local work is zero.
The bet is that in 2026, the winning agent isn’t the smartest one in the cloud — it’s the one you can actually trust with your documents. Portable Computer is the most complete attempt yet to build that.
Sources
- [1] https://www.perplexity.ai/hub/blog/introducing-portable-computer-for-local-first-ai
- [2] https://www.perplexity.ai/hub/blog/a-local-first-agent-for-private-and-cost-effective-knowledge-work
- [3] https://www.marktechpost.com/2026/08/25/perplexity-ships-portable-computer-on-nvidia-dgx-spark-local-harness-os-enforced-sandbox-and-zero-per-token-cost-for-local-steps/
- [4] https://venturebeat.com/infrastructure/perplexity-partners-with-nvidia-to-launch-portable-computer-a-fully-local-ai-agent-with-zero-token-costs
- [5] https://www.zdnet.com/article/portable-computer-perplexity-local-ai-agent/
- [6] https://www.perplexity.ai/hub/products/portable-computer